Tools · Open Biostatistics
Sample size for comparing paired means (before and after)
Enter the mean change you want to be able to detect and the variability of the differences (directly, or from the standard deviation of each measurement and the correlation between them), and get how many pairs are needed, with the power curve, a plain-language interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-medias-pareadas
This link does not include the pasted data: they are too long for a URL.
Results
SD of the differences (σ_d)
1.5
entered
z of the significance level
1.960
z of the power
0.842
Cohen's d (Δ/σ_d)
0.333
Pairs needed
73
non-central t · unrounded: 72.58
Pairs (normal approximation)
73
normal with Guenther's correction · unrounded: 72.56
Pairs adjusted for losses
81
10% · unrounded: 80.65
Power with the available pairs
not applicable
Power with the available pairs (normal approximation)
not applicable
Interpretation
To detect a mean change of 0.5 (in the units of the scale; Cohen's d = 0.333) with 80% power and α = 0.050 (Two-sided), 73 pairs are required.
The standard deviation of the differences was taken directly from what you entered: 1.5.
The normal formula with Guenther's correction gives 73 pairs; the exact non-central t solution, identical to R's power.t.test(type = "paired"), is the one reported and the one the calculation rests on.
Allowing for 10% losses, you should recruit 81 pairs to finish the study with the 73 the calculation requires.
If you already know how many complete pairs you will have, type it into "inverse mode" and you will see what power that figure would reach.
- The figures depend entirely on the assumptions you entered (the change that matters and the variability of the differences): take them from previous studies or a pilot and state their origin in the protocol.
Explanation
In a paired design each participant is their own control: measured before and after, or with the two techniques being compared. What gets analysed is not two groups but a single column of differences, and the one-sample t test asks whether their mean is zero. That is why the sample size is counted in PAIRS, not in measurements.
The quantity that rules is not the variability of the measurements but that of the differences. If the two measurements on the same person are very alike, the differences vary little and very few pairs are needed; if they are nearly independent, the paired design loses its advantage. The link between them is σ_d = σ·√(2(1 − ρ)): at ρ = 0.5 the SD of the differences equals the SD of each measurement, and above 0.5 the pairing starts to pay off.
This calculator accepts both routes. If the paper you start from reports the standard deviation of the differences, type it in: it is the exact figure and it overrides anything else. If it only reports the SD of each measurement, enter it together with the assumed correlation and the calculator derives σ_d. That ρ is an assumption, not a datum, so justify it and stress-test the result with neighbouring values.
The main result is the exact solution based on the non-central t distribution, the same one power.t.test(type = "paired") returns in R. The normal approximation with Guenther's (1981) correction is shown next to it as a teaching row; its correction is half the two-sample one because here there is only one variance to estimate. And since any study will have incomplete pairs, the adjustment for losses divides the total by (1 − L).
Equations
- mean of the differences to be detected
- standard deviation of the differences
- number of pairs
- non-centrality parameter
- non-central t distribution with n − 1 degrees of freedom
- 2 if the test is two-sided, 1 if it is one-sided
- standard deviation of each measurement
- correlation between the two measurements
- normal quantile of the significance level
- normal quantile of the power
- standard normal distribution function
- 2 if the test is two-sided, 1 if it is one-sided
- proportion of expected losses
- pairs adjusted for losses
- ceiling: rounding up is applied only once, at the end
R code
# Sample size for paired means, exact non-central t - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
delta <- 0.5 # mean of the within-pair differences that matters
de_dif <- 1.5 # SD of the DIFFERENCES; 0 = not captured, then it is derived below
sigma <- 0 # SD of each measurement; only used when de_dif is not given (0 = not captured)
rho <- 0 # correlation between the two measurements; only used together with sigma
alfa <- 0.05
lateralidad <- "bilateral" # "bilateral" or "unilateral"
poder <- 0.8 # target power (1 - beta)
perdidas <- 0.1 # expected losses, 0 to 0.5
n_dado <- 0 # inverse mode: pairs already available; 0 = not given
# The paired t test only needs the SD of the differences. When it is not reported,
# it follows from the SD of each measurement and their correlation: var(A - B) = 2*sigma^2*(1 - rho).
de_dif <- if (de_dif > 0) de_dif else sigma * sqrt(2 * (1 - rho))
tside <- if (lateralidad == "bilateral") 2 else 1
z_alfa <- qnorm(alfa/tside, lower.tail = FALSE)
z_beta <- qnorm(poder)
d_cohen <- abs(delta)/de_dif
# Exact power of the paired t test: non-central t with nu = n - 1 and ncp = sqrt(n)*|delta|/de_dif.
# It is the body power.t.test evaluates with type = "paired" (tsample = 1).
poder_de <- function(n) {
nu <- pmax(1e-07, n - 1)
pt(qt(alfa/tside, nu, lower.tail = FALSE), nu, ncp = sqrt(n) * abs(delta)/de_dif, lower.tail = FALSE)
}
# Smallest number of PAIRS reaching the target power. Two guards before calling R:
# fewer than 2 pairs is not a sample size, and an effect so small that even 1e7 pairs
# falls short is reported as "not defined" instead of a made-up number.
n <- if (poder_de(2) >= poder) 2 else
if (poder_de(1e7) < poder) NA_real_ else
power.t.test(delta = abs(delta), sd = de_dif, sig.level = alfa, power = poder,
type = "paired",
alternative = if (tside == 2) "two.sided" else "one.sided",
tol = 1e-10)$n
# Classical normal approximation with Guenther's (1981) one-sample correction, as a teaching row
n_normal <- (z_alfa + z_beta)^2 * de_dif^2 / delta^2 + z_alfa^2/2
n_ajustado <- n / (1 - perdidas) # Lwanga & Lemeshow 1991; the interface takes the ceiling
# Inverse mode: power actually reached with n_dado pairs
poder_dado <- if (n_dado >= 2) poder_de(n_dado) else NA_real_
poder_dado_normal <- if (n_dado >= 2)
pnorm(sqrt(n_dado) * abs(delta)/de_dif - z_alfa) else NA_real_
res <- list(de_dif = de_dif, z_alfa = z_alfa, z_beta = z_beta, d_cohen = d_cohen,
n = n, n_normal = n_normal, n_ajustado = n_ajustado,
poder_dado = poder_dado, poder_dado_normal = poder_dado_normal)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run here): pwr::pwr.t.test(d = d_cohen, power = poder, type = "paired")
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The sample size was computed to detect a mean change of 0.5 (in the units of the scale) in a paired design, where a standard deviation of the differences of 1.5 was assumed (Cohen's d = 0.333), with α = 0.050 (Two-sided) and 80% power, using the exact solution based on the non-central t distribution [1], equivalent to R's power.t.test(type = "paired"); the normal approximation with Guenther's correction [2] is reported for comparison and the effect-size convention follows Cohen [3]. This requires 73 pairs. A further 10% was added for expected losses, bringing the total number of pairs to recruit up to 81. The procedure follows standard design recommendations [4, 5]. Calculations used the "Sample size for paired means" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-medias-pareadas), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Original source
- 02 Guenther WC. Sample size formulas for normal theory T tests. The American Statistician. 1981;35(4):243–244. doi:10.1080/00031305.1981.10479363 Original source
- 03 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Complementary
- 04 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading
- 05 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading