← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

The smallest mean change between the first and the second measurement that would alter a clinical decision. The sign does not change the result.

Standard deviation of the within-pair differences. This is the figure that rules: if you have it, leave σ and ρ empty.

Standard deviation of each measurement on its own. It is only used if you did not enter the SD of the differences, and then the correlation is needed too.

Expected correlation between the first and the second measurement on the same participant. Only used together with σ. Values from 0 to 0.98 are allowed.

Probability of concluding there is a change when there is none. Conventionally 0.05; values from 0.001 to 0.20 are allowed.

Two-sided detects changes in either direction and is the usual choice. One-sided only makes sense if a single direction is of interest and that was decided before seeing the data.

Probability of detecting the change if it really exists. Conventionally 0.80 or 0.90; values from 0.50 to 0.99 are allowed.

Proportion of pairs expected to be lost (one of the two measurements missing). The total is divided by (1 − losses).

Optional. If you already know how many complete pairs you will have, type it in and the calculator will tell you what power you would reach. Leave it empty (or 0) for the usual calculation.

Example loaded

Illustrative example: detecting a mean drop of 0.5 points on a clinical scale measured before and after, with a standard deviation of the differences of 1.5 points, α = 0.05 two-sided, 80% power and 10% expected losses (fictitious data).

Illustrative data, not real.

Results

SD of the differences (σ_d)

1.5

entered

z of the significance level

1.960

z of the power

0.842

Cohen's d (Δ/σ_d)

0.333

Pairs needed

73

non-central t · unrounded: 72.58

Pairs (normal approximation)

73

normal with Guenther's correction · unrounded: 72.56

Pairs adjusted for losses

81

10% · unrounded: 80.65

Power with the available pairs

not applicable

Power with the available pairs (normal approximation)

not applicable

Interpretation

To detect a mean change of 0.5 (in the units of the scale; Cohen's d = 0.333) with 80% power and α = 0.050 (Two-sided), 73 pairs are required.

The standard deviation of the differences was taken directly from what you entered: 1.5.

The normal formula with Guenther's correction gives 73 pairs; the exact non-central t solution, identical to R's power.t.test(type = "paired"), is the one reported and the one the calculation rests on.

Allowing for 10% losses, you should recruit 81 pairs to finish the study with the 73 the calculation requires.

If you already know how many complete pairs you will have, type it into "inverse mode" and you will see what power that figure would reach.

  • The figures depend entirely on the assumptions you entered (the change that matters and the variability of the differences): take them from previous studies or a pilot and state their origin in the protocol.
Power against the number of pairsPairs: 73; Power: 80%.PowerExact power (non-central t)0%25%50%75%100%50100PairsTarget power: 80%chosen n: 73
Power against the number of pairs

Explanation

In a paired design each participant is their own control: measured before and after, or with the two techniques being compared. What gets analysed is not two groups but a single column of differences, and the one-sample t test asks whether their mean is zero. That is why the sample size is counted in PAIRS, not in measurements.

The quantity that rules is not the variability of the measurements but that of the differences. If the two measurements on the same person are very alike, the differences vary little and very few pairs are needed; if they are nearly independent, the paired design loses its advantage. The link between them is σ_d = σ·√(2(1 − ρ)): at ρ = 0.5 the SD of the differences equals the SD of each measurement, and above 0.5 the pairing starts to pay off.

This calculator accepts both routes. If the paper you start from reports the standard deviation of the differences, type it in: it is the exact figure and it overrides anything else. If it only reports the SD of each measurement, enter it together with the assumed correlation and the calculator derives σ_d. That ρ is an assumption, not a datum, so justify it and stress-test the result with neighbouring values.

The main result is the exact solution based on the non-central t distribution, the same one power.t.test(type = "paired") returns in R. The normal approximation with Guenther's (1981) correction is shown next to it as a teaching row; its correction is half the two-sample one because here there is only one variance to estimate. And since any study will have incomplete pairs, the adjustment for losses divides the total by (1 − L).

Equations

1−β  =  P ⁣(Tn−1,λ′>tn−1,  1−α/k),λ=n ∣δ∣σd1-\beta \;=\; P\!\left(T'_{n-1,\lambda} > t_{n-1,\;1-\alpha/k}\right), \qquad \lambda = \frac{\sqrt{n}\,|\delta|}{\sigma_d}
δ\delta
mean of the differences to be detected
σd\sigma_d
standard deviation of the differences
nn
number of pairs
λ\lambda
non-centrality parameter
Tn−1,λ′T'_{n-1,\lambda}
non-central t distribution with n − 1 degrees of freedom
kk
2 if the test is two-sided, 1 if it is one-sided
Exact solution: n is the smallest number of pairs satisfying the equality. It matches R's power.t.test(type = "paired", tol = 1e-10), which solves this same equation.
σd  =  σ2 (1−ρ)\sigma_d \;=\; \sigma\sqrt{2\,(1-\rho)}
σ\sigma
standard deviation of each measurement
ρ\rho
correlation between the two measurements
Alternative route when the source study does not report the SD of the differences. At ρ = 0.5 it gives σ_d = σ; as ρ → 1 the differences vanish, which is why ρ ≥ 0.99 is blocked.
n  =  (z1−α/k+z1−β)2σd2δ2  +  z1−α/k22,1−β  ≈  Φ ⁣(n ∣δ∣σd−z1−α/k)naj  =  ⌈n1−L⌉\begin{gathered} n \;=\; \frac{\left(z_{1-\alpha/k}+z_{1-\beta}\right)^2\sigma_d^2}{\delta^2} \;+\; \frac{z_{1-\alpha/k}^2}{2}, \qquad 1-\beta \;\approx\; \Phi\!\left(\frac{\sqrt{n}\,|\delta|}{\sigma_d} - z_{1-\alpha/k}\right) \\ n_{aj} \;=\; \left\lceil\frac{n}{1-L}\right\rceil \end{gathered}
z1−α/kz_{1-\alpha/k}
normal quantile of the significance level
z1−βz_{1-\beta}
normal quantile of the power
Φ\Phi
standard normal distribution function
kk
2 if the test is two-sided, 1 if it is one-sided
LL
proportion of expected losses
najn_{aj}
pairs adjusted for losses
⌈  ⌉\lceil\;\rceil
ceiling: rounding up is applied only once, at the end
Guenther's (1981) one-sample normal approximation and its inverse mode. The adjustment for losses follows Lwanga and Lemeshow (1991); the ceiling is applied only once, at the end, and the unrounded value stays in the cell detail. Pairs are a single group, so there are no ceilings to add up here.

R code

# Sample size for paired means, exact non-central t - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

delta       <- 0.5         # mean of the within-pair differences that matters
de_dif      <- 1.5        # SD of the DIFFERENCES; 0 = not captured, then it is derived below
sigma       <- 0         # SD of each measurement; only used when de_dif is not given (0 = not captured)
rho         <- 0           # correlation between the two measurements; only used together with sigma
alfa        <- 0.05
lateralidad <- "bilateral"   # "bilateral" or "unilateral"
poder       <- 0.8         # target power (1 - beta)
perdidas    <- 0.1      # expected losses, 0 to 0.5
n_dado      <- 0        # inverse mode: pairs already available; 0 = not given

# The paired t test only needs the SD of the differences. When it is not reported,
# it follows from the SD of each measurement and their correlation: var(A - B) = 2*sigma^2*(1 - rho).
de_dif <- if (de_dif > 0) de_dif else sigma * sqrt(2 * (1 - rho))

tside   <- if (lateralidad == "bilateral") 2 else 1
z_alfa  <- qnorm(alfa/tside, lower.tail = FALSE)
z_beta  <- qnorm(poder)
d_cohen <- abs(delta)/de_dif

# Exact power of the paired t test: non-central t with nu = n - 1 and ncp = sqrt(n)*|delta|/de_dif.
# It is the body power.t.test evaluates with type = "paired" (tsample = 1).
poder_de <- function(n) {
  nu <- pmax(1e-07, n - 1)
  pt(qt(alfa/tside, nu, lower.tail = FALSE), nu, ncp = sqrt(n) * abs(delta)/de_dif, lower.tail = FALSE)
}

# Smallest number of PAIRS reaching the target power. Two guards before calling R:
# fewer than 2 pairs is not a sample size, and an effect so small that even 1e7 pairs
# falls short is reported as "not defined" instead of a made-up number.
n <- if (poder_de(2) >= poder) 2 else
     if (poder_de(1e7) < poder) NA_real_ else
     power.t.test(delta = abs(delta), sd = de_dif, sig.level = alfa, power = poder,
                  type = "paired",
                  alternative = if (tside == 2) "two.sided" else "one.sided",
                  tol = 1e-10)$n

# Classical normal approximation with Guenther's (1981) one-sample correction, as a teaching row
n_normal <- (z_alfa + z_beta)^2 * de_dif^2 / delta^2 + z_alfa^2/2

n_ajustado <- n / (1 - perdidas)   # Lwanga & Lemeshow 1991; the interface takes the ceiling

# Inverse mode: power actually reached with n_dado pairs
poder_dado <- if (n_dado >= 2) poder_de(n_dado) else NA_real_
poder_dado_normal <- if (n_dado >= 2)
  pnorm(sqrt(n_dado) * abs(delta)/de_dif - z_alfa) else NA_real_

res <- list(de_dif = de_dif, z_alfa = z_alfa, z_beta = z_beta, d_cohen = d_cohen,
            n = n, n_normal = n_normal, n_ajustado = n_ajustado,
            poder_dado = poder_dado, poder_dado_normal = poder_dado_normal)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio (not run here): pwr::pwr.t.test(d = d_cohen, power = poder, type = "paired")

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

The sample size was computed to detect a mean change of 0.5 (in the units of the scale) in a paired design, where a standard deviation of the differences of 1.5 was assumed (Cohen's d = 0.333), with α = 0.050 (Two-sided) and 80% power, using the exact solution based on the non-central t distribution [1], equivalent to R's power.t.test(type = "paired"); the normal approximation with Guenther's correction [2] is reported for comparison and the effect-size convention follows Cohen [3]. This requires 73 pairs. A further 10% was added for expected losses, bringing the total number of pairs to recruit up to 81. The procedure follows standard design recommendations [4, 5]. Calculations used the "Sample size for paired means" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-medias-pareadas), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Original source
  2. 02 Guenther WC. Sample size formulas for normal theory T tests. The American Statistician. 1981;35(4):243–244. doi:10.1080/00031305.1981.10479363 Original source
  3. 03 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Complementary
  4. 04 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading
  5. 05 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading