Tools · Open Biostatistics
Sample size for comparing two independent means
Enter the difference in means you want to be able to detect, the expected standard deviation, the significance level and the power, and get how many participants each group needs, with the power curve, a plain-language interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-dos-medias
This link does not include the pasted data: they are too long for a URL.
Results
z of the significance level
1.960
z of the power
0.842
Cohen's d (Δ/σ)
0.333
n in group 1
143
non-central t · unrounded: 142.25
n in group 2
143
non-central t · unrounded: 142.25
Total
286
unrounded: 284.49
Per group according to power.t.test
143
unrounded: 142.25
n in group 1 (normal approximation)
143
normal with Guenther's correction · unrounded: 142.24
n in group 2 (normal approximation)
143
normal with Guenther's correction · unrounded: 142.24
Total adjusted for losses
318
10% · unrounded: 316.10
Power with the available n
not applicable
Power with the available n (normal approximation)
not applicable
Interpretation
To detect a mean difference of 1 (in the units of the outcome; Cohen's d = 0.333) with 80% power and α = 0.050 (Two-sided), 143 participants are required in group 1 and 143 in group 2, that is 286 in total.
The normal formula with Guenther's correction gives 143 per group; the exact non-central t solution is the one reported, and the one the calculation rests on.
With equal groups the result matches R's power.t.test(delta = 1, sd = 3, sig.level = 0.050, power = 0.80, type = "two.sample"), which returns 143 per group.
Allowing for 10% losses, the recruitment target rises to 318 participants in total, 159 and 159 per group, to finish the study with the 286 the calculation requires.
If you already know how many participants you can recruit, type it into "inverse mode" and you will see what power that figure would reach.
- The figures depend entirely on the assumptions you entered (the difference that matters and the standard deviation): take them from previous studies or a pilot and state their origin in the protocol.
Explanation
Planning the sample size of a study that compares two means answers one concrete question: if a difference of at least Δ units existed in the population, how many participants would I need for my test to detect it with a reasonable probability? That probability is the power. The usual choices are 80% or 90%, with α set at 0.05.
The calculation needs three assumptions that do not come from this calculator but from the literature or from a pilot study: the smallest difference that matters clinically, the standard deviation of the outcome and the form of the test (two-sided or one-sided). The smallest meaningful difference is not the one you hope to find but the smallest one that would change a clinical decision; asking the study to detect anything smaller makes recruitment more expensive for no gain.
The textbook formula uses the normal distribution and underestimates slightly, because it forgets that the standard deviation is also estimated from the data. The exact solution acknowledges that detail: under the alternative hypothesis the t statistic follows a NON-central t distribution, and the sample size is the smallest n that leaves above the critical value the fraction of mass the power demands. This is exactly what power.t.test does in R, and it is what this calculator reports as the main result; the normal approximation with Guenther's (1981) correction appears next to it as a teaching row, and usually falls one or two units short.
The allocation ratio r = n₂/n₁ is for when the groups will not be the same size: r = 2 means two controls per case. Allocating unequally costs participants overall (the smallest total always comes from equal groups), but sometimes it is the only option. And since any real study will have dropouts, the adjustment for losses divides the total by (1 − L): asking for 10% losses is not pessimism, it is planning.
Equations
- difference in means to be detected
- common standard deviation of the outcome
- degrees of freedom of the t test
- non-centrality parameter
- non-central t distribution
- power required of the study
- 2 if the test is two-sided, 1 if it is one-sided
- allocation ratio n₂/n₁
- normal quantile of the significance level
- normal quantile of the power
- 2 if the test is two-sided, 1 if it is one-sided
- standard normal distribution function
- proportion of expected losses
- sample size adjusted for losses
- 2 if the test is two-sided, 1 if it is one-sided
- ceiling: rounding up is applied only once, at the end
R code
# Sample size for two independent means, exact non-central t - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
delta <- 1 # smallest difference between means that matters
sigma <- 3 # common standard deviation of the outcome
alfa <- 0.05
lateralidad <- "bilateral" # "bilateral" or "unilateral"
poder <- 0.8 # target power (1 - beta)
r <- 1 # allocation ratio n2/n1
perdidas <- 0.1 # expected losses, 0 to 0.5
n_dado <- 0 # inverse mode: n1 already available; 0 = not given
tside <- if (lateralidad == "bilateral") 2 else 1
z_alfa <- qnorm(alfa/tside, lower.tail = FALSE)
z_beta <- qnorm(poder)
d_cohen <- abs(delta)/sigma
# Exact power of Student's two-sample t test: under the alternative the statistic
# follows a NON-CENTRAL t with nu = n1 + n2 - 2 and ncp = |delta| / (sigma * sqrt(1/n1 + 1/n2)).
# This is the same function power.t.test evaluates, written for any allocation ratio r.
poder_de <- function(n1) {
nu <- pmax(1e-07, n1 * (1 + r) - 2)
ncp <- abs(delta) / (sigma * sqrt(1/n1 + 1/(r * n1)))
pt(qt(alfa/tside, nu, lower.tail = FALSE), nu, ncp = ncp, lower.tail = FALSE)
}
# Smallest n1 reaching the target power, with the bracket and the tolerance of power.t.test.
# Two guards: fewer than 2 per group is not a sample size, and an effect so small that
# even 1e7 per group falls short is reported as "not defined" instead of a made-up number.
n1 <- if (poder_de(2) >= poder) 2 else
if (poder_de(1e7) < poder) NA_real_ else
uniroot(function(n1) poder_de(n1) - poder, c(2, 1e7), tol = 1e-10)$root
n2 <- r * n1
n_total <- n1 + n2
# power.t.test only covers equal groups; with r != 1 it does not apply and the equation above is the answer.
# It needs the SAME two guards as n1: power.t.test extends the bracket (extendInt = "upX") and
# would answer with a sample size below 2, or beyond the 1e7 where the equation above stops.
n_ptt <- if (r != 1) NA_real_ else
if (poder_de(2) >= poder) 2 else
if (poder_de(1e7) < poder) NA_real_ else
power.t.test(delta = abs(delta), sd = sigma, sig.level = alfa, power = poder,
type = "two.sample",
alternative = if (tside == 2) "two.sided" else "one.sided",
tol = 1e-10)$n
# Classical normal approximation with Guenther's (1981) correction, shown as a teaching row
n1_normal <- (1 + 1/r) * (z_alfa + z_beta)^2 * sigma^2 / delta^2 + z_alfa^2/4
n2_normal <- r * n1_normal
n_ajustado <- n_total / (1 - perdidas) # Lwanga & Lemeshow 1991; the interface takes the ceiling
# Inverse mode: power actually reached with n1 = n_dado and n2 = r * n_dado
poder_dado <- if (n_dado >= 2) poder_de(n_dado) else NA_real_
poder_dado_normal <- if (n_dado >= 2)
pnorm(abs(delta) / (sigma * sqrt(1/n_dado + 1/(r * n_dado))) - z_alfa) else NA_real_
res <- list(z_alfa = z_alfa, z_beta = z_beta, d_cohen = d_cohen,
n1 = n1, n2 = n2, n_total = n_total, n_ptt = n_ptt,
n1_normal = n1_normal, n2_normal = n2_normal, n_ajustado = n_ajustado,
poder_dado = poder_dado, poder_dado_normal = poder_dado_normal)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run here): pwr::pwr.t.test(d = d_cohen, power = poder, type = "two.sample")
# and, for unequal groups, pwr::pwr.t2n.test(n1 = ..., n2 = ..., d = d_cohen)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The sample size was computed to compare the means of two independent groups, assuming a minimum difference of interest of 1 in the units of the outcome and a common standard deviation of 3 (Cohen's d = 0.333), with α = 0.050 (Two-sided) and 80% power [3], using the exact solution based on the non-central t distribution [1], equivalent to R's power.t.test; the normal approximation with Guenther's correction [2] is reported for comparison and the effect-size convention follows Cohen [4, 5]. This requires 143 and 143 participants per group (286 in total). A further 10% was added for expected losses, bringing the total to recruit up to 318. The procedure follows standard design recommendations [6, 7, 8]. Calculations used the "Sample size for two means" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-dos-medias), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Original source
- 02 Guenther WC. Sample size formulas for normal theory T tests. The American Statistician. 1981;35(4):243–244. doi:10.1080/00031305.1981.10479363 Original source
- 03 Neyman J, Pearson ES. On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character. 1933;231:289–337. doi:10.1098/rsta.1933.0009 Complementary
- 04 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Complementary
- 05 Lachin JM. Introduction to sample size determination and power analysis for clinical trials. Controlled Clinical Trials. 1981;2(2):93–113. doi:10.1016/0197-2456(81)90001-5 PMID: 7273794 Complementary
- 06 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading
- 07 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
- 08 Chow SC, Shao J, Wang H, Lokhnygina Y. Sample Size Calculations in Clinical Research. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2017. doi:10.1201/9781315183084 Complementary