← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

The smallest difference between the two group means that would change a clinical decision, in the units of the outcome. The sign does not change the result.

Standard deviation of the outcome, assumed equal in both groups. Take it from a previous study or a pilot and state where it came from.

Probability of concluding there is a difference when there is none. Conventionally 0.05; values from 0.001 to 0.20 are allowed.

Two-sided detects differences in either direction and is the usual choice. One-sided only makes sense if a single direction is of interest and that was decided before collecting the data.

Probability of detecting the difference if it really exists. Conventionally 0.80 or 0.90; values from 0.50 to 0.99 are allowed.

How many group 2 participants per group 1 participant. With 1 the groups are equal, which is the most efficient; with 2 there are twice as many in group 2.

Proportion of participants expected to be lost during follow-up. The total is divided by (1 − losses). Leave it empty if you do not want the adjustment.

Optional. If you already know how many participants you can recruit into group 1, type it here and the calculator will tell you what power you would reach. Leave it empty (or 0) for the usual calculation.

Example loaded

Illustrative example: detecting a 1-day difference in mean hospital stay between two treatments, with a standard deviation of 3 days, α = 0.05 two-sided, 80% power and 10% expected losses (fictitious data).

Illustrative data, not real.

Results

z of the significance level

1.960

z of the power

0.842

Cohen's d (Δ/σ)

0.333

n in group 1

143

non-central t · unrounded: 142.25

n in group 2

143

non-central t · unrounded: 142.25

Total

286

unrounded: 284.49

Per group according to power.t.test

143

unrounded: 142.25

n in group 1 (normal approximation)

143

normal with Guenther's correction · unrounded: 142.24

n in group 2 (normal approximation)

143

normal with Guenther's correction · unrounded: 142.24

Total adjusted for losses

318

10% · unrounded: 316.10

Power with the available n

not applicable

Power with the available n (normal approximation)

not applicable

Interpretation

To detect a mean difference of 1 (in the units of the outcome; Cohen's d = 0.333) with 80% power and α = 0.050 (Two-sided), 143 participants are required in group 1 and 143 in group 2, that is 286 in total.

The normal formula with Guenther's correction gives 143 per group; the exact non-central t solution is the one reported, and the one the calculation rests on.

With equal groups the result matches R's power.t.test(delta = 1, sd = 3, sig.level = 0.050, power = 0.80, type = "two.sample"), which returns 143 per group.

Allowing for 10% losses, the recruitment target rises to 318 participants in total, 159 and 159 per group, to finish the study with the 286 the calculation requires.

If you already know how many participants you can recruit, type it into "inverse mode" and you will see what power that figure would reach.

  • The figures depend entirely on the assumptions you entered (the difference that matters and the standard deviation): take them from previous studies or a pilot and state their origin in the protocol.
Power against the size of group 1Participants in group 1: 143; Power: 80%.PowerExact power (non-central t)0%25%50%75%100%50100150200250Participants in group 1Target power: 80%chosen n: 143
Power against the size of group 1

Explanation

Planning the sample size of a study that compares two means answers one concrete question: if a difference of at least Δ units existed in the population, how many participants would I need for my test to detect it with a reasonable probability? That probability is the power. The usual choices are 80% or 90%, with α set at 0.05.

The calculation needs three assumptions that do not come from this calculator but from the literature or from a pilot study: the smallest difference that matters clinically, the standard deviation of the outcome and the form of the test (two-sided or one-sided). The smallest meaningful difference is not the one you hope to find but the smallest one that would change a clinical decision; asking the study to detect anything smaller makes recruitment more expensive for no gain.

The textbook formula uses the normal distribution and underestimates slightly, because it forgets that the standard deviation is also estimated from the data. The exact solution acknowledges that detail: under the alternative hypothesis the t statistic follows a NON-central t distribution, and the sample size is the smallest n that leaves above the critical value the fraction of mass the power demands. This is exactly what power.t.test does in R, and it is what this calculator reports as the main result; the normal approximation with Guenther's (1981) correction appears next to it as a teaching row, and usually falls one or two units short.

The allocation ratio r = n₂/n₁ is for when the groups will not be the same size: r = 2 means two controls per case. Allocating unequally costs participants overall (the smallest total always comes from equal groups), but sometimes it is the only option. And since any real study will have dropouts, the adjustment for losses divides the total by (1 − L): asking for 10% losses is not pessimism, it is planning.

Equations

1−β  =  P ⁣(Tν,λ′>tν,  1−α/k),ν=n1+n2−2,λ=∣δ∣σ1n1+1n21-\beta \;=\; P\!\left(T'_{\nu,\lambda} > t_{\nu,\;1-\alpha/k}\right), \qquad \nu = n_1 + n_2 - 2, \qquad \lambda = \frac{|\delta|}{\sigma\sqrt{\dfrac{1}{n_1}+\dfrac{1}{n_2}}}
δ\delta
difference in means to be detected
σ\sigma
common standard deviation of the outcome
ν\nu
degrees of freedom of the t test
λ\lambda
non-centrality parameter
Tν,λ′T'_{\nu,\lambda}
non-central t distribution
1−β1-\beta
power required of the study
kk
2 if the test is two-sided, 1 if it is one-sided
Exact solution: n₁ is the smallest value satisfying the equality, found by root finding on the non-central t. With equal groups (r = 1) it matches R's power.t.test(type = "two.sample"); with r ≠ 1 the same equation is solved with n₂ = r·n₁.
n1  =  (1+1r)(z1−α/k+z1−β)2σ2δ2  +  z1−α/k24,n2=r n1n_1 \;=\; \frac{\left(1+\dfrac{1}{r}\right)\left(z_{1-\alpha/k}+z_{1-\beta}\right)^2\sigma^2}{\delta^2} \;+\; \frac{z_{1-\alpha/k}^2}{4}, \qquad n_2 = r\,n_1
rr
allocation ratio n₂/n₁
z1−α/kz_{1-\alpha/k}
normal quantile of the significance level
z1−βz_{1-\beta}
normal quantile of the power
kk
2 if the test is two-sided, 1 if it is one-sided
Classical normal approximation. The final term is Guenther's (1981) correction, which compensates for having treated the standard deviation as known; without it the formula falls even shorter.
1−β  ≈  Φ ⁣(∣δ∣σ1n1+1n2−z1−α/k),naj=⌈n11−L⌉+⌈n21−L⌉1-\beta \;\approx\; \Phi\!\left(\frac{|\delta|}{\sigma\sqrt{\dfrac{1}{n_1}+\dfrac{1}{n_2}}} - z_{1-\alpha/k}\right), \qquad n_{aj} = \left\lceil\dfrac{n_1}{1-L}\right\rceil + \left\lceil\dfrac{n_2}{1-L}\right\rceil
Φ\Phi
standard normal distribution function
LL
proportion of expected losses
najn_{aj}
sample size adjusted for losses
kk
2 if the test is two-sided, 1 if it is one-sided
⌈  ⌉\lceil\;\rceil
ceiling: rounding up is applied only once, at the end
Inverse mode: the power an already available sample size would reach. The calculator also reports the exact power from the non-central t, which is the one to quote. The adjustment for losses follows Lwanga and Lemeshow (1991) and is rounded PER GROUP, just like the sample size: the headline adds the two ceilings and the unrounded value stays in the cell detail.

R code

# Sample size for two independent means, exact non-central t - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

delta       <- 1         # smallest difference between means that matters
sigma       <- 3         # common standard deviation of the outcome
alfa        <- 0.05
lateralidad <- "bilateral"   # "bilateral" or "unilateral"
poder       <- 0.8         # target power (1 - beta)
r           <- 1             # allocation ratio n2/n1
perdidas    <- 0.1      # expected losses, 0 to 0.5
n_dado      <- 0        # inverse mode: n1 already available; 0 = not given

tside   <- if (lateralidad == "bilateral") 2 else 1
z_alfa  <- qnorm(alfa/tside, lower.tail = FALSE)
z_beta  <- qnorm(poder)
d_cohen <- abs(delta)/sigma

# Exact power of Student's two-sample t test: under the alternative the statistic
# follows a NON-CENTRAL t with nu = n1 + n2 - 2 and ncp = |delta| / (sigma * sqrt(1/n1 + 1/n2)).
# This is the same function power.t.test evaluates, written for any allocation ratio r.
poder_de <- function(n1) {
  nu  <- pmax(1e-07, n1 * (1 + r) - 2)
  ncp <- abs(delta) / (sigma * sqrt(1/n1 + 1/(r * n1)))
  pt(qt(alfa/tside, nu, lower.tail = FALSE), nu, ncp = ncp, lower.tail = FALSE)
}

# Smallest n1 reaching the target power, with the bracket and the tolerance of power.t.test.
# Two guards: fewer than 2 per group is not a sample size, and an effect so small that
# even 1e7 per group falls short is reported as "not defined" instead of a made-up number.
n1 <- if (poder_de(2) >= poder) 2 else
      if (poder_de(1e7) < poder) NA_real_ else
      uniroot(function(n1) poder_de(n1) - poder, c(2, 1e7), tol = 1e-10)$root
n2      <- r * n1
n_total <- n1 + n2

# power.t.test only covers equal groups; with r != 1 it does not apply and the equation above is the answer.
# It needs the SAME two guards as n1: power.t.test extends the bracket (extendInt = "upX") and
# would answer with a sample size below 2, or beyond the 1e7 where the equation above stops.
n_ptt <- if (r != 1) NA_real_ else
         if (poder_de(2) >= poder) 2 else
         if (poder_de(1e7) < poder) NA_real_ else
         power.t.test(delta = abs(delta), sd = sigma, sig.level = alfa, power = poder,
                      type = "two.sample",
                      alternative = if (tside == 2) "two.sided" else "one.sided",
                      tol = 1e-10)$n

# Classical normal approximation with Guenther's (1981) correction, shown as a teaching row
n1_normal <- (1 + 1/r) * (z_alfa + z_beta)^2 * sigma^2 / delta^2 + z_alfa^2/4
n2_normal <- r * n1_normal

n_ajustado <- n_total / (1 - perdidas)   # Lwanga & Lemeshow 1991; the interface takes the ceiling

# Inverse mode: power actually reached with n1 = n_dado and n2 = r * n_dado
poder_dado <- if (n_dado >= 2) poder_de(n_dado) else NA_real_
poder_dado_normal <- if (n_dado >= 2)
  pnorm(abs(delta) / (sigma * sqrt(1/n_dado + 1/(r * n_dado))) - z_alfa) else NA_real_

res <- list(z_alfa = z_alfa, z_beta = z_beta, d_cohen = d_cohen,
            n1 = n1, n2 = n2, n_total = n_total, n_ptt = n_ptt,
            n1_normal = n1_normal, n2_normal = n2_normal, n_ajustado = n_ajustado,
            poder_dado = poder_dado, poder_dado_normal = poder_dado_normal)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio (not run here): pwr::pwr.t.test(d = d_cohen, power = poder, type = "two.sample")
# and, for unequal groups, pwr::pwr.t2n.test(n1 = ..., n2 = ..., d = d_cohen)

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

The sample size was computed to compare the means of two independent groups, assuming a minimum difference of interest of 1 in the units of the outcome and a common standard deviation of 3 (Cohen's d = 0.333), with α = 0.050 (Two-sided) and 80% power [3], using the exact solution based on the non-central t distribution [1], equivalent to R's power.t.test; the normal approximation with Guenther's correction [2] is reported for comparison and the effect-size convention follows Cohen [4, 5]. This requires 143 and 143 participants per group (286 in total). A further 10% was added for expected losses, bringing the total to recruit up to 318. The procedure follows standard design recommendations [6, 7, 8]. Calculations used the "Sample size for two means" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-dos-medias), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Student. The probable error of a mean. Biometrika. 1908;6(1):1–25. doi:10.1093/biomet/6.1.1 Original source
  2. 02 Guenther WC. Sample size formulas for normal theory T tests. The American Statistician. 1981;35(4):243–244. doi:10.1080/00031305.1981.10479363 Original source
  3. 03 Neyman J, Pearson ES. On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character. 1933;231:289–337. doi:10.1098/rsta.1933.0009 Complementary
  4. 04 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Complementary
  5. 05 Lachin JM. Introduction to sample size determination and power analysis for clinical trials. Controlled Clinical Trials. 1981;2(2):93–113. doi:10.1016/0197-2456(81)90001-5 PMID: 7273794 Complementary
  6. 06 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading
  7. 07 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
  8. 08 Chow SC, Shao J, Wang H, Lokhnygina Y. Sample Size Calculations in Clinical Research. 3rd ed. Boca Raton: Chapman & Hall/CRC; 2017. doi:10.1201/9781315183084 Complementary