← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

Proportion of the outcome you expect in the first group (control or usual care). Take it from previous studies or from a pilot.

Proportion you expect in the second group. The difference between p₁ and p₂ is what weighs most on the sample size.

Probability of declaring a difference that does not exist. Write it as a decimal (0.05, not 5).

A two-sided test detects differences in either direction and is the usual choice. A one-sided test needs fewer participants but must be justified before seeing the data.

Probability of detecting the difference if it really exists. The conventional value is 0.80; 0.90 or 0.95 for confirmatory studies.

Participants in group 2 for each one in group 1. With 1 the groups are equal, which is the most efficient split; with 2 there are twice as many in group 2.

Proportion of participants you expect to lose during follow-up. The total to recruit is divided by one minus this figure.

Inverse mode: if you already know how many people you can study in group 1, write it here to see the power you reach. Leave it at 0 to skip it.

Example loaded

Illustrative example: a clinical trial comparing two treatments in patients with dengue with warning signs; cure is expected in 70 % with usual care and in 85 % with the new treatment, with α = 0.05 two-sided, 80 % power, equal groups and 10 % anticipated losses (fictitious data).

Illustrative data, not real.

Results

z for α (z₁₋α/₂)

1.960

z for the power (z₁₋β)

0.842

n₁ from Fleiss without correction

121

Fleiss formula · exact 120.47

n₂ from Fleiss without correction

121

Fleiss formula · exact 120.47

n₁ with continuity correction

134

Fleiss, Tytun and Ury · exact 133.47

n₂ with continuity correction

134

Fleiss, Tytun and Ury · exact 133.47

Participants in group 1 (n₁)

134

reported figure · exact 133.47

Participants in group 2 (n₂)

134

reported figure · exact 133.47

Total to be studied

268

sum of the two rounded groups · exact 266.94

Total to recruit given the expected losses

298

sum of the two rounded groups · exact 296.61

Comparison: R's power.prop.test

121

equal groups only · exact 120.47

Cohen's h (arcsine transformation)

-0.364

Comparison: pwr::pwr.2p.test from h

119

equal groups only · exact 118.55

Power attainable with the available participants

Not applicable

without continuity correction

Interpretation

To detect a proportion of 70.0% versus one of 85.0% with 80% power and α = 0.05 (two-sided), 134 participants are required in group 1 and 134 in group 2, 268 in total.

The figure reported carries the continuity correction of Fleiss, Tytun and Ury, which approximates what Fisher's exact test will require: 134 and 134 participants. Uncorrected, Fleiss's formula — which approximates Pearson's chi-squared — would ask for 121 and 121. Report whichever matches the test you will run.

With equal groups, R's power.prop.test solves the same uncorrected equation and returns 121 per group. Cohen's alternative starts from the arcsine transformation: h = -0.364 and 119 per group. They are two approximations to the same problem, so a difference of a few participants is expected.

Anticipating 10% losses during follow-up, the recruitment target rises to 298 participants in total, 149 and 149 per group.

If you already know how many people you can study, write that number in the field for available participants and the calculator will report the power you reach with it.

  • These figures depend entirely on the proportions you assumed: take them from previous studies, from a pilot or from the smallest clinically important difference, and document their origin in Methods. A small change in the expected difference changes the sample size a great deal.
Power of the test against the number of participants in group 1, with the required size and the target power markedPower of the test: Participants in group 1 (n₁) = 134 Participants in group 2 (n₂) = 134; Power (1 − β) = 80%.PowerPower of the test0%25%50%75%100%50100150200250Participants in group 1 (n₁)target power: 80%n₁ required: 134
Power of the test against the number of participants in group 1, with the required size and the target power marked

Explanation

Before a trial or any comparative study you have to decide how many people to study. The question this calculator answers is: if the outcome occurs in a proportion p₁ of one group and a proportion p₂ of the other, how many participants do I need in each group for the statistical test to detect that difference with the probability I require? That probability is the power, and the probability of declaring a difference that does not exist is α.

The classical formula is Fleiss's: it adds the error tolerated under the null hypothesis (the pooled variance, with both proportions mixed together) to the error tolerated under the alternative (the two variances separately), and divides everything by the square of the difference to be detected. That difference is what dominates: going from 0.70 to 0.85 takes a few hundred participants; going from 0.70 to 0.72 takes tens of thousands.

The continuity correction of Fleiss, Tytun and Ury adds a margin because the real variable is discrete (people are counted) while the formula uses a continuous curve. The corrected size approximates what Fisher's exact test will require; the uncorrected size, what Pearson's chi-squared will require. Report the one that matches the test you will actually run and say so in Methods; both are shown here.

The allocation ratio r = n₂/n₁ allows unequal groups (two controls per case, for instance). Splitting in half always gives the smallest total: any other split demands more participants overall for the same power. Expected losses are covered by dividing by 1 − L at the end, and rounding up is applied once, per group, because you cannot recruit half a person.

All of these figures depend entirely on the proportions you assume. They are not data: they are an informed bet that must come from previous studies, from a pilot or from the smallest difference you consider clinically important. Document where they came from; a sample size without that justification is not reproducible.

Equations

n1=[z1−α/k(r+1) pˉ qˉ+z1−βr p1q1+p2q2]2r (p1−p2)2,pˉ=p1+r p2r+1,qˉ=1−pˉ,n2=r n1n_1=\frac{\left[z_{1-\alpha/k}\sqrt{(r+1)\,\bar p\,\bar q}+z_{1-\beta}\sqrt{r\,p_1q_1+p_2q_2}\right]^2}{r\,(p_1-p_2)^2},\qquad \bar p=\frac{p_1+r\,p_2}{r+1},\qquad \bar q=1-\bar p,\qquad n_2=r\,n_1
p1, p2p_1,\ p_2
proportions expected in group 1 and in group 2
q1=1−p1q_1=1-p_1
complement of p₁ (and q₂ = 1 − p₂)
r=n2/n1r=n_2/n_1
allocation ratio between the two groups
pˉ\bar p
pooled proportion under the null hypothesis
z1−α/kz_{1-\alpha/k}
normal quantile of the significance level
kk
2 if the test is two-sided, 1 if it is one-sided
z1−βz_{1-\beta}
normal quantile of the power
Fleiss's formula without continuity correction, valid for any allocation ratio. With r = 1 it is exactly what R's power.prop.test solves, because (p₁ + p₂)(1 − (p₁ + p₂)/2) = 2p̄q̄.
n1′=n14[1+1+2(r+1)r n1 ∣p1−p2∣]2,n2′=r n1′n_1'=\frac{n_1}{4}\left[1+\sqrt{1+\frac{2(r+1)}{r\,n_1\,\lvert p_1-p_2\rvert}}\right]^2,\qquad n_2'=r\,n_1'
n1n_1
uncorrected size given by the previous equation
Continuity correction of Fleiss, Tytun and Ury (1980). The corrected size approximates Fisher's exact test; the uncorrected size, Pearson's chi-squared.
1−β=Φ ⁣(∣p1−p2∣r n1−z1−α/k(r+1) pˉ qˉr p1q1+p2q2)1-\beta=\Phi\!\left(\frac{\lvert p_1-p_2\rvert\sqrt{r\,n_1}-z_{1-\alpha/k}\sqrt{(r+1)\,\bar p\,\bar q}}{\sqrt{r\,p_1q_1+p_2q_2}}\right)
Φ\Phi
cumulative distribution function of the standard normal
kk
2 if the test is two-sided, 1 if it is one-sided
Inverse mode: power attainable with a sample size that is already fixed. It is the same equation solved for the power, without continuity correction, and it is also the curve drawn in the plot.
h=2arcsin⁡p1−2arcsin⁡p2,n=2(z1−α/k+z1−β)2h2h=2\arcsin\sqrt{p_1}-2\arcsin\sqrt{p_2},\qquad n=\frac{2\left(z_{1-\alpha/k}+z_{1-\beta}\right)^2}{h^2}
hh
Cohen's effect size for two proportions
kk
2 if the test is two-sided, 1 if it is one-sided
Cohen's alternative (1988): the arcsine transformation stabilises the variance, so the sample size depends on h alone. This is the route taken by pwr::pwr.2p.test and it gives figures very close to, but not identical with, Fleiss's.
naj=⌈n11−L⌉+⌈n21−L⌉n_{\mathrm{aj}}=\left\lceil\frac{n_1}{1-L}\right\rceil+\left\lceil\frac{n_2}{1-L}\right\rceil
LL
proportion of losses expected during follow-up
Adjustment for losses from Lwanga and Lemeshow (1991). Rounding up is applied once and per group, at the end, because you cannot recruit half a person: the totals shown are the sum of the two already rounded groups, and the exact unrounded value appears in the detail of each cell.

R code

# Sample size for two independent proportions (Fleiss) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
library(pwr)

p1 <- 0.7; p2 <- 0.85       # proportions expected in group 1 and in group 2
alfa        <- 0.05        # significance level
lateralidad <- "bilateral" # "bilateral" or "unilateral"
poder       <- 0.8       # target power, 1 - beta
r           <- 1           # allocation ratio n2/n1 (1 = equal groups)
correccion  <- "si"  # "si" = report the continuity-corrected n
perdidas    <- 0.1    # expected losses to follow-up, 0 to 0.5
n_dado      <- 0      # inverse mode: n1 already available; 0 = not used

alternative <- if (lateralidad == "bilateral") "two.sided" else "one.sided"
tside  <- if (lateralidad == "bilateral") 2 else 1
z_alfa <- qnorm(alfa / tside, lower.tail = FALSE)
z_beta <- qnorm(poder)

# Fleiss (1981; Fleiss, Levin & Paik 2003, section 4.2) written out by hand so that
# any allocation ratio is allowed: pooled variance under H0, separate variances under H1.
d    <- abs(p1 - p2)
q1   <- 1 - p1
q2   <- 1 - p2
pbar <- (p1 + r * p2) / (r + 1)
qbar <- 1 - pbar
n1_fleiss <- (z_alfa * sqrt((r + 1) * pbar * qbar) + z_beta * sqrt(r * p1 * q1 + p2 * q2))^2 / (r * d^2)
n2_fleiss <- r * n1_fleiss

# Continuity correction of Fleiss, Tytun & Ury (1980): the corrected n approximates
# Fisher's exact test, the uncorrected one approximates Pearson's chi-squared.
n1_cc <- n1_fleiss / 4 * (1 + sqrt(1 + 2 * (r + 1) / (r * n1_fleiss * d)))^2
n2_cc <- r * n1_cc

n1 <- if (correccion == "si") n1_cc else n1_fleiss
n2 <- r * n1
n_total    <- n1 + n2
n_ajustado <- n_total / (1 - perdidas)   # Lwanga & Lemeshow 1991; the interface takes the ceiling

# Same problem solved by base R. It assumes equal groups, so with r != 1 there is
# nothing to compare against; tryCatch keeps an unsolvable design from stopping the script.
n_ppt <- if (r == 1) {
  tryCatch(power.prop.test(p1 = p1, p2 = p2, sig.level = alfa, power = poder,
                           alternative = alternative, tol = 1e-10)$n,
           error = function(e) NA_real_)
} else NA_real_

# Cohen's h (arcsine transformation) and the n that pwr derives from it. pwr has no
# "one.sided" level: the one-sided test is asked for with "greater" and h > 0, so the
# magnitude of h is what travels, exactly as pwr itself does with "two.sided".
h_cohen <- ES.h(p1, p2)
alt_pwr <- if (lateralidad == "bilateral") "two.sided" else "greater"
n_pwr_h <- if (r == 1) {
  tryCatch(pwr.2p.test(h = abs(h_cohen), sig.level = alfa, power = poder,
                       alternative = alt_pwr)$n,
           error = function(e) NA_real_)
} else NA_real_

# Inverse mode: power attainable with n1 = n_dado and n2 = r * n_dado, by the same
# normal formula of Fleiss without the continuity correction.
poder_dado <- if (n_dado >= 2) {
  pnorm((d * sqrt(r * n_dado) - z_alfa * sqrt((r + 1) * pbar * qbar)) / sqrt(r * p1 * q1 + p2 * q2))
} else NA_real_

res <- list(z_alfa = z_alfa, z_beta = z_beta,
            n1_fleiss = n1_fleiss, n2_fleiss = n2_fleiss,
            n1_cc = n1_cc, n2_cc = n2_cc,
            n1 = n1, n2 = n2, n_total = n_total, n_ajustado = n_ajustado,
            n_ppt = n_ppt, h_cohen = h_cohen, n_pwr_h = n_pwr_h,
            poder_dado = poder_dado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio (not run in the browser):
# epiR::epi.sscompb(treat = 0.85, control = 0.70, n = NA, power = 0.80, r = 1, conf.level = 0.95)

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

The sample size was computed to compare two independent proportions (70.0% versus 85.0%, allocation ratio 1), assuming those proportions come from previous studies or from a pilot, with α = 0.05 (two-sided) and 80% power, using the Fleiss formula for two independent proportions with the continuity correction of Fleiss, Tytun and Ury [1,2]; 134 and 134 participants per group are required, 268 in total. A further 10% was added for anticipated losses [7], so that the recruitment target is 298 participants. Calculations were carried out with the «Sample size (two proportions)» calculator from Open Biostatistics (Academic Body UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-dos-proporciones), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Fleiss JL, Tytun A, Ury HK. A simple approximation for calculating sample sizes for comparing independent proportions. Biometrics. 1980;36(2):343–346. doi:10.2307/2529990 PMID: 26625475 Original source
  2. 02 Fleiss JL, Levin B, Paik MC. Statistical Methods for Rates and Proportions. 3rd ed. Hoboken, NJ: John Wiley & Sons; 2003. doi:10.1002/0471445428 Original source
  3. 03 Casagrande JT, Pike MC, Smith PG. An improved approximate formula for calculating sample sizes for comparing two binomial distributions. Biometrics. 1978;34(3):483–486. doi:10.2307/2530613 PMID: 719125 Complementary
  4. 04 Lachin JM. Introduction to sample size determination and power analysis for clinical trials. Controlled Clinical Trials. 1981;2(2):93–113. doi:10.1016/0197-2456(81)90001-5 PMID: 7273794 Complementary
  5. 05 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Complementary
  6. 06 Champely S. pwr: Basic Functions for Power Analysis. R package version 1.3-0. CRAN; 2020. Complementary
  7. 07 Lwanga SK, Lemeshow S. Sample Size Determination in Health Studies: A Practical Manual. Geneva: World Health Organization; 1991. Didactic reading
  8. 08 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading