← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

Proportion you expect to find, between 0 and 1 (or as a percentage). With no previous estimate, use 0.5: it is the worst case and demands the largest sample.

Half-width of the confidence interval you want. With d = 0.05 the proportion will be estimated within ±5 percentage points.

Number of people in the population you will sample from, if it is finite and known. Leave it empty (or 0) otherwise.

Proportion of participants you expect to lose (dropout, loss to follow-up, unusable samples). Leave it empty if you expect none.

Inverse mode: if you already have a given number of participants, enter it here to see the precision you would reach. Leave it empty (or 0) if you are not using it.

Example loaded

Illustrative example: estimating the prevalence of dengue among febrile patients, expected to be around 30%, within ±5 percentage points, in a service that sees 2,000 febrile patients a year and anticipating 10% losses (fictitious data).

Illustrative data, not real.

Results

Normal quantile (z)

1.960

95%

Uncorrected size (n₀)

323

exact value: 322.68

Required size (n)

278

exact value: 277.97

Recruitment with losses

309

exact value: 308.86

Precision reachable with the available n

not defined

enter an available sample size to compute it

Interpretation

With 278 participants, an observed proportion near 30.0% would be estimated with a 95% confidence interval of about ±5.0%.

The source population has 2,000 people: the finite population correction lowers the requirement from 323 to 278 participants.

With 10% expected losses, you need to recruit 309 participants so that 278 finish.

If you already have a given number of participants, enter it under "Sample size already available (n, optional)" to see how precisely you could estimate the proportion.

  • The figures depend on the assumptions (expected proportion, precision) taken from previous studies or pilots; document where they come from.
  • The declared population (N = 2000) is less than ten times the uncorrected size: the finite population correction changes the result substantially. Check that N really is the frame you will sample from.
Reachable precision by sample sizeReachable precision con 278 Participants (n): ±5.0%; target precision ±5.0%.Half-width of the CIReachable precision0.0%5.0%10.0%100200300400500Participants (n)target precision ±5.0%computed n: 278
Reachable precision by sample size

Explanation

When the aim of the study is descriptive (what proportion of febrile patients has dengue, what percentage of a cohort develops a complication), sample size is not computed to detect a difference but to make the confidence interval narrow enough. The absolute precision d is the half-width of that interval: with d = 0.05 the proportion will be estimated within ±5 percentage points.

The formula needs an expected proportion p, taken from previous studies or a pilot. Because the variance of a proportion is p(1−p), the required size is largest when p = 0.5: with no prior information, using 0.5 is the conservative choice, since no other proportion will demand a larger sample. The further p is from 0.5, the fewer participants are needed for the same precision.

If the population being sampled is finite and known (patients seen in a year, records in a department), the finite population correction reduces the size: sampling 300 out of 2,000 people carries more information than sampling 300 out of a million. The correction matters when the population is less than ten times the uncorrected size; with large populations it is irrelevant and can be left out.

Two adjustments close the calculation. Expected losses L raise recruitment to n/(1−L), so that the planned sample is the one that finishes. And the inverse mode answers the opposite question, the most common one in practice: if I already have a given number of participants, how precisely can I estimate the proportion? Every sample size is rounded up once, at the end.

Equations

n0=z1−α/22  p (1−p)d2n_0=\dfrac{z_{1-\alpha/2}^{2}\;p\,(1-p)}{d^{2}}
pp
expected proportion in the population (0.5 if unknown)
dd
absolute precision: half-width of the confidence interval
z1−α/2z_{1-\alpha/2}
normal quantile for the confidence level (1.96 at 95%)
Cochran's (1977) normal formula for estimating a proportion with absolute precision.
n=n01+n0−1Nn=\dfrac{n_0}{1+\dfrac{n_0-1}{N}}
NN
size of the finite population being sampled
Finite population correction; applied to the unrounded n₀. With no N declared, n = n₀.
naj=⌈n1−L⌉n_{aj}=\left\lceil\dfrac{n}{1-L}\right\rceil
LL
expected proportion of losses (0 to 0.5)
Adjustment for expected losses (Lwanga and Lemeshow 1991): recruitment rises so that n participants finish. The ceiling is applied once, at the end; the unrounded value is shown in the cell's detail.
d(n)=z1−α/2p (1−p)n0(n),n0(n)=n (N−1)N−nd(n)=z_{1-\alpha/2}\sqrt{\dfrac{p\,(1-p)}{n_0(n)}},\qquad n_0(n)=\dfrac{n\,(N-1)}{N-n}
n0(n)n_0(n)
uncorrected size equivalent to n in a population of N
Inverse mode: the same n(d) relationship solved for d. With no N declared, n₀(n) = n.

R code

# Sample size to estimate one proportion (absolute precision) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

p <- 0.3; d <- 0.05; nivel <- 0.95
poblacion <- 2000   # size of the finite population; 0 = population not bounded
perdidas <- 0.1     # expected losses to follow-up, 0 to 0.5
n_dado <- 0         # inverse mode: sample size already available; 0 = not used

z <- qnorm(1 - (1 - nivel) / 2)
n0 <- z^2 * p * (1 - p) / d^2                                   # Cochran 1977
n <- if (poblacion >= 2) n0 / (1 + (n0 - 1) / poblacion) else n0     # finite population correction
n_ajustado <- n / (1 - perdidas)                                # Lwanga & Lemeshow 1991
# Inverse mode: the very same n(d) solved for d, with the same correction
n0_dado <- if (poblacion >= 2) n_dado * (poblacion - 1) / (poblacion - n_dado) else n_dado
d_dado <- if (n_dado >= 2) z * sqrt(p * (1 - p) / n0_dado) else NA_real_

res <- list(z = z, n0 = n0, n = n, n_ajustado = n_ajustado, d_dado = d_dado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# The ceiling is applied once, by the interface: ceiling(n0), ceiling(n), ceiling(n_ajustado).
# Equivalent in RStudio (not run in the browser):
# presize::prec_prop(p, conf.width = 2 * d, conf.level = nivel, method = "wald")
# epiR::epi.sssimpleestb(N = poblacion, Py = p, epsilon = d, error = "absolute", conf.level = nivel)

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

Sample size was computed to estimate a proportion with absolute precision, assuming an expected proportion of 30.0% and an interval half-width of 5.0%, with a confidence level of 95%, using Cochran's normal formula [1], corrected for a finite population of 2,000 people; 10% was added for expected losses [2], giving a recruitment target of 309 participants. Calculations used the "Sample size for one proportion" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-una-proporcion), verified against R.

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Cochran WG. Sampling Techniques. 3rd ed. New York: John Wiley & Sons; 1977. Original source
  2. 02 Lwanga SK, Lemeshow S. Sample Size Determination in Health Studies: A Practical Manual. Geneva: World Health Organization; 1991. Original source
  3. 03 Newcombe RG. Two-sided confidence intervals for the single proportion: comparison of seven methods. Statistics in Medicine. 1998;17(8):857–872. doi:10.1002/(SICI)1097-0258(19980430)17:8<857::AID-SIM777>3.0.CO;2-E PMID: 9595616 Complementary
  4. 04 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
  5. 05 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading