Tools · Open Biostatistics
Sample size to estimate one proportion, with finite population correction
Enter the proportion you expect to find and how precisely you want to estimate it; get the sample size, the effect of a finite population, the recruitment needed if you expect losses, and the precision you would reach with a sample size you already have.
https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-una-proporcion
This link does not include the pasted data: they are too long for a URL.
Results
Normal quantile (z)
1.960
95%
Uncorrected size (n₀)
323
exact value: 322.68
Required size (n)
278
exact value: 277.97
Recruitment with losses
309
exact value: 308.86
Precision reachable with the available n
not defined
enter an available sample size to compute it
Interpretation
With 278 participants, an observed proportion near 30.0% would be estimated with a 95% confidence interval of about ±5.0%.
The source population has 2,000 people: the finite population correction lowers the requirement from 323 to 278 participants.
With 10% expected losses, you need to recruit 309 participants so that 278 finish.
If you already have a given number of participants, enter it under "Sample size already available (n, optional)" to see how precisely you could estimate the proportion.
- The figures depend on the assumptions (expected proportion, precision) taken from previous studies or pilots; document where they come from.
- The declared population (N = 2000) is less than ten times the uncorrected size: the finite population correction changes the result substantially. Check that N really is the frame you will sample from.
Explanation
When the aim of the study is descriptive (what proportion of febrile patients has dengue, what percentage of a cohort develops a complication), sample size is not computed to detect a difference but to make the confidence interval narrow enough. The absolute precision d is the half-width of that interval: with d = 0.05 the proportion will be estimated within ±5 percentage points.
The formula needs an expected proportion p, taken from previous studies or a pilot. Because the variance of a proportion is p(1−p), the required size is largest when p = 0.5: with no prior information, using 0.5 is the conservative choice, since no other proportion will demand a larger sample. The further p is from 0.5, the fewer participants are needed for the same precision.
If the population being sampled is finite and known (patients seen in a year, records in a department), the finite population correction reduces the size: sampling 300 out of 2,000 people carries more information than sampling 300 out of a million. The correction matters when the population is less than ten times the uncorrected size; with large populations it is irrelevant and can be left out.
Two adjustments close the calculation. Expected losses L raise recruitment to n/(1−L), so that the planned sample is the one that finishes. And the inverse mode answers the opposite question, the most common one in practice: if I already have a given number of participants, how precisely can I estimate the proportion? Every sample size is rounded up once, at the end.
Equations
- expected proportion in the population (0.5 if unknown)
- absolute precision: half-width of the confidence interval
- normal quantile for the confidence level (1.96 at 95%)
- size of the finite population being sampled
- expected proportion of losses (0 to 0.5)
- uncorrected size equivalent to n in a population of N
R code
# Sample size to estimate one proportion (absolute precision) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
p <- 0.3; d <- 0.05; nivel <- 0.95
poblacion <- 2000 # size of the finite population; 0 = population not bounded
perdidas <- 0.1 # expected losses to follow-up, 0 to 0.5
n_dado <- 0 # inverse mode: sample size already available; 0 = not used
z <- qnorm(1 - (1 - nivel) / 2)
n0 <- z^2 * p * (1 - p) / d^2 # Cochran 1977
n <- if (poblacion >= 2) n0 / (1 + (n0 - 1) / poblacion) else n0 # finite population correction
n_ajustado <- n / (1 - perdidas) # Lwanga & Lemeshow 1991
# Inverse mode: the very same n(d) solved for d, with the same correction
n0_dado <- if (poblacion >= 2) n_dado * (poblacion - 1) / (poblacion - n_dado) else n_dado
d_dado <- if (n_dado >= 2) z * sqrt(p * (1 - p) / n0_dado) else NA_real_
res <- list(z = z, n0 = n0, n = n, n_ajustado = n_ajustado, d_dado = d_dado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# The ceiling is applied once, by the interface: ceiling(n0), ceiling(n), ceiling(n_ajustado).
# Equivalent in RStudio (not run in the browser):
# presize::prec_prop(p, conf.width = 2 * d, conf.level = nivel, method = "wald")
# epiR::epi.sssimpleestb(N = poblacion, Py = p, epsilon = d, error = "absolute", conf.level = nivel)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
Sample size was computed to estimate a proportion with absolute precision, assuming an expected proportion of 30.0% and an interval half-width of 5.0%, with a confidence level of 95%, using Cochran's normal formula [1], corrected for a finite population of 2,000 people; 10% was added for expected losses [2], giving a recruitment target of 309 participants. Calculations used the "Sample size for one proportion" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-una-proporcion), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Cochran WG. Sampling Techniques. 3rd ed. New York: John Wiley & Sons; 1977. Original source
- 02 Lwanga SK, Lemeshow S. Sample Size Determination in Health Studies: A Practical Manual. Geneva: World Health Organization; 1991. Original source
- 03 Newcombe RG. Two-sided confidence intervals for the single proportion: comparison of seven methods. Statistics in Medicine. 1998;17(8):857–872. doi:10.1002/(SICI)1097-0258(19980430)17:8<857::AID-SIM777>3.0.CO;2-E PMID: 9595616 Complementary
- 04 Hulley SB, Cummings SR, Browner WS, Grady DG, Newman TB. Designing Clinical Research. 4th ed. Philadelphia: Lippincott Williams & Wilkins; 2013. Didactic reading
- 05 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading