Tools · Open Biostatistics
Confidence interval for a proportion: Wilson, Clopper-Pearson, Agresti-Coull, Jeffreys and Wald
Enter how many cases show the characteristic (x) out of a total (n) and get the proportion with its confidence interval by six methods, a comparison among them, a plain-language interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/ic-proporcion
This link does not include the pasted data: they are too long for a URL.
Results
Observed proportion
85.0%
Wilson (score)
85.0%
75.6% to 91.2%
Interval width: 15.6%
Wilson with continuity correction
85.0%
74.9% to 91.7%
Interval width: 16.8%
Clopper-Pearson (exact)
85.0%
75.3% to 92.0%
Interval width: 16.7%
Agresti-Coull
85.0%
75.4% to 91.4%
Interval width: 15.9%
Jeffreys
85.0%
76.0% to 91.5%
Interval width: 15.5%
Wald (asymptotic)
85.0%
77.2% to 92.8%
Interval width: 15.6%
Interpretation
68 of 80 observations show the characteristic: 85.0%. With 95% confidence, the population proportion lies between 75.6% and 91.2% (Wilson method, recommended).
With n = 80 and a proportion away from 0 and 1, the six methods agree almost completely: widths range from 15.5% to 16.8% percentage points. The choice of method barely changes the conclusion.
Explanation
An observed proportion (x of n) is only an estimate of the proportion in the population. The confidence interval expresses the uncertainty of that estimate: if the study were repeated many times, 95% of the intervals built this way would contain the true proportion.
The Wald interval, the one in almost every textbook, misbehaves when the sample is small or the proportion is close to 0 or 1: it can fall outside [0, 1] and, with x = 0 or x = n, it collapses to a single point. That is why current recommendations (Newcombe 1998; Brown, Cai and DasGupta 2001) favour the Wilson method, or Jeffreys, and reserve Clopper-Pearson for when guaranteed coverage is required.
This calculator shows all six methods side by side so you can see where they agree and where they do not; the width of each interval lets you compare them at a glance.
Equations
- observed proportion, x/n
- normal quantile for the confidence level (1.96 at 95%)
- sample size
- quantile q of the beta distribution with parameters a and b
R code
# Confidence interval for a proportion (six methods) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(binom) # Wilson (default), Clopper-Pearson ("exact"), Agresti-Coull and Wald ("asymptotic")
library(jsonlite)
x <- 68; n <- 80; nivel <- 0.95
alfa <- 1 - nivel; z <- qnorm(1 - alfa / 2); p <- x / n
ic <- function(m) { r <- binom.confint(x, n, conf.level = nivel, methods = m); c(p, r$lower, r$upper) }
wilson <- ic("wilson")
clopper_pearson <- ic("exact")
agresti_coull <- ic("agresti-coull")
wald <- ic("asymptotic")
# Wilson with continuity correction (Newcombe 1998, method 4); 0 and 1 at the boundaries
wcc_lo <- if (x == 0) 0 else max(0, (2 * n * p + z^2 - 1 - z * sqrt(z^2 - 2 - 1 / n + 4 * p * (n * (1 - p) + 1))) / (2 * (n + z^2)))
wcc_hi <- if (x == n) 1 else min(1, (2 * n * p + z^2 + 1 + z * sqrt(z^2 + 2 - 1 / n + 4 * p * (n * (1 - p) - 1))) / (2 * (n + z^2)))
wilson_cc <- c(p, wcc_lo, wcc_hi)
# Jeffreys, equal-tailed (Brown, Cai & DasGupta 2001); binom's "bayes" method is HPD and differs
jeffreys <- c(p, if (x == 0) 0 else qbeta(alfa / 2, x + 0.5, n - x + 0.5),
if (x == n) 1 else qbeta(1 - alfa / 2, x + 0.5, n - x + 0.5))
res <- list(p = p, wilson = wilson, wilson_cc = wilson_cc, clopper_pearson = clopper_pearson,
agresti_coull = agresti_coull, jeffreys = jeffreys, wald = wald)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run in the browser):
# DescTools::BinomCI(x, n, conf.level = nivel, method = c("wilson", "wilsoncc", "clopper-pearson", "agresti-coull", "jeffreys", "wald"))
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The 95% confidence interval for the proportion (68/80 = 85.0%) was computed with the Wilson method [1,4]: 75.6% to 91.2%. Alternative methods (Clopper-Pearson [2], Agresti-Coull [3], Jeffreys [5]) were computed for comparison. Calculations used the "CI for a proportion" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/ic-proporcion), verified against R (binom package).
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Wilson EB. Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association. 1927;22(158):209–212. doi:10.1080/01621459.1927.10502953 Original source
- 02 Clopper CJ, Pearson ES. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika. 1934;26(4):404–413. doi:10.1093/biomet/26.4.404 Original source
- 03 Agresti A, Coull BA. Approximate is better than “exact” for interval estimation of binomial proportions. The American Statistician. 1998;52(2):119–126. doi:10.1080/00031305.1998.10480550 Original source
- 04 Newcombe RG. Two-sided confidence intervals for the single proportion: comparison of seven methods. Statistics in Medicine. 1998;17(8):857–872. doi:10.1002/(SICI)1097-0258(19980430)17:8<857::AID-SIM777>3.0.CO;2-E PMID: 9595616 Original source
- 05 Brown LD, Cai TT, DasGupta A. Interval estimation for a binomial proportion. Statistical Science. 2001;16(2):101–133. doi:10.1214/ss/1009213286 Complementary
- 06 Altman DG, Machin D, Bryant TN, Gardner MJ. Statistics with Confidence: Confidence Intervals and Statistical Guidelines. 2nd ed. London: BMJ Books; 2000. Didactic reading