Tools · Open Biostatistics
Sample size to estimate the sensitivity and specificity of a diagnostic test (Buderer 1996)
Enter the sensitivity and specificity you expect, the prevalence of the disease where the test will be used and the precision with which you want to estimate them, and get how many diseased people, how many non-diseased people and how many consecutive patients are needed, the interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-prueba-diagnostica
This link does not include the pasted data: they are too long for a URL.
Results
Normal quantile (z)
1.960
Diseased people needed
196
exact value 195.91
Patients to recruit for sensitivity
654
exact value 653.05
Non-diseased people needed
73
exact value 72.99
Patients to recruit for specificity
105
exact value 104.27
Patients to recruit (total)
654
exact value 653.05
Total with expected losses
726
exact value 725.61
Interpretation
To estimate a sensitivity near 85.0% with a 95% CI of ±0.05, 196 patients with the disease are needed, that is, 654 consecutive patients if prevalence is 30.0%; for specificity, 73 people without the disease are needed (105 patients). Recruiting 654 is recommended.
Sensitivity is the binding requirement: gathering 196 diseased people with a prevalence of 30.0% means assessing more patients (654) than the 105 that specificity asks for. With this design, specificity will be estimated more precisely than ±0.05.
Allowing for 10% losses or exclusions, 726 patients must be recruited in order to end with 654 evaluable ones.
- The figures depend entirely on the sensitivity, specificity and prevalence that are assumed: take them from previous studies or a pilot and document their source in the protocol.
- With a sensitivity or specificity of 95% or more, the Wald interval behind this formula is too narrow and the resulting size falls short: check the result with a Wilson or Clopper-Pearson interval.
Explanation
A diagnostic accuracy study estimates two different proportions in two different groups: sensitivity is estimated only among diseased people and specificity only among non-diseased people. That is why the sample size is not one number but two, and the study needs the larger of them.
The calculation has two steps (Buderer 1996). First it asks how many diseased people are needed for the confidence interval of sensitivity to have the desired half-width: that is the classic Wald interval for a proportion. The same, separately, for specificity among the non-diseased.
Second: in a consecutive series nobody recruits by disease status, everyone who arrives is assessed. If the expected prevalence is 30%, gathering 196 diseased people requires assessing about 654 patients. Prevalence is what turns "diseased people needed" into "patients to be recruited", and it is the reason why a rare disease makes estimating sensitivity well so expensive.
The figures depend entirely on the sensitivity, specificity and prevalence that are assumed: take them from a previous study, a pilot or the literature, and state them in the protocol. The interval behind this formula is Wald's; with values close to 1 it is worth checking the result with a Wilson or Clopper-Pearson interval, which is what will finally be reported.
Equations
- people with the disease needed to estimate sensitivity
- consecutive patients who must be assessed to gather those diseased people
- normal quantile of the confidence level (1.96 at 95%)
- half-width of the confidence interval (absolute precision)
- expected prevalence where the test will be applied
- people without the disease needed to estimate specificity
- consecutive patients who must be assessed to gather those non-diseased people
- study size: the larger of the two requirements
- expected losses, as a proportion
- ceiling: rounding up is applied only once, at the end
R code
# Sample size for sensitivity and specificity (Buderer 1996) - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
sn <- 0.85; sp <- 0.95 # expected sensitivity and specificity (strictly between 0 and 1)
prev <- 0.3 # expected prevalence where the test will be used
w <- 0.05 # half-width of the confidence interval (absolute precision)
nivel <- 0.95 # confidence level
perdidas <- 0.1 # expected losses (0-0.5)
z <- qnorm(1 - (1 - nivel) / 2)
# Buderer (1996): each half of the design is a Wald interval for a proportion.
n_d <- z^2 * sn * (1 - sn) / w^2 # diseased needed to estimate Sn with precision +-w
n_nd <- z^2 * sp * (1 - sp) / w^2 # non-diseased needed to estimate Sp with precision +-w
# In a consecutive series nobody recruits by disease status: the prevalence
# decides how many patients must be screened to reach each of those groups.
n_sn <- n_d / prev
n_sp <- n_nd / (1 - prev)
n_total <- max(n_sn, n_sp) # the binding requirement of the two
n_ajustado <- n_total / (1 - perdidas) # Lwanga & Lemeshow (1991)
# Values travel WITHOUT rounding: the interface takes the ceiling once, at the end.
res <- list(z = z, n_d = n_d, n_sn = n_sn, n_nd = n_nd, n_sp = n_sp,
n_total = n_total, n_ajustado = n_ajustado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio: presize::prec_sens(sens = sn, prev = prev, conf.width = 2*w, conf.level = nivel, method = "wald"); presize::prec_spec(spec = sp, prev = prev, conf.width = 2*w, conf.level = nivel, method = "wald")
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The sample size was computed to estimate sensitivity and specificity with an absolute precision of ±0.05 and 95% confidence, assuming a sensitivity of 85.0%, a specificity of 95.0% and a prevalence of 30.0%, using Buderer's formula [1]: 196 participants with the disease (654 consecutive patients) and 73 without it (105 patients); the larger of the two requirements was taken, 654 patients. An additional 10% was added for expected losses or exclusions, bringing planned recruitment to 726 patients. Calculations used the "Sample size for a diagnostic test" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/muestra-prueba-diagnostica), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Buderer NM. Statistical methodology: I. Incorporating the prevalence of disease into the sample size calculation for sensitivity and specificity. Academic Emergency Medicine. 1996;3(9):895–900. doi:10.1111/j.1553-2712.1996.tb03538.x PMID: 8870764 Original source
- 02 Simel DL, Samsa GP, Matchar DB. Likelihood ratios with confidence: sample size estimation for diagnostic test studies. Journal of Clinical Epidemiology. 1991;44(8):763–770. doi:10.1016/0895-4356(91)90128-V PMID: 1941027 Complementary
- 03 Bujang MA, Adnan TH. Requirements for minimum sample size for sensitivity and specificity analysis. Journal of Clinical and Diagnostic Research. 2016;10(10):YE01–YE06. doi:10.7860/JCDR/2016/18129.8744 PMID: 27891446 Didactic reading
- 04 Newcombe RG. Two-sided confidence intervals for the single proportion: comparison of seven methods. Statistics in Medicine. 1998;17(8):857–872. doi:10.1002/(SICI)1097-0258(19980430)17:8<857::AID-SIM777>3.0.CO;2-E PMID: 9595616 Complementary
- 05 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading