Tools · Open Biostatistics
Diagnostic test (2×2 table): sensitivity, specificity, predictive values, likelihood ratios and DOR
Enter the four counts of the 2×2 table (test versus reference standard) and get Sn, Sp, PPV, NPV, prevalence, accuracy, Youden's index, LR+, LR− and DOR with their confidence intervals, a plain-language interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/prueba-diagnostica-2x2
This link does not include the pasted data: they are too long for a URL.
Results
Total participants (n)
200
Sensitivity (Sn)
85.0%
75.6% to 91.2%
95% CI · Wilson (score)
Specificity (Sp)
95.0%
89.5% to 97.7%
95% CI · Wilson (score)
Positive predictive value (PPV)
91.9%
83.4% to 96.2%
95% CI · Wilson (score)
Negative predictive value (NPV)
90.5%
84.1% to 94.5%
95% CI · Wilson (score)
Sample prevalence
40.0%
33.5% to 46.9%
95% CI · Wilson (score)
Accuracy
91.0%
86.2% to 94.2%
95% CI · Wilson (score)
Youden's index (J)
0.80
0.71 to 0.89
95% CI · delta method
LR+
17.00
7.75 to 37.28
95% CI · Simel's log method
LR−
0.16
0.0936 to 0.27
95% CI · Simel's log method
Diagnostic odds ratio (DOR)
107.67
38.63 to 300.07
95% CI · Woolf's log method
Interpretation
Of 200 participants, 68 + 12 had the disease and 6 + 114 did not (sample prevalence: 40.0%). The test detects 85.0% of the people with the disease (sensitivity; 95% CI: 75.6% to 91.2%) and correctly rules out 95.0% of those without it (specificity; 95% CI: 89.5% to 97.7%). Overall accuracy: 91.0% (95% CI: 86.2% to 94.2%).
LR+ = 17.00 (95% CI: 7.75 to 37.28): a positive result multiplies the pre-test odds by 17.00, a large and often conclusive change in the probability of disease (Jaeschke 1994).
LR− = 0.16 (95% CI: 0.0936 to 0.27): a negative result moderately lowers the probability of disease (Jaeschke 1994).
In this sample, 91.9% of positives had the disease (PPV; 95% CI: 83.4% to 96.2%) and 90.5% of negatives did not (NPV; 95% CI: 84.1% to 94.5%). Both figures depend on prevalence (40.0%) and do not transfer to other populations; in a case-control design they are not valid. For another prevalence use the "Predictive values" calculator.
DOR = 107.67 (95% CI: 38.63 to 300.07): the odds of a positive result are 107.67 times higher in people with the disease than in those without it. Youden's index J = 0.80 (95% CI: 0.71 to 0.89); J = 1 would be a perfect test and J = 0 a worthless one.
Explanation
A diagnostic test is evaluated against a reference standard in the same population. The 2×2 table crosses the test result (positive or negative) with the true state (diseased or not): true positives (TP), false positives (FP), false negatives (FN) and true negatives (TN).
Sensitivity is the proportion of diseased people with a positive test and specificity the proportion of non-diseased people with a negative test (Yerushalmy 1947). Both describe the test and do not depend on prevalence. Predictive values answer the clinical question (how likely is disease given a positive result?) but change with prevalence (Vecchio 1966): the ones shown here hold for this sample's prevalence.
Likelihood ratios (LR+ and LR−) summarise both properties in a single number per result and are what carries over to the post-test probability through Bayes' theorem (Fagan's nomogram). The diagnostic odds ratio (DOR) condenses overall discrimination into one figure, useful for comparing tests, not for deciding about a patient.
Each proportion carries the confidence interval of the chosen method (Wilson by default, as recommended by Newcombe 1998); the ratios carry the logarithmic intervals of Simel (1991) and Woolf (1955). When a cell is 0 a ratio becomes 0 or infinite and has no interval; the Haldane-Anscombe correction (adding 0.5 to each cell) yields a finite estimate and must be reported.
Equations
- true positives (TP)
- false positives (FP)
- false negatives (FN)
- true negatives (TN)
- total, a + b + c + d
- normal quantile for the confidence level (1.96 at 95%)
- observed proportion, x/m
- denominator of the proportion (a + c, b + d, a + b, c + d or n)
R code
# Diagnostic test accuracy from a 2x2 table - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(binom) # CIs for proportions: "wilson" (default), "exact" (Clopper-Pearson), "agresti-coull", "asymptotic" (Wald)
library(jsonlite)
vp <- 68; fp <- 6; fn <- 12; vn <- 114 # rows = test result (+, -), columns = reference standard (diseased, not diseased)
nivel <- 0.95
metodo <- "wilson" # CI method for proportions: "wilson", "clopper-pearson", "agresti-coull", "jeffreys" or "wald"
corr <- 0 # 0.5 = Haldane-Anscombe correction for the ratios (LR, DOR) when a cell is 0; 0 = none
n <- vp + fp + fn + vn
z <- qnorm(1 - (1 - nivel) / 2)
# Proportion with CI -> c(estimate, lower, upper); Jeffreys is the equal-tailed interval (Brown, Cai & DasGupta 2001)
ic_prop <- function(x, m) {
if (metodo == "jeffreys") {
lo <- if (x == 0) 0 else qbeta((1 - nivel) / 2, x + 0.5, m - x + 0.5)
hi <- if (x == m) 1 else qbeta(1 - (1 - nivel) / 2, x + 0.5, m - x + 0.5)
return(c(x / m, lo, hi))
}
m_binom <- switch(metodo, wilson = "wilson", "clopper-pearson" = "exact", "agresti-coull" = "agresti-coull", wald = "asymptotic")
ci <- binom.confint(x, m, conf.level = nivel, methods = m_binom)
c(x / m, ci$lower, ci$upper)
}
sn <- ic_prop(vp, vp + fn) # sensitivity
sp <- ic_prop(vn, vn + fp) # specificity
vpp <- ic_prop(vp, vp + fp) # positive predictive value (at this sample's prevalence)
vpn <- ic_prop(vn, vn + fn) # negative predictive value
prev <- ic_prop(vp + fn, n) # prevalence in the sample
exactitud <- ic_prop(vp + vn, n) # accuracy
# Youden's J = Sn + Sp - 1 with a delta-method (Wald) CI
j <- sn[1] + sp[1] - 1
j_ee <- sqrt(sn[1] * (1 - sn[1]) / (vp + fn) + sp[1] * (1 - sp[1]) / (vn + fp))
youden <- c(j, j - z * j_ee, j + z * j_ee)
# Ratios with the log-method CI: exp(log(est) -/+ z * SE); undefined (NA) when a cell of the (corrected) table is 0
a <- vp + corr; b <- fp + corr; c <- fn + corr; d <- vn + corr
ic_log <- function(est, ee) {
if (is.finite(log(est)) && is.finite(ee)) c(est, exp(log(est) - z * ee), exp(log(est) + z * ee)) else c(est, NA, NA)
}
ee_lr <- function(x1, n1, x2, n2) sqrt((1 - x1 / n1) / x1 + (1 - x2 / n2) / x2) # Simel, Samsa & Matchar 1991
lr_pos <- ic_log(a / (a + c) / (b / (b + d)), ee_lr(a, a + c, b, b + d))
lr_neg <- ic_log(c / (a + c) / (d / (b + d)), ee_lr(c, a + c, d, b + d))
dor <- ic_log((a * d) / (b * c), sqrt(1 / a + 1 / b + 1 / c + 1 / d)) # Woolf 1955
res <- list(n = n, sn = sn, sp = sp, vpp = vpp, vpn = vpn, prev = prev, exactitud = exactitud,
youden = youden, lr_pos = lr_pos, lr_neg = lr_neg, dor = dor)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run in the browser):
# epiR::epi.tests(as.table(matrix(c(vp, fp, fn, vn), nrow = 2, byrow = TRUE)), method = "wilson", conf.level = nivel)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
Sensitivity, specificity, predictive values, accuracy, Youden's index, likelihood ratios (LR+ and LR−) and the diagnostic odds ratio (DOR) were computed from the 2×2 table against the reference standard [1,2]. The 95% confidence intervals of the proportions were obtained with the Wilson (score) method [6,7]; those of the likelihood ratios with the logarithmic method of Simel et al. [3]; that of the DOR with Woolf's method [4,5], and that of Youden's index by the delta method. Calculations used the "Diagnostic test (2×2 table)" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/prueba-diagnostica-2x2), verified against R (binom package).
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Yerushalmy J. Statistical problems in assessing methods of medical diagnosis, with special reference to X-ray techniques. Public Health Reports. 1947;62(40):1432–1449. doi:10.2307/4586294 PMID: 20340527 Original source
- 02 Youden WJ. Index for rating diagnostic tests. Cancer. 1950;3(1):32–35. doi:10.1002/1097-0142(1950)3:1<32::AID-CNCR2820030106>3.0.CO;2-3 PMID: 15405679 Original source
- 03 Simel DL, Samsa GP, Matchar DB. Likelihood ratios with confidence: sample size estimation for diagnostic test studies. Journal of Clinical Epidemiology. 1991;44(8):763–770. doi:10.1016/0895-4356(91)90128-V PMID: 1941027 Original source
- 04 Woolf B. On estimating the relation between blood group and disease. Annals of Human Genetics. 1955;19(4):251–253. doi:10.1111/j.1469-1809.1955.tb01348.x PMID: 14388528 Original source
- 05 Glas AS, Lijmer JG, Prins MH, Bonsel GJ, Bossuyt PMM. The diagnostic odds ratio: a single indicator of test performance. Journal of Clinical Epidemiology. 2003;56(11):1129–1135. doi:10.1016/S0895-4356(03)00177-X PMID: 14615004 Complementary
- 06 Wilson EB. Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association. 1927;22(158):209–212. doi:10.1080/01621459.1927.10502953 Original source
- 07 Newcombe RG. Two-sided confidence intervals for the single proportion: comparison of seven methods. Statistics in Medicine. 1998;17(8):857–872. doi:10.1002/(SICI)1097-0258(19980430)17:8<857::AID-SIM777>3.0.CO;2-E PMID: 9595616 Original source
- 08 Clopper CJ, Pearson ES. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika. 1934;26(4):404–413. doi:10.1093/biomet/26.4.404 Complementary
- 09 Agresti A, Coull BA. Approximate is better than “exact” for interval estimation of binomial proportions. The American Statistician. 1998;52(2):119–126. doi:10.1080/00031305.1998.10480550 Complementary
- 10 Brown LD, Cai TT, DasGupta A. Interval estimation for a binomial proportion. Statistical Science. 2001;16(2):101–133. doi:10.1214/ss/1009213286 Complementary
- 11 Altman DG, Machin D, Bryant TN, Gardner MJ. Statistics with Confidence: Confidence Intervals and Statistical Guidelines. 2nd ed. London: BMJ Books; 2000. Complementary
- 12 Haldane JBS. The estimation and significance of the logarithm of a ratio of frequencies. Annals of Human Genetics. 1956;20(4):309–311. doi:10.1111/j.1469-1809.1955.tb01285.x PMID: 13314400 Original source
- 13 Anscombe FJ. On estimating binomial response relations. Biometrika. 1956;43(3-4):461–464. doi:10.1093/biomet/43.3-4.461 Original source
- 14 Jaeschke R, Guyatt GH, Sackett DL. Users' guides to the medical literature. III. How to use an article about a diagnostic test. B. What are the results and will they help me in caring for my patients? JAMA. 1994;271(9):703–707. doi:10.1001/jama.271.9.703 PMID: 8309035 Didactic reading
- 15 Altman DG, Bland JM. Diagnostic tests 1: sensitivity and specificity. BMJ. 1994;308(6943):1552. doi:10.1136/bmj.308.6943.1552 PMID: 8019315 Didactic reading
- 16 Vecchio TJ. Predictive value of a single diagnostic test in unselected populations. New England Journal of Medicine. 1966;274(21):1171–1173. doi:10.1056/NEJM196605262742104 PMID: 5934954 Didactic reading
- 17 Deeks JJ, Altman DG. Diagnostic tests 4: likelihood ratios. BMJ. 2004;329(7458):168–169. doi:10.1136/bmj.329.7458.168 PMID: 15258077 Didactic reading