← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

Test result Reference standard Total
Diseased Not diseased
Positive 74
Negative 126
Total 80 120 200
  1. True positives (TP)
  2. False positives (FP)
  3. False negatives (FN)
  4. True negatives (TN)

Wilson is the recommended method (Newcombe 1998). Wald is for teaching purposes only.

Use it when a cell is 0: adds 0.5 to the four cells before computing LR+, LR− and DOR (it does not affect Sn, Sp or the predictive values).

Example loaded

Illustrative example: rapid NS1 antigen test versus RT-PCR for dengue in 200 febrile patients: 68 true positives, 6 false positives, 12 false negatives and 114 true negatives (fictitious data).

Illustrative data, not real.

Results

Total participants (n)

200

Sensitivity (Sn)

85.0%

75.6% to 91.2%

95% CI · Wilson (score)

Specificity (Sp)

95.0%

89.5% to 97.7%

95% CI · Wilson (score)

Positive predictive value (PPV)

91.9%

83.4% to 96.2%

95% CI · Wilson (score)

Negative predictive value (NPV)

90.5%

84.1% to 94.5%

95% CI · Wilson (score)

Sample prevalence

40.0%

33.5% to 46.9%

95% CI · Wilson (score)

Accuracy

91.0%

86.2% to 94.2%

95% CI · Wilson (score)

Youden's index (J)

0.80

0.71 to 0.89

95% CI · delta method

LR+

17.00

7.75 to 37.28

95% CI · Simel's log method

LR−

0.16

0.0936 to 0.27

95% CI · Simel's log method

Diagnostic odds ratio (DOR)

107.67

38.63 to 300.07

95% CI · Woolf's log method

Interpretation

Of 200 participants, 68 + 12 had the disease and 6 + 114 did not (sample prevalence: 40.0%). The test detects 85.0% of the people with the disease (sensitivity; 95% CI: 75.6% to 91.2%) and correctly rules out 95.0% of those without it (specificity; 95% CI: 89.5% to 97.7%). Overall accuracy: 91.0% (95% CI: 86.2% to 94.2%).

LR+ = 17.00 (95% CI: 7.75 to 37.28): a positive result multiplies the pre-test odds by 17.00, a large and often conclusive change in the probability of disease (Jaeschke 1994).

LR− = 0.16 (95% CI: 0.0936 to 0.27): a negative result moderately lowers the probability of disease (Jaeschke 1994).

In this sample, 91.9% of positives had the disease (PPV; 95% CI: 83.4% to 96.2%) and 90.5% of negatives did not (NPV; 95% CI: 84.1% to 94.5%). Both figures depend on prevalence (40.0%) and do not transfer to other populations; in a case-control design they are not valid. For another prevalence use the "Predictive values" calculator.

DOR = 107.67 (95% CI: 38.63 to 300.07): the odds of a positive result are 107.67 times higher in people with the disease than in those without it. Youden's index J = 0.80 (95% CI: 0.71 to 0.89); J = 1 would be a perfect test and J = 0 a worthless one.

Estimates with their confidence interval: proportions (top) and ratios on a log scale (bottom)Sensitivity (Sn): 85.0% (75.6% to 91.2%); Specificity (Sp): 95.0% (89.5% to 97.7%); LR+: 17.00; LR−: 0.16Sensitivity (Sn)Specificity (Sp)Positive predictive value(PPV)Negative predictive value(NPV)Accuracy0%20%40%60%80%100%Ratios (log scale; reference at 1)LR+LR−Diagnostic odds ratio(DOR)0.010.11101001,000
Estimates with their confidence interval: proportions (top) and ratios on a log scale (bottom)

Explanation

A diagnostic test is evaluated against a reference standard in the same population. The 2×2 table crosses the test result (positive or negative) with the true state (diseased or not): true positives (TP), false positives (FP), false negatives (FN) and true negatives (TN).

Sensitivity is the proportion of diseased people with a positive test and specificity the proportion of non-diseased people with a negative test (Yerushalmy 1947). Both describe the test and do not depend on prevalence. Predictive values answer the clinical question (how likely is disease given a positive result?) but change with prevalence (Vecchio 1966): the ones shown here hold for this sample's prevalence.

Likelihood ratios (LR+ and LR−) summarise both properties in a single number per result and are what carries over to the post-test probability through Bayes' theorem (Fagan's nomogram). The diagnostic odds ratio (DOR) condenses overall discrimination into one figure, useful for comparing tests, not for deciding about a patient.

Each proportion carries the confidence interval of the chosen method (Wilson by default, as recommended by Newcombe 1998); the ratios carry the logarithmic intervals of Simel (1991) and Woolf (1955). When a cell is 0 a ratio becomes 0 or infinite and has no interval; the Haldane-Anscombe correction (adding 0.5 to each cell) yields a finite estimate and must be reported.

Equations

Sn=aa+c,Sp=db+d,PPV=aa+b,NPV=dc+d,Accuracy=a+dn\mathrm{Sn}=\frac{a}{a+c},\qquad \mathrm{Sp}=\frac{d}{b+d},\qquad \mathrm{PPV}=\frac{a}{a+b},\qquad \mathrm{NPV}=\frac{d}{c+d},\qquad \mathrm{Accuracy}=\frac{a+d}{n}
aa
true positives (TP)
bb
false positives (FP)
cc
false negatives (FN)
dd
true negatives (TN)
nn
total, a + b + c + d
Sensitivity and specificity (Yerushalmy 1947); predictive values at the sample prevalence, (a + c)/n.
LR+=Sn1−Sp,LR−=1−SnSp,DOR=adbc=LR+LR−\mathrm{LR}^{+}=\frac{\mathrm{Sn}}{1-\mathrm{Sp}},\qquad \mathrm{LR}^{-}=\frac{1-\mathrm{Sn}}{\mathrm{Sp}},\qquad \mathrm{DOR}=\frac{ad}{bc}=\frac{\mathrm{LR}^{+}}{\mathrm{LR}^{-}}
Likelihood ratios and diagnostic odds ratio (Glas 2003).
CI(LR+)=exp⁡ ⁣[ln⁡LR+±z1−α/21a−1a+c+1b−1b+d],CI(LR−)=exp⁡ ⁣[ln⁡LR−±z1c−1a+c+1d−1b+d]\mathrm{CI}(\mathrm{LR}^{+})=\exp\!\left[\ln \mathrm{LR}^{+}\pm z_{1-\alpha/2}\sqrt{\frac{1}{a}-\frac{1}{a+c}+\frac{1}{b}-\frac{1}{b+d}}\right],\qquad \mathrm{CI}(\mathrm{LR}^{-})=\exp\!\left[\ln \mathrm{LR}^{-}\pm z\sqrt{\frac{1}{c}-\frac{1}{a+c}+\frac{1}{d}-\frac{1}{b+d}}\right]
z1−α/2z_{1-\alpha/2}
normal quantile for the confidence level (1.96 at 95%)
Logarithmic intervals of Simel, Samsa and Matchar (1991). With the Haldane-Anscombe correction, a, b, c and d carry 0.5 added.
CI(DOR)=exp⁡ ⁣[ln⁡DOR±z1a+1b+1c+1d]\mathrm{CI}(\mathrm{DOR})=\exp\!\left[\ln \mathrm{DOR}\pm z\sqrt{\frac{1}{a}+\frac{1}{b}+\frac{1}{c}+\frac{1}{d}}\right]
Woolf's (1955) logarithmic interval.
J=Sn+Sp−1,CI(J)=J±zSn (1−Sn)a+c+Sp (1−Sp)b+dJ=\mathrm{Sn}+\mathrm{Sp}-1,\qquad \mathrm{CI}(J)=J\pm z\sqrt{\frac{\mathrm{Sn}\,(1-\mathrm{Sn})}{a+c}+\frac{\mathrm{Sp}\,(1-\mathrm{Sp})}{b+d}}
Youden's index (1950) with a delta-method Wald interval.
p^+z22m±zp^ (1−p^)m+z24m21+z2m\frac{\hat p+\frac{z^2}{2m}\pm z\sqrt{\frac{\hat p\,(1-\hat p)}{m}+\frac{z^2}{4m^2}}}{1+\frac{z^2}{m}}
p^\hat p
observed proportion, x/m
mm
denominator of the proportion (a + c, b + d, a + b, c + d or n)
Wilson (1927) interval, the default for every proportion; the selector offers Clopper-Pearson, Agresti-Coull, Jeffreys and Wald (see the "CI for a proportion" calculator).

R code

# Diagnostic test accuracy from a 2x2 table - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(binom)      # CIs for proportions: "wilson" (default), "exact" (Clopper-Pearson), "agresti-coull", "asymptotic" (Wald)
library(jsonlite)

vp <- 68; fp <- 6; fn <- 12; vn <- 114   # rows = test result (+, -), columns = reference standard (diseased, not diseased)
nivel  <- 0.95
metodo <- "wilson"   # CI method for proportions: "wilson", "clopper-pearson", "agresti-coull", "jeffreys" or "wald"
corr   <- 0     # 0.5 = Haldane-Anscombe correction for the ratios (LR, DOR) when a cell is 0; 0 = none
n <- vp + fp + fn + vn
z <- qnorm(1 - (1 - nivel) / 2)

# Proportion with CI -> c(estimate, lower, upper); Jeffreys is the equal-tailed interval (Brown, Cai & DasGupta 2001)
ic_prop <- function(x, m) {
  if (metodo == "jeffreys") {
    lo <- if (x == 0) 0 else qbeta((1 - nivel) / 2, x + 0.5, m - x + 0.5)
    hi <- if (x == m) 1 else qbeta(1 - (1 - nivel) / 2, x + 0.5, m - x + 0.5)
    return(c(x / m, lo, hi))
  }
  m_binom <- switch(metodo, wilson = "wilson", "clopper-pearson" = "exact", "agresti-coull" = "agresti-coull", wald = "asymptotic")
  ci <- binom.confint(x, m, conf.level = nivel, methods = m_binom)
  c(x / m, ci$lower, ci$upper)
}
sn   <- ic_prop(vp, vp + fn)    # sensitivity
sp   <- ic_prop(vn, vn + fp)    # specificity
vpp  <- ic_prop(vp, vp + fp)    # positive predictive value (at this sample's prevalence)
vpn  <- ic_prop(vn, vn + fn)    # negative predictive value
prev <- ic_prop(vp + fn, n)     # prevalence in the sample
exactitud <- ic_prop(vp + vn, n)   # accuracy

# Youden's J = Sn + Sp - 1 with a delta-method (Wald) CI
j    <- sn[1] + sp[1] - 1
j_ee <- sqrt(sn[1] * (1 - sn[1]) / (vp + fn) + sp[1] * (1 - sp[1]) / (vn + fp))
youden <- c(j, j - z * j_ee, j + z * j_ee)

# Ratios with the log-method CI: exp(log(est) -/+ z * SE); undefined (NA) when a cell of the (corrected) table is 0
a <- vp + corr; b <- fp + corr; c <- fn + corr; d <- vn + corr
ic_log <- function(est, ee) {
  if (is.finite(log(est)) && is.finite(ee)) c(est, exp(log(est) - z * ee), exp(log(est) + z * ee)) else c(est, NA, NA)
}
ee_lr <- function(x1, n1, x2, n2) sqrt((1 - x1 / n1) / x1 + (1 - x2 / n2) / x2)   # Simel, Samsa & Matchar 1991
lr_pos <- ic_log(a / (a + c) / (b / (b + d)), ee_lr(a, a + c, b, b + d))
lr_neg <- ic_log(c / (a + c) / (d / (b + d)), ee_lr(c, a + c, d, b + d))
dor    <- ic_log((a * d) / (b * c), sqrt(1 / a + 1 / b + 1 / c + 1 / d))         # Woolf 1955

res <- list(n = n, sn = sn, sp = sp, vpp = vpp, vpn = vpn, prev = prev, exactitud = exactitud,
            youden = youden, lr_pos = lr_pos, lr_neg = lr_neg, dor = dor)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio (not run in the browser):
# epiR::epi.tests(as.table(matrix(c(vp, fp, fn, vn), nrow = 2, byrow = TRUE)), method = "wilson", conf.level = nivel)

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

Sensitivity, specificity, predictive values, accuracy, Youden's index, likelihood ratios (LR+ and LR−) and the diagnostic odds ratio (DOR) were computed from the 2×2 table against the reference standard [1,2]. The 95% confidence intervals of the proportions were obtained with the Wilson (score) method [6,7]; those of the likelihood ratios with the logarithmic method of Simel et al. [3]; that of the DOR with Woolf's method [4,5], and that of Youden's index by the delta method. Calculations used the "Diagnostic test (2×2 table)" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/prueba-diagnostica-2x2), verified against R (binom package).

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Yerushalmy J. Statistical problems in assessing methods of medical diagnosis, with special reference to X-ray techniques. Public Health Reports. 1947;62(40):1432–1449. doi:10.2307/4586294 PMID: 20340527 Original source
  2. 02 Youden WJ. Index for rating diagnostic tests. Cancer. 1950;3(1):32–35. doi:10.1002/1097-0142(1950)3:1<32::AID-CNCR2820030106>3.0.CO;2-3 PMID: 15405679 Original source
  3. 03 Simel DL, Samsa GP, Matchar DB. Likelihood ratios with confidence: sample size estimation for diagnostic test studies. Journal of Clinical Epidemiology. 1991;44(8):763–770. doi:10.1016/0895-4356(91)90128-V PMID: 1941027 Original source
  4. 04 Woolf B. On estimating the relation between blood group and disease. Annals of Human Genetics. 1955;19(4):251–253. doi:10.1111/j.1469-1809.1955.tb01348.x PMID: 14388528 Original source
  5. 05 Glas AS, Lijmer JG, Prins MH, Bonsel GJ, Bossuyt PMM. The diagnostic odds ratio: a single indicator of test performance. Journal of Clinical Epidemiology. 2003;56(11):1129–1135. doi:10.1016/S0895-4356(03)00177-X PMID: 14615004 Complementary
  6. 06 Wilson EB. Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association. 1927;22(158):209–212. doi:10.1080/01621459.1927.10502953 Original source
  7. 07 Newcombe RG. Two-sided confidence intervals for the single proportion: comparison of seven methods. Statistics in Medicine. 1998;17(8):857–872. doi:10.1002/(SICI)1097-0258(19980430)17:8<857::AID-SIM777>3.0.CO;2-E PMID: 9595616 Original source
  8. 08 Clopper CJ, Pearson ES. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika. 1934;26(4):404–413. doi:10.1093/biomet/26.4.404 Complementary
  9. 09 Agresti A, Coull BA. Approximate is better than “exact” for interval estimation of binomial proportions. The American Statistician. 1998;52(2):119–126. doi:10.1080/00031305.1998.10480550 Complementary
  10. 10 Brown LD, Cai TT, DasGupta A. Interval estimation for a binomial proportion. Statistical Science. 2001;16(2):101–133. doi:10.1214/ss/1009213286 Complementary
  11. 11 Altman DG, Machin D, Bryant TN, Gardner MJ. Statistics with Confidence: Confidence Intervals and Statistical Guidelines. 2nd ed. London: BMJ Books; 2000. Complementary
  12. 12 Haldane JBS. The estimation and significance of the logarithm of a ratio of frequencies. Annals of Human Genetics. 1956;20(4):309–311. doi:10.1111/j.1469-1809.1955.tb01285.x PMID: 13314400 Original source
  13. 13 Anscombe FJ. On estimating binomial response relations. Biometrika. 1956;43(3-4):461–464. doi:10.1093/biomet/43.3-4.461 Original source
  14. 14 Jaeschke R, Guyatt GH, Sackett DL. Users' guides to the medical literature. III. How to use an article about a diagnostic test. B. What are the results and will they help me in caring for my patients? JAMA. 1994;271(9):703–707. doi:10.1001/jama.271.9.703 PMID: 8309035 Didactic reading
  15. 15 Altman DG, Bland JM. Diagnostic tests 1: sensitivity and specificity. BMJ. 1994;308(6943):1552. doi:10.1136/bmj.308.6943.1552 PMID: 8019315 Didactic reading
  16. 16 Vecchio TJ. Predictive value of a single diagnostic test in unselected populations. New England Journal of Medicine. 1966;274(21):1171–1173. doi:10.1056/NEJM196605262742104 PMID: 5934954 Didactic reading
  17. 17 Deeks JJ, Altman DG. Diagnostic tests 4: likelihood ratios. BMJ. 2004;329(7458):168–169. doi:10.1136/bmj.329.7458.168 PMID: 15258077 Didactic reading