Tools · Open Biostatistics
Association and effect in a 2×2 table: relative risk, odds ratio, ARR, RRR and NNT
Enter the four counts of the 2×2 table (exposure or treatment versus outcome) and get the risk in each group, the relative risk, the odds ratio, the absolute and relative risk reduction and the number needed to treat with their confidence intervals, a plain-language interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/efecto-2x2
This link does not include the pasted data: they are too long for a URL.
Results
Total participants (n)
200
Risk in the exposed / treated (p₁)
12.0%
7.0% to 19.8%
95% CI · Wilson's method
Risk in the unexposed / controls (p₀)
30.0%
21.9% to 39.6%
95% CI · Wilson's method
Relative risk (RR) / prevalence ratio in a cross-sectional design
0.40
0.22 to 0.74
95% CI · Katz's log method
Odds ratio (OR)
0.32
0.15 to 0.67
95% CI · Woolf's log method
Absolute risk reduction (ARR, percentage points)
-18.0
-28.8 to -6.7
95% CI · Newcombe hybrid (method 10)
Relative risk reduction (RRR)
60.0%
26.4% to 78.3%
95% CI · Katz's log method
Number needed to treat (NNT)
5.56
3.47 to 14.83
95% CI · reciprocal of the ARR limits (Altman)
Interpretation
Of 100 exposed or treated people, 12 had the outcome (12.0%; 95% CI: 7.0% to 19.8%); of 100 unexposed or controls, 30 had it (30.0%; 95% CI: 21.9% to 39.6%). The absolute risk difference is -18.0 percentage points (95% CI: -28.8 to -6.7) and the relative change in risk (RRR = 1 − RR) is 60.0% (95% CI: 26.4% to 78.3%).
RR = 0.40 (95% CI: 0.22 to 0.74): the risk in the exposed or treated group is 0.40 times that of the reference group and the interval excludes 1, so the data are hardly compatible with no association.
OR = 0.32 (95% CI: 0.15 to 0.67). The odds ratio compares the odds of the outcome between the exposed or treated group and the reference group: it approximates the relative risk only when the outcome is rare and, with common outcomes, it moves further from 1 and exaggerates the effect. With a cell at 0 it becomes 0 or infinite and has no interval, unless the Haldane-Anscombe correction is applied.
6 people have to be treated or exposed to avoid one additional event (NNTB; exact value 5.56; 95% CI: 3.47 to 14.83).
- The risk in the reference group is above 10%: the odds ratio moves away from the relative risk and exaggerates the effect. The OR approximates the RR only when the outcome is rare; report the relative risk as the main measure.
Explanation
An association 2×2 table crosses an exposure or treatment (rows) with an outcome (columns). It yields two risks: that of the exposed or treated, p₁ = a/(a + b), and that of the unexposed or controls, p₀ = c/(c + d). Everything else is a way of comparing those two numbers: dividing them (relative risk), subtracting them (absolute risk reduction) or comparing their odds (odds ratio).
The relative risk says how many times more (or less) frequent the outcome is in the exposed group; the absolute reduction says how many events are avoided (or caused) per 100 people, and its reciprocal is the number needed to treat: how many people must be treated or exposed for one extra event to happen or be prevented. The same 50% relative reduction can mean an NNT of 10 or of 1,000 depending on how common the outcome is: that is why the three measures are reported together.
The odds ratio is not a relative risk. They agree only when the outcome is rare (risk in the reference group below 10%); when the outcome is common the odds ratio moves away from 1 much more than the relative risk and exaggerates the effect. Design matters too: in a cohort or a clinical trial all five measures are interpretable; in a cross-sectional study the ratio of proportions is a prevalence ratio, not an incidence ratio; and in a case-control study the sampling fixes how many cases and how many controls there are, so risks, relative risk, absolute reduction and NNT cannot be estimated and only the odds ratio is valid.
When a cell is 0, the relative risk and the odds ratio become 0 or infinite and their logarithmic interval does not exist. The Haldane-Anscombe correction (adding 0.5 to the four cells) yields a finite estimate and must be declared in Methods; here it is an explicit option and affects only the relative risk, the odds ratio and the relative risk reduction, never the risks, the absolute difference or the NNT. If the interval of the absolute difference includes 0, the NNT has no two finite limits of the same sign and is reported with Altman's notation: "NNTB x to ∞ to NNTH y".
Equations
- exposed or treated with the outcome
- exposed or treated without the outcome
- unexposed or controls with the outcome
- unexposed or controls without the outcome
- total exposed or treated, a + b
- total unexposed or controls, c + d
- observed proportion, x/m (here p₁ or p₀)
- denominator of the proportion (n₁ or n₀)
- normal quantile for the confidence level (1.96 at 95%)
- two-sided normal quantile for the confidence level
- Wilson limits of p₁
- Wilson limits of p₀
- limits of the relative risk interval
- limits of the absolute risk reduction interval
R code
# Measures of association and effect from a 2x2 table - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
# rows = exposure or treatment, columns = outcome:
# a = exposed with the outcome, b = exposed without it, c = unexposed with it, d = unexposed without it
a <- 12; b <- 88; c <- 30; d <- 70
nivel <- 0.95
diseno <- "cohorte" # "cohorte", "casos_controles" or "transversal": changes only how the results are read
metodo_rra <- "newcombe" # CI for the absolute risk reduction: "newcombe" (hybrid score, method 10) or "wald"
corr <- 0 # 0.5 = Haldane-Anscombe correction for the ratios; it never touches p1, p0, ARR or NNT
n <- a + b + c + d
z <- qnorm(1 - (1 - nivel) / 2)
# Risks with Wilson's (1927) score interval, written out to keep the same order of operations
# as binom::binom.confint(methods = "wilson"); the correction is not applied here on purpose.
n1 <- a + b; n0 <- c + d
ic_wilson <- function(x, m) {
p <- x / m
w1 <- p + 0.5 * z^2 / m
w2 <- z * sqrt((p * (1 - p) + 0.25 * z^2 / m) / m)
w3 <- 1 + z^2 / m
c(p, (w1 - w2) / w3, (w1 + w2) / w3)
}
p1 <- ic_wilson(a, n1) # risk in the exposed / treated
p0 <- ic_wilson(c, n0) # risk in the unexposed / controls
# Ratios on the (possibly corrected) cells, with the log-method CI: exp(log(est) -/+ z * SE);
# undefined (NA) when log(est) or the SE is not finite, that is, when a cell of the table is 0.
ac <- a + corr; bc <- b + corr; cc <- c + corr; dc <- d + corr
n1c <- ac + bc; n0c <- cc + dc
ic_log <- function(est, ee) {
if (is.finite(log(est)) && is.finite(ee)) c(est, exp(log(est) - z * ee), exp(log(est) + z * ee)) else c(est, NA, NA)
}
rr <- ic_log((ac / n1c) / (cc / n0c), sqrt(1/ac - 1/n1c + 1/cc - 1/n0c)) # Katz 1978
or <- ic_log((ac * dc) / (bc * cc), sqrt(1/ac + 1/bc + 1/cc + 1/dc)) # Woolf 1955
# Absolute risk reduction: hybrid score interval (Newcombe 1998, method 10) or Wald
rra <- p1[1] - p0[1]
if (metodo_rra == "newcombe") {
rra_ic <- c(rra - sqrt((p1[1] - p1[2])^2 + (p0[3] - p0[1])^2),
rra + sqrt((p1[3] - p1[1])^2 + (p0[1] - p0[2])^2))
} else {
rra_ee <- sqrt(p1[1] * (1 - p1[1]) / n1 + p0[1] * (1 - p0[1]) / n0)
rra_ic <- c(rra - z * rra_ee, rra + z * rra_ee)
}
# Relative risk reduction and number needed to treat (reciprocal of the ARR and of its limits)
rrr <- c(1 - rr[1], 1 - rr[3], 1 - rr[2])
nnt <- 1 / abs(rra); nnt_ic <- sort(1 / abs(rra_ic)) # Altman 1998; if the ARR CI crosses 0, read it as NNTB x to Inf to NNTH y
res <- list(n = n, p1 = p1, p0 = p0, rr = rr, or = or,
rra = c(rra, rra_ic[1], rra_ic[2]), rrr = rrr, nnt = c(nnt, nnt_ic[1], nnt_ic[2]))
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Alternative (score interval, Miettinen-Nurminen type): PropCIs::diffscoreci(a, n1, c, n0, nivel)
# Equivalent in RStudio (not run in the browser):
# epiR::epi.2by2(as.table(matrix(c(a, b, c, d), nrow = 2, byrow = TRUE)), method = "cohort.count", conf.level = nivel)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The relative risk (95% CI by Katz's method) [3], the odds ratio (Woolf's method) [2], the absolute risk reduction (Newcombe's hybrid score interval, method 10 [4]) and the number needed to treat (reciprocal of the ARR and of its limits [5,6]) were estimated. Calculations used the "Association and effect (2×2 table)" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/efecto-2x2), verified against R.
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Cornfield J. A method of estimating comparative rates from clinical data. Applications to cancer of the lung, breast, and cervix. Journal of the National Cancer Institute. 1951;11(6):1269–1275. doi:10.1093/jnci/11.6.1269 PMID: 14861651 Original source
- 02 Woolf B. On estimating the relation between blood group and disease. Annals of Human Genetics. 1955;19(4):251–253. doi:10.1111/j.1469-1809.1955.tb01348.x PMID: 14388528 Original source
- 03 Katz D, Baptista J, Azen SP, Pike MC. Obtaining confidence intervals for the risk ratio in cohort studies. Biometrics. 1978;34(3):469–474. doi:10.2307/2530610 Original source
- 04 Newcombe RG. Interval estimation for the difference between independent proportions: comparison of eleven methods. Statistics in Medicine. 1998;17(8):873–890. doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I PMID: 9595617 Original source
- 05 Laupacis A, Sackett DL, Roberts RS. An assessment of clinically useful measures of the consequences of treatment. New England Journal of Medicine. 1988;318(26):1728–1733. doi:10.1056/NEJM198806303182605 PMID: 3374545 Original source
- 06 Altman DG. Confidence intervals for the number needed to treat. BMJ. 1998;317(7168):1309–1312. doi:10.1136/bmj.317.7168.1309 PMID: 9804726 Original source
- 07 Cook RJ, Sackett DL. The number needed to treat: a clinically useful measure of treatment effect. BMJ. 1995;310(6977):452–454. doi:10.1136/bmj.310.6977.452 PMID: 7873954 Complementary
- 08 Haldane JBS. The estimation and significance of the logarithm of a ratio of frequencies. Annals of Human Genetics. 1956;20(4):309–311. doi:10.1111/j.1469-1809.1955.tb01285.x PMID: 13314400 Original source
- 09 Anscombe FJ. On estimating binomial response relations. Biometrika. 1956;43(3-4):461–464. doi:10.1093/biomet/43.3-4.461 Original source
- 10 Wilson EB. Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association. 1927;22(158):209–212. doi:10.1080/01621459.1927.10502953 Complementary
- 11 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading
- 12 Altman DG, Machin D, Bryant TN, Gardner MJ. Statistics with Confidence: Confidence Intervals and Statistical Guidelines. 2nd ed. London: BMJ Books; 2000. Didactic reading