Tools · Open Biostatistics
Paired proportions: McNemar's test, paired difference and paired odds ratio
Enter the four cells of the paired 2×2 table (two tests or two time points on the same subjects) and get McNemar's test in its three versions, the paired difference of proportions with its confidence interval and the paired odds ratio, with a plain-language interpretation and the equivalent R code.
https://udgca1190.com.mx/en/herramientas/bioestadistica/mcnemar
This link does not include the pasted data: they are too long for a URL.
Results
Pairs analysed (n)
150
Discordant pairs (b + c)
20
the only ones that enter the test
Positive with test A
36.7%
Positive with test B
30.0%
McNemar's chi-squared (uncorrected)
5.000
1 degree of freedom · no correction
p of the uncorrected chi-squared
0.025
1 degree of freedom · no correction
Chi-squared with Edwards' correction
4.050
1 degree of freedom · Edwards' correction
p of the chi-squared with Edwards' correction
0.044
1 degree of freedom · Edwards' correction
p of the exact (binomial) test
0.041
two-sided binomial on b + c
Paired difference (b − c)/n, percentage points
6.7
0.9 to 12.4
95% CI · Wald
Paired odds ratio (b/c)
3.00
1.04 to 10.55
95% CI · derived from Clopper-Pearson
Interpretation
Among the 150 pairs, 15 changed in one direction and 5 in the other: test A was positive in 36.7% and test B in 30.0% (difference 6.7 percentage points, 95% CI: 0.9 to 12.4).
Changes were more frequent in the A+ / B− direction (15 pairs against 5): test A was positive more often than test B.
Exact McNemar test (two-sided binomial on the 20 discordant pairs): p = 0.041.
p = 0.041; at α = 0.05 the null hypothesis that both measurements give the same proportion of positives is rejected. The size of the change (6.7 percentage points, 95% CI: 0.9 to 12.4) matters more than the p value.
Paired odds ratio b/c = 3.00 (95% CI: 1.04 to 10.55): changes in one direction (15 pairs) were 3.00 times more frequent than the opposite ones (5 pairs).
The three versions of the test on the same 20 discordant pairs: uncorrected chi-squared = 5.000 (p = 0.025); chi-squared with Edwards' correction = 4.050 (p = 0.044); exact binomial, p = 0.041. Reported here: the exact binomial test.
- There are only 20 discordant pairs (fewer than 25): the exact binomial test is reported as the main one, because with so few the chi-squared approximation is unreliable.
- The test uses only the discordant pairs (20): pairs where both measurements agreed carry no information about change.
Explanation
When the two proportions being compared come from the same subjects (two tests on the same patient, or the same measurement before and after), the observations come in pairs and are not independent. The 2×2 table no longer counts people per group: it counts pairs. Cells a and d hold the pairs where both measurements agreed (concordant) and cells b and c the pairs where they disagreed (discordant).
McNemar's test (1947) looks only at the discordant pairs: if the two measurements were equivalent, a change in one direction would be as likely as the opposite one, that is, b and c would be two halves of b + c. Concordant pairs carry no information about change and so do not enter the statistic; adding a hundred concordant pairs changes neither the chi-squared nor the p value. The classic version compares (b − c)²/(b + c) with a chi-squared distribution on one degree of freedom; Edwards' continuity correction (1948) subtracts 1 from |b − c| before squaring, and the exact version computes the binomial probability of a split at least as extreme. With fewer than 25 discordant pairs the exact test is reported; from 25 upwards, the corrected chi-squared.
The p value says whether the change can be told apart from chance, not how large it is. The magnitude lives in the paired difference of proportions, δ = (b − c)/n, read as percentage points of the total number of pairs. Its Wald interval is the usual one and works well with many pairs; with few pairs, or when δ approaches ±1, the interval of Agresti and Min (2005) —equivalent to adding half an observation to each cell and clipping the result to [−1, 1]— has better coverage. The reported point estimate is always (b − c)/n; the selector changes only the interval.
The paired odds ratio b/c summarises the asymmetry of the change on a ratio scale: how many times more frequent changes in one direction were than in the other. Its interval comes from carrying the Clopper-Pearson interval of the proportion b/(b + c) to the odds scale. When b or c is 0 the ratio becomes 0 or infinite and stops reading as "so many times more"; when there are no discordant pairs at all, neither the test nor the odds ratio is defined.
Equations
- pairs positive on the first measurement and negative on the second, (+,−)
- pairs negative on the first and positive on the second, (−,+)
- discordant pairs; the concordant ones, a and d, do not enter
- number of changes in one direction under the null hypothesis
- difference of paired proportions
- total number of pairs, a + b + c + d
- adjusted estimator that centres the interval (half an observation per cell)
- Clopper-Pearson limits of the proportion b/(b + c)
R code
# McNemar test for paired proportions - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
# Rows = first measurement (test A, or "before"), columns = second one (test B, or "after").
# a = (+,+) and d = (-,-) are the concordant pairs; b = (+,-) and c = (-,+), the discordant ones.
a <- 40; b <- 15; c <- 5; d <- 90
metodo_delta <- "wald" # CI for the paired difference: "wald" or "agresti-min"
nivel <- 0.95
x <- matrix(c(a, b, c, d), nrow = 2, byrow = TRUE)
n <- a + b + c + d
n_disc <- b + c # only the discordant pairs carry information about change
z <- qnorm(1 - (1 - nivel) / 2)
p_a <- (a + b) / n # proportion positive with the first measurement
p_b <- (a + c) / n # proportion positive with the second one
# McNemar's chi-squared with 1 df, without and with Edwards' (1948) continuity
# correction. Note that mcnemar.test() skips the correction when b == c, and
# that both statistics (and their p values) are NaN when b + c = 0.
mc <- mcnemar.test(x, correct = FALSE)
mce <- mcnemar.test(x, correct = TRUE)
# Exact two-sided binomial test on the discordant pairs (exact McNemar)
p_exacta <- if (n_disc > 0) binom.test(b, n_disc, p = 0.5)$p.value else NA_real_
# Paired difference. The point estimate is always (b - c)/n; only the interval
# changes with `metodo_delta`.
est <- (b - c) / n
if (metodo_delta == "agresti-min") {
# Agresti & Min (2005): interval centred on the adjusted estimator (b - c)/(n + 2)
# and clipped to [-1, 1], as in PropCIs::diffpropci.mp (which reports (c - b)/n).
est_am <- (b - c) / (n + 2)
se_am <- sqrt((b + c + 1) - (b - c)^2 / (n + 2)) / (n + 2)
ll <- max(-1, est_am - z * se_am)
ul <- min(1, est_am + z * se_am)
delta <- c(est, ll, ul)
} else {
se <- sqrt((b + c) - (b - c)^2 / n) / n
delta <- c(est, est - z * se, est + z * se)
}
# Paired odds ratio b/c: the Clopper-Pearson interval for b/(b + c) carried to
# the odds scale, p/(1 - p). Undefined (NA) without discordant pairs.
or_pareado <- if (n_disc > 0) {
alpha <- 1 - nivel
pl <- if (b == 0) 0 else qbeta(alpha / 2, b, c + 1)
pu <- if (c == 0) 1 else qbeta(1 - alpha / 2, b + 1, c)
c(b / c, pl / (1 - pl), pu / (1 - pu))
} else c(NA, NA, NA)
res <- list(n = n, n_disc = n_disc, p_a = p_a, p_b = p_b,
chi2 = unname(mc$statistic), p_chi2 = mc$p.value,
chi2_edwards = unname(mce$statistic), p_edwards = mce$p.value,
p_exacta = p_exacta, delta = delta, or_pareado = or_pareado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run in the browser):
# exact2x2::mcnemar.exact(x)
# PropCIs::diffpropci.Wald.mp(b, c, n, nivel) # note: PropCIs reports (c - b)/n
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
Paired proportions were compared with McNemar's test [1] (the exact binomial version when there were fewer than 25 discordant pairs; otherwise the chi-squared with Edwards' continuity correction [2]); for these data we report the exact binomial test. The difference of paired proportions is reported with its 95% CI by the Wald method, and the paired odds ratio with the interval derived from the Clopper-Pearson interval [5] of b/(b + c). Calculations used the "McNemar's test (paired)" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/mcnemar), verified against R (stats).
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 McNemar Q. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika. 1947;12(2):153–157. doi:10.1007/BF02295996 PMID: 20254758 Original source
- 02 Edwards AL. Note on the “correction for continuity” in testing the significance of the difference between correlated proportions. Psychometrika. 1948;13(3):185–187. doi:10.1007/BF02289261 PMID: 18885738 Original source
- 03 Agresti A, Min Y. Simple improved confidence intervals for comparing matched proportions. Statistics in Medicine. 2005;24(5):729–740. doi:10.1002/sim.1781 PMID: 15696504 Original source
- 04 Newcombe RG. Improved confidence intervals for the difference between binomial proportions based on paired data. Statistics in Medicine. 1998;17(22):2635–2650. doi:10.1002/(SICI)1097-0258(19981130)17:22<2635::AID-SIM954>3.0.CO;2-C PMID: 9839354 Complementary
- 05 Clopper CJ, Pearson ES. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika. 1934;26(4):404–413. doi:10.1093/biomet/26.4.404 Complementary
- 06 Fleiss JL, Levin B, Paik MC. Statistical Methods for Rates and Proportions. 3rd ed. Hoboken, NJ: John Wiley & Sons; 2003. doi:10.1002/0471445428 Didactic reading
- 07 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading