← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

Test A (or before) Test B (or after) Total
Positive Negative
Positive 55
Negative 95
Total 45 105 150
  1. Positive on both (a)
  2. Positive on A, negative on B (b) Pairs in which the first measurement was positive and the second negative.
  3. Negative on A, positive on B (c) Pairs in which the first measurement was negative and the second positive.
  4. Negative on both (d)

Wald is the usual interval. Agresti-Min (2005) has better coverage with few pairs or when the difference approaches ±1.

Example loaded

Illustrative example: two rapid tests applied to the same 150 patients; 40 pairs were positive on both, 15 only on test A, 5 only on test B and 90 negative on both (fictitious data).

Illustrative data, not real.

Results

Pairs analysed (n)

150

Discordant pairs (b + c)

20

the only ones that enter the test

Positive with test A

36.7%

Positive with test B

30.0%

McNemar's chi-squared (uncorrected)

5.000

1 degree of freedom · no correction

p of the uncorrected chi-squared

0.025

1 degree of freedom · no correction

Chi-squared with Edwards' correction

4.050

1 degree of freedom · Edwards' correction

p of the chi-squared with Edwards' correction

0.044

1 degree of freedom · Edwards' correction

p of the exact (binomial) test

0.041

two-sided binomial on b + c

Paired difference (b − c)/n, percentage points

6.7

0.9 to 12.4

95% CI · Wald

Paired odds ratio (b/c)

3.00

1.04 to 10.55

95% CI · derived from Clopper-Pearson

Interpretation

Among the 150 pairs, 15 changed in one direction and 5 in the other: test A was positive in 36.7% and test B in 30.0% (difference 6.7 percentage points, 95% CI: 0.9 to 12.4).

Changes were more frequent in the A+ / B− direction (15 pairs against 5): test A was positive more often than test B.

Exact McNemar test (two-sided binomial on the 20 discordant pairs): p = 0.041.

p = 0.041; at α = 0.05 the null hypothesis that both measurements give the same proportion of positives is rejected. The size of the change (6.7 percentage points, 95% CI: 0.9 to 12.4) matters more than the p value.

Paired odds ratio b/c = 3.00 (95% CI: 1.04 to 10.55): changes in one direction (15 pairs) were 3.00 times more frequent than the opposite ones (5 pairs).

The three versions of the test on the same 20 discordant pairs: uncorrected chi-squared = 5.000 (p = 0.025); chi-squared with Edwards' correction = 4.050 (p = 0.044); exact binomial, p = 0.041. Reported here: the exact binomial test.

  • There are only 20 discordant pairs (fewer than 25): the exact binomial test is reported as the main one, because with so few the chi-squared approximation is unreliable.
  • The test uses only the discordant pairs (20): pairs where both measurements agreed carry no information about change.
Discordant pairs in each direction, with the line of the split expected under the null hypothesisA+ / B− (b): 15; A− / B+ (c): 5; Expected under H₀: (b + c)/2: 10.0Number of pairs051015A+ / B− (b)A− / B+ (c)Expected under H₀: (b + c)/2155
Discordant pairs in each direction, with the line of the split expected under the null hypothesis

Explanation

When the two proportions being compared come from the same subjects (two tests on the same patient, or the same measurement before and after), the observations come in pairs and are not independent. The 2×2 table no longer counts people per group: it counts pairs. Cells a and d hold the pairs where both measurements agreed (concordant) and cells b and c the pairs where they disagreed (discordant).

McNemar's test (1947) looks only at the discordant pairs: if the two measurements were equivalent, a change in one direction would be as likely as the opposite one, that is, b and c would be two halves of b + c. Concordant pairs carry no information about change and so do not enter the statistic; adding a hundred concordant pairs changes neither the chi-squared nor the p value. The classic version compares (b − c)²/(b + c) with a chi-squared distribution on one degree of freedom; Edwards' continuity correction (1948) subtracts 1 from |b − c| before squaring, and the exact version computes the binomial probability of a split at least as extreme. With fewer than 25 discordant pairs the exact test is reported; from 25 upwards, the corrected chi-squared.

The p value says whether the change can be told apart from chance, not how large it is. The magnitude lives in the paired difference of proportions, δ = (b − c)/n, read as percentage points of the total number of pairs. Its Wald interval is the usual one and works well with many pairs; with few pairs, or when δ approaches ±1, the interval of Agresti and Min (2005) —equivalent to adding half an observation to each cell and clipping the result to [−1, 1]— has better coverage. The reported point estimate is always (b − c)/n; the selector changes only the interval.

The paired odds ratio b/c summarises the asymmetry of the change on a ratio scale: how many times more frequent changes in one direction were than in the other. Its interval comes from carrying the Clopper-Pearson interval of the proportion b/(b + c) to the odds scale. When b or c is 0 the ratio becomes 0 or infinite and stops reading as "so many times more"; when there are no discordant pairs at all, neither the test nor the odds ratio is defined.

Equations

χ2=(b−c)2b+c ∼ χ12\chi^2=\frac{(b-c)^2}{b+c}\ \sim\ \chi^2_{1}
bb
pairs positive on the first measurement and negative on the second, (+,−)
cc
pairs negative on the first and positive on the second, (−,+)
b+cb+c
discordant pairs; the concordant ones, a and d, do not enter
McNemar's (1947) statistic without correction, on one degree of freedom.
χc2=(∣b−c∣−1)2b+c\chi^2_{c}=\frac{(\lvert b-c\rvert-1)^2}{b+c}
Edwards' (1948) continuity correction. As in `mcnemar.test(correct = TRUE)`, it is not applied when b = c (the statistic stays at 0); with b + c = 0 it is undefined.
pexact=min⁡ ⁣(1, 2∑k=0min⁡(b, c)(b+ck)(12)b+c)p_{\mathrm{exact}}=\min\!\left(1,\ 2\sum_{k=0}^{\min(b,\,c)}\binom{b+c}{k}\left(\frac{1}{2}\right)^{b+c}\right)
kk
number of changes in one direction under the null hypothesis
Two-sided binomial test on the discordant pairs, with p = 1/2 under the null hypothesis.
δ^=b−cn,SE⁡(δ^)=(b+c)−(b−c)2/nn\hat\delta=\frac{b-c}{n},\qquad \se(\hat\delta)=\frac{\sqrt{(b+c)-(b-c)^2/n}}{n}
δ^\hat\delta
difference of paired proportions
nn
total number of pairs, a + b + c + d
Paired difference with its Wald standard error; the interval is δ ± z·SE(δ).
δ~=b−cn+2,SE⁡(δ~)=(b+c+1)−(b−c)2/(n+2)n+2\tilde\delta=\frac{b-c}{n+2},\qquad \se(\tilde\delta)=\frac{\sqrt{(b+c+1)-(b-c)^2/(n+2)}}{n+2}
δ~\tilde\delta
adjusted estimator that centres the interval (half an observation per cell)
Agresti and Min (2005) interval: δ̃ ± z·SE(δ̃), clipped to [−1, 1]. The reported value is still (b − c)/n.
ORpaired=bc,CI=[πL1−πL, πU1−πU]\OR_{\mathrm{paired}}=\frac{b}{c},\qquad \CI=\left[\frac{\pi_{L}}{1-\pi_{L}},\ \frac{\pi_{U}}{1-\pi_{U}}\right]
πL, πU\pi_{L},\ \pi_{U}
Clopper-Pearson limits of the proportion b/(b + c)
Paired odds ratio and its interval carried from the proportion scale to the odds scale.

R code

# McNemar test for paired proportions - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)

# Rows = first measurement (test A, or "before"), columns = second one (test B, or "after").
# a = (+,+) and d = (-,-) are the concordant pairs; b = (+,-) and c = (-,+), the discordant ones.
a <- 40; b <- 15; c <- 5; d <- 90
metodo_delta <- "wald"   # CI for the paired difference: "wald" or "agresti-min"
nivel <- 0.95
x <- matrix(c(a, b, c, d), nrow = 2, byrow = TRUE)
n <- a + b + c + d
n_disc <- b + c        # only the discordant pairs carry information about change
z <- qnorm(1 - (1 - nivel) / 2)

p_a <- (a + b) / n     # proportion positive with the first measurement
p_b <- (a + c) / n     # proportion positive with the second one

# McNemar's chi-squared with 1 df, without and with Edwards' (1948) continuity
# correction. Note that mcnemar.test() skips the correction when b == c, and
# that both statistics (and their p values) are NaN when b + c = 0.
mc  <- mcnemar.test(x, correct = FALSE)
mce <- mcnemar.test(x, correct = TRUE)

# Exact two-sided binomial test on the discordant pairs (exact McNemar)
p_exacta <- if (n_disc > 0) binom.test(b, n_disc, p = 0.5)$p.value else NA_real_

# Paired difference. The point estimate is always (b - c)/n; only the interval
# changes with `metodo_delta`.
est <- (b - c) / n
if (metodo_delta == "agresti-min") {
  # Agresti & Min (2005): interval centred on the adjusted estimator (b - c)/(n + 2)
  # and clipped to [-1, 1], as in PropCIs::diffpropci.mp (which reports (c - b)/n).
  est_am <- (b - c) / (n + 2)
  se_am <- sqrt((b + c + 1) - (b - c)^2 / (n + 2)) / (n + 2)
  ll <- max(-1, est_am - z * se_am)
  ul <- min(1, est_am + z * se_am)
  delta <- c(est, ll, ul)
} else {
  se <- sqrt((b + c) - (b - c)^2 / n) / n
  delta <- c(est, est - z * se, est + z * se)
}

# Paired odds ratio b/c: the Clopper-Pearson interval for b/(b + c) carried to
# the odds scale, p/(1 - p). Undefined (NA) without discordant pairs.
or_pareado <- if (n_disc > 0) {
  alpha <- 1 - nivel
  pl <- if (b == 0) 0 else qbeta(alpha / 2, b, c + 1)
  pu <- if (c == 0) 1 else qbeta(1 - alpha / 2, b + 1, c)
  c(b / c, pl / (1 - pl), pu / (1 - pu))
} else c(NA, NA, NA)

res <- list(n = n, n_disc = n_disc, p_a = p_a, p_b = p_b,
            chi2 = unname(mc$statistic), p_chi2 = mc$p.value,
            chi2_edwards = unname(mce$statistic), p_edwards = mce$p.value,
            p_exacta = p_exacta, delta = delta, or_pareado = or_pareado)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio (not run in the browser):
# exact2x2::mcnemar.exact(x)
# PropCIs::diffpropci.Wald.mp(b, c, n, nivel)   # note: PropCIs reports (c - b)/n

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

Paired proportions were compared with McNemar's test [1] (the exact binomial version when there were fewer than 25 discordant pairs; otherwise the chi-squared with Edwards' continuity correction [2]); for these data we report the exact binomial test. The difference of paired proportions is reported with its 95% CI by the Wald method, and the paired odds ratio with the interval derived from the Clopper-Pearson interval [5] of b/(b + c). Calculations used the "McNemar's test (paired)" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/mcnemar), verified against R (stats).

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 McNemar Q. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika. 1947;12(2):153–157. doi:10.1007/BF02295996 PMID: 20254758 Original source
  2. 02 Edwards AL. Note on the “correction for continuity” in testing the significance of the difference between correlated proportions. Psychometrika. 1948;13(3):185–187. doi:10.1007/BF02289261 PMID: 18885738 Original source
  3. 03 Agresti A, Min Y. Simple improved confidence intervals for comparing matched proportions. Statistics in Medicine. 2005;24(5):729–740. doi:10.1002/sim.1781 PMID: 15696504 Original source
  4. 04 Newcombe RG. Improved confidence intervals for the difference between binomial proportions based on paired data. Statistics in Medicine. 1998;17(22):2635–2650. doi:10.1002/(SICI)1097-0258(19981130)17:22<2635::AID-SIM954>3.0.CO;2-C PMID: 9839354 Complementary
  5. 05 Clopper CJ, Pearson ES. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika. 1934;26(4):404–413. doi:10.1093/biomet/26.4.404 Complementary
  6. 06 Fleiss JL, Levin B, Paik MC. Statistical Methods for Rates and Proportions. 3rd ed. Hoboken, NJ: John Wiley & Sons; 2003. doi:10.1002/0471445428 Didactic reading
  7. 07 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading