Tools · Open Biostatistics
2×2 tests of independence: Pearson's chi-squared, Yates' correction, the N−1 variant, Fisher's exact test and the phi coefficient
Enter the four counts of the 2×2 table and get the expected frequencies, Pearson's chi-squared with and without Yates' correction, the N−1 variant, Fisher's two-sided exact test, the signed phi coefficient and the conditional odds ratio with its interval, plus the test that Cochran's rule recommends for these counts.
https://udgca1190.com.mx/en/herramientas/bioestadistica/chi-cuadrada-fisher
This link does not include the pasted data: they are too long for a URL.
Results
Total participants (n)
200
Expected in a
21.00
under independence
Expected in b
79.00
under independence
Expected in c
21.00
under independence
Expected in d
79.00
under independence
Smallest expected frequency
21.00
Cochran's rule: at least 5
Pearson's χ²
9.765
1 df, no correction
p (Pearson's χ²)
0.002
1 df, no correction
χ² with Yates' correction
8.710
1 df, correction bounded as in R
p (Yates' correction)
0.003
1 df, correction bounded as in R
χ² "N−1" (Campbell)
9.716
1 df, χ² × (n − 1)/n
p (N−1 variant)
0.002
1 df, χ² × (n − 1)/n
p (Fisher's exact test, two-sided)
0.003
two-sided "minlike", as fisher.test
Conditional odds ratio (Fisher)
0.32
0.14 to 0.70
95% CI · conditional maximum likelihood
Phi coefficient
-0.221
signed; |φ| = √(χ²/n)
Interpretation
The proportion with the outcome was 12.0% in group 1 versus 30.0% in group 2 (n = 200).
Recommended test: Pearson's chi-squared without correction, p = 0.002. Every expected frequency reaches 5 and the total is at least 20, so the chi-squared approximation is adequate (Cochran's rule).
p = 0.002; at α = 0.05 the null hypothesis of independence between group and outcome is rejected. The size of the association and its confidence interval matter more than the p value.
φ = -0.221: a small association, negative (the outcome is less frequent in group 1), by Cohen's convention (|φ| between 0.1 and 0.3). These are conventional cut-offs, not results derived from theory.
Conditional odds ratio: 0.32 (95% CI: 0.14 to 0.70). It is the maximum likelihood estimate that accompanies Fisher's exact test, the one coherent with its p value. For the usual odds ratio (ad/bc) with Woolf's interval, use the "Association and effect (2×2 table)" calculator.
The three versions of the chi-squared test give: Pearson without correction χ² = 9.765 (p = 0.002); with Yates' correction χ² = 8.710 (p = 0.003); the N−1 variant χ² = 9.716 (p = 0.002). Yates' correction is conservative: it lowers the statistic and raises the p value more than necessary, so it is not recommended as a matter of routine. With small expected frequencies Fisher's exact test is preferable and, if an approximation is wanted, Campbell's N−1 variant (2007) matches the nominal size better.
Explanation
A 2×2 table crosses two dichotomous variables: here, the group or exposure (rows) and the outcome (columns). These tests ask whether the two observed proportions are compatible with independence between the variables, that is, with the outcome occurring equally often in both groups. Under that null hypothesis each cell would have the expected frequency E = (its row total) × (its column total) / n, and Pearson's chi-squared (1900) measures how far the observed frequencies fall from those expected ones.
The chi-squared distribution is an approximation and stops being reliable with small counts. Cochran's rule (1954) asks that every expected frequency reach 5 and that the total be at least 20; below that, Fisher's exact test is preferred, because it approximates nothing: it enumerates every table with the same marginal totals, computes the hypergeometric probability of each and adds up those no more likely than the observed one. (With n below 20 the smallest expected frequency can never reach 5, because it cannot exceed n/4: the two conditions are stated separately, but requiring every expected frequency to reach 5 already forces n to be at least 20.)
Yates' continuity correction (1934) was meant to bring the approximation closer to the exact result, but it is conservative: it produces p values that are too large. It is shown here because many programs apply it by default, not because it is recommended. Campbell's N−1 variant (2007) multiplies chi-squared by (n−1)/n and matches the nominal size better than either of the others when the smallest expected frequency lies between 1 and 5. This calculator computes Yates' correction exactly as R does, bounding it to min(0.5, |O − E|): if |ad − bc| is smaller than n/2, the corrected statistic stays at 0 instead of turning negative.
A p value does not measure the strength of the association; the phi coefficient (Yule 1912) does. It is the correlation between two dichotomous variables, runs from −1 to +1 and its sign gives the direction. In a 2×2 table, |φ| = √(χ²/n), and it is read with Cohen's convention (1988): 0.1 small, 0.3 medium, 0.5 large, always with the caveat that these are conventional cut-offs. The odds ratio that accompanies Fisher's test is the conditional one: the value that maximises the likelihood of the non-central hypergeometric distribution with the marginal totals fixed, with the interval formed by the parameter values leaving α/2 of probability in each tail. It is not the ad/bc odds ratio of the "Association and effect (2×2 table)" calculator, which is the unconditional estimate with Woolf's logarithmic interval; the two are close with large counts and differ with small ones.
Equations
- observed frequency of the cell
- expected frequency of the cell under independence
- row 1 total (a + b) and row 2 total (c + d)
- column 1 total (a + c) and column 2 total (b + d)
- total, a + b + c + d
- bounded continuity correction, exactly as R applies it
- possible value of cell a with the same marginal totals
R code
# 2x2 tests of independence: chi-squared (Pearson, Yates, N-1), Fisher's exact test and phi
# Bioestadistica abierta, UDG-CA-1190. Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(jsonlite)
a <- 12; b <- 88; c <- 30; d <- 70 # rows = group or exposure (1, 2), columns = outcome (yes, no)
nivel <- 0.95 # confidence level of the conditional odds ratio interval
m <- matrix(c(a, b, c, d), nrow = 2, byrow = TRUE)
n <- sum(m)
esperados <- outer(rowSums(m), colSums(m)) / n # E_ij = (row i total) * (column j total) / n
chi <- suppressWarnings(chisq.test(m, correct = FALSE)) # Pearson's chi-squared, 1 df
yates <- suppressWarnings(chisq.test(m, correct = TRUE)) # Yates' correction, bounded to min(0.5, |O - E|) as R does
# "N - 1" chi-squared (Campbell 2007): Pearson's statistic times (n - 1)/n
chi2_n1 <- unname(chi$statistic) * (n - 1) / n
p_n1 <- pchisq(chi2_n1, 1, lower.tail = FALSE)
# Two-sided "minlike" p (sum of the tables no more likely than the observed one) and the
# conditional maximum likelihood odds ratio with its interval (solved by uniroot, tol ~ 1.2e-4)
fisher <- fisher.test(m, conf.level = nivel)
phi <- (a * d - b * c) / sqrt(prod(rowSums(m)) * prod(colSums(m))) # signed phi coefficient
res <- list(n = n,
e_a = esperados[1, 1], e_b = esperados[1, 2], e_c = esperados[2, 1], e_d = esperados[2, 2],
e_min = min(esperados),
chi2 = unname(chi$statistic), p_chi2 = chi$p.value,
chi2_yates = unname(yates$statistic), p_yates = yates$p.value,
chi2_n1 = chi2_n1, p_n1 = p_n1,
p_fisher = fisher$p.value,
or_cond = c(unname(fisher$estimate), fisher$conf.int[1], fisher$conf.int[2]),
phi = phi)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run in the browser):
# exact2x2::exact2x2(m, tsmethod = "minlike") # same p as fisher.test
# DescTools::Phi(m)
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
The association between the two variables of the 2×2 table (Group or exposure and Outcome) was assessed with Pearson's chi-squared test [1] (Yates' continuity correction [2] and the N−1 variant [6] are also reported) or with Fisher's two-sided exact test [3,4] when any expected frequency was below 5 or the total was below 20 (Cochran's rule [5]); for these data Pearson's chi-squared test without correction is recommended. The strength of the association was summarised with the phi coefficient [7], read with Cohen's convention [8], and with the conditional odds ratio of Fisher's test and its 95% confidence interval [4]. Calculations used the "2×2 independence (chi-squared and Fisher)" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/chi-cuadrada-fisher), verified against R (chisq.test, fisher.test).
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Pearson K. X. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, Series 5. 1900;50(302):157–175. doi:10.1080/14786440009463897 Original source
- 02 Yates F. Contingency tables involving small numbers and the χ² test. Supplement to the Journal of the Royal Statistical Society. 1934;1(2):217–235. doi:10.2307/2983604 Original source
- 03 Fisher RA. Statistical Methods for Research Workers. 5th ed. Edinburgh: Oliver & Boyd; 1934. Original source
- 04 Fisher RA. The logic of inductive inference. Journal of the Royal Statistical Society. 1935;98(1):39–82. doi:10.2307/2342435 Original source
- 05 Cochran WG. Some methods for strengthening the common χ² tests. Biometrics. 1954;10(4):417–451. doi:10.2307/3001616 Original source
- 06 Campbell I. Chi-squared and Fisher-Irwin tests of two-by-two tables with small sample recommendations. Statistics in Medicine. 2007;26(19):3661–3675. doi:10.1002/sim.2832 PMID: 17315184 Complementary
- 07 Yule GU. On the methods of measuring association between two attributes. Journal of the Royal Statistical Society. 1912;75(6):579–652. doi:10.2307/2340126 Original source
- 08 Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates; 1988. Didactic reading
- 09 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading