Tools · Open Biostatistics
Cohen's kappa, unweighted and weighted: agreement between two raters with confidence interval, PABAK and Byrt's indices
Paste the agreement table of two raters, from 2 × 2 up to 10 × 10, and get Cohen's kappa with its confidence interval, the Landis and Koch band, the largest kappa the marginals allow and, when there are two categories, the PABAK and the prevalence and bias indices that explain why high agreement can still give a low kappa.
https://udgca1190.com.mx/en/herramientas/bioestadistica/kappa
This link does not include the pasted data: they are too long for a URL.
Results
Cases classified (n)
100
Observed agreement (p₀)
76.0%
Unweighted
Agreement expected by chance (pₑ)
38.7%
Unweighted
Cohen's kappa (κ)
0.61
0.47 to 0.74
95% CI · Fleiss, Cohen and Everitt
Standard error of κ
0.0690
Fleiss, Cohen and Everitt
z against κ = 0
8.14
variance under κ = 0
p against κ = 0
< 0.001
variance under κ = 0
Largest κ these marginals allow
0.95
limit imposed by the marginals
PABAK
—
two categories only
Prevalence index
—
two categories only
Bias index
—
two categories only
Interpretation
The raters agreed on 76.0% of the 100 cases; the agreement expected by chance was 38.7%. κ = 0.61 (95% CI: 0.47 to 0.74): substantial agreement beyond chance.
On the Landis and Koch bands, κ = 0.61 is substantial agreement (0.61 to 0.80). The interval 0.47 to 0.74 shows how far the uncertainty reaches.
Unweighted κ: every disagreement counts the same, whether between neighbouring categories or between the extremes. That is the right choice when the 3 categories have no order.
z = 8.14, p = < 0.001 against κ = 0. The test says whether agreement exceeds chance, not whether agreement is good enough: with large samples a small kappa also comes out significant.
With these marginals the largest possible κ is 0.95: no arrangement of the same row and column totals would give more. The further it falls below 1, the more the result is limited by the two raters using the categories with different frequencies.
PABAK and the prevalence and bias indices are defined only with two categories; this table has 3.
Observed agreement (76.0%) and κ (0.61) point the same way: there is no sign of the kappa paradox, which appears when one category holds almost every case and chance agreement shoots up.
- The Landis and Koch bands (poor, slight, fair, moderate, substantial, almost perfect) are conventional cut-offs, not derived from theory, and they mark no threshold of validity. Interpret κ together with its confidence interval and with the observed agreement.
Explanation
Two people classify the same cases and agree on 90% of them. Is that good agreement? It depends on how much they would have agreed without looking: if 95% of the cases belong to a single category, two raters answering at random with those same frequencies would already agree almost always. Cohen's kappa answers exactly that: it compares observed agreement with the agreement expected by chance and expresses how much of the remaining room was actually covered. It is 1 for perfect agreement, 0 when agreement is what chance would produce, and below 0 when it is worse than chance.
When the categories are ordered (absent, mild, moderate, severe) not every disagreement is the same: confusing "mild" with "moderate" is not the same as confusing "mild" with "severe". Weighted kappa captures that idea by giving partial credit to the cells near the diagonal. With linear weights the credit falls in proportion to the distance between categories; with quadratic weights large disagreements count far more than small ones, and with equally spaced categories that version coincides with the intraclass correlation coefficient. Weighting categories that have no order (germ types, say) means nothing: there the right kappa is the unweighted one.
Kappa has a quirk worth knowing before you report it, the so-called kappa paradox. When one category holds almost every case, the agreement expected by chance rises so much that kappa collapses even though the raters agree on nearly everything. Two indices from Byrt, Bishop and Carlin help diagnose it with two categories: the prevalence index measures how unequal the two agreement cells are, and the bias index, how much the two raters differ in how often they use each category. PABAK is the kappa you would get if both indices were zero, that is, the observed agreement rescaled, and it is reported alongside kappa, never instead of it.
The confidence interval uses the asymptotic standard error of Fleiss, Cohen and Everitt (1969) and is truncated to [−1, 1]. The test against κ = 0 uses a different variance, the one that corresponds to the null hypothesis of chance agreement, and it is the one irr::kappa2 reports in R; it answers only whether agreement exceeds chance, not whether agreement is good enough, so with large samples a small kappa also comes out significant. The Landis and Koch bands (poor, slight, fair, moderate, substantial and almost perfect) are conventions quoted since 1977, not thresholds derived from theory: always read them with the interval alongside.
Equations
- number of categories in the table (from 2 to 10)
- agreement weight of the cell in row i, column j
- linear weights
- quadratic weights
- proportion of cases in the cell in row i, column j
- marginal of row i (rater A)
- marginal of column j (rater B)
- observed (weighted) agreement
- agreement expected by chance (weighted)
- weighted mean of row i, using the column marginals as weights
- weighted mean of column j, using the row marginals as weights
- number of cases classified by both raters
- standard normal quantile (1.96 at 95%)
- standard error under the null hypothesis of chance agreement
- highest kappa compatible with the observed marginals
- maximum agreement attainable with those marginals
- north-west corner allocation: the table with those same marginals that maximises weighted agreement
- agreement cells of a 2 × 2 table
- disagreement cells of a 2 × 2 table
- prevalence index
- bias index
R code
# Cohen's kappa (unweighted and weighted) for a k x k agreement table
# - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(irr)
library(jsonlite)
# The pasted table, row by row: rows = rater A, columns = rater B,
# the same categories in the same order.
x <- matrix(c(40, 8, 2, 6, 25, 4, 1, 3, 11), nrow = 3, byrow = TRUE)
ponderacion <- "ninguna" # "ninguna" | "lineal" | "cuadratica"
nivel <- 0.95
tipo <- switch(ponderacion, ninguna = "unweighted", lineal = "equal", cuadratica = "squared")
# A category nobody used (empty row AND empty column) is dropped first: it adds no
# agreement, and the weights depend on how many categories there are.
vacias <- rowSums(x) == 0 & colSums(x) == 0
x <- x[!vacias, !vacias, drop = FALSE]
k <- nrow(x); n <- sum(x)
# Agreement weights w_ij (Cohen 1968); with k = 2 the three schemes are the identity.
w <- switch(tipo, unweighted = diag(k), equal = 1 - abs(outer(1:k, 1:k, "-"))/(k - 1),
squared = 1 - (outer(1:k, 1:k, "-")/(k - 1))^2)
p <- x/n; pi <- rowSums(p); pj <- colSums(p)
po <- sum(w*p); pe <- sum(w*outer(pi, pj)); kappa <- (po - pe)/(1 - pe)
# Asymptotic standard error of Fleiss, Cohen & Everitt (1969), from the OBSERVED cells,
# and its Wald interval truncated to [-1, 1]. max(0, .) only absorbs the rounding of
# perfect agreement, where the radicand is 0 up to one ulp.
wi <- as.vector(w %*% pj); wj <- as.vector(t(w) %*% pi) # weighted row and column means
ee <- sqrt(max(0, sum(p*(w - outer(wi, wj, "+")*(1 - kappa))^2) - (kappa - pe*(1 - kappa))^2))/((1 - pe)*sqrt(n))
z <- qnorm(1 - (1 - nivel)/2)
kappa_ic <- c(kappa, max(-1, kappa - z*ee), min(1, kappa + z*ee))
# z and p against kappa = 0 exactly as irr::kappa2, whose variance is the one under the
# null hypothesis of chance agreement (EXPECTED cells), not the one of the interval.
# kappa2 needs the ratings pair by pair, so they are rebuilt from the table; the
# categories are numbered from 11 because kappa2 sorts its levels as text, and with
# 1 to 10 it would read them "1", "10", "2" and shift the weights of a 10 x 10 table.
niveles <- 10 + 1:k
ratings <- cbind(rep(niveles, rowSums(x)), unlist(lapply(1:k, function(i) rep(niveles, x[i, ]))))
k2 <- kappa2(ratings, weight = tipo)
# Largest kappa these marginals allow. Unweighted that maximum is sum(min(pi, pj));
# with linear or quadratic weights the north-west corner rule reaches it, because
# |i - j| and (i - j)^2 satisfy Monge's condition.
esquina <- function(a, b) {
m <- matrix(0, length(a), length(b)); i <- 1; j <- 1
while (i <= length(a) && j <= length(b)) {
v <- min(a[i], b[j]); m[i, j] <- m[i, j] + v; a[i] <- a[i] - v; b[j] <- b[j] - v
if (a[i] <= 0) i <- i + 1 else j <- j + 1
}
m
}
po_max <- if (tipo == "unweighted") sum(pmin(pi, pj)) else sum(w*esquina(pi, pj))
kappa_max <- (po_max - pe)/(1 - pe)
# Only with two categories (Byrt, Bishop & Carlin 1993): PABAK removes the influence
# of prevalence and bias, and the two indices measure them.
pabak <- if (k == 2) 2*po - 1 else NA_real_
indice_prevalencia <- if (k == 2) abs(x[1, 1] - x[2, 2])/n else NA_real_
indice_sesgo <- if (k == 2) abs(x[1, 2] - x[2, 1])/n else NA_real_
res <- list(n = n, po = po, pe = pe, kappa = kappa_ic, ee = ee,
z_h0 = unname(k2$statistic), p_h0 = k2$p.value, kappa_max = kappa_max,
pabak = pabak, indice_prevalencia = indice_prevalencia,
indice_sesgo = indice_sesgo)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))
# Equivalent in RStudio (not run in the browser):
# DescTools::CohenKappa(x, weights = "Unweighted", conf.level = nivel) # "Equal-Spacing", "Fleiss-Cohen"
This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.
In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.
Methods for a manuscript
Interobserver agreement was assessed with Cohen's kappa coefficient [1] (unweighted) with a 95% CI based on the asymptotic standard error of Fleiss, Cohen and Everitt [3]; its magnitude was interpreted following Landis and Koch [4]; the largest κ compatible with the observed marginals is also reported. 100 cases were classified into 3 categories and κ = 0.61 (95% CI 0.47 to 0.74) was obtained, with an observed agreement of 76.0%. Calculations used the "Cohen's kappa" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/kappa), verified against R (irr::kappa2).
A paragraph ready for the Methods section; bracketed numbers refer to the reference list.
References
- 01 Cohen J. A coefficient of agreement for nominal scales. Educational and Psychological Measurement. 1960;20(1):37–46. doi:10.1177/001316446002000104 Original source
- 02 Cohen J. Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin. 1968;70(4):213–220. doi:10.1037/h0026256 PMID: 19673146 Original source
- 03 Fleiss JL, Cohen J, Everitt BS. Large sample standard errors of kappa and weighted kappa. Psychological Bulletin. 1969;72(5):323–327. doi:10.1037/h0028106 Original source
- 04 Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–174. doi:10.2307/2529310 PMID: 843571 Original source
- 05 Byrt T, Bishop J, Carlin JB. Bias, prevalence and kappa. Journal of Clinical Epidemiology. 1993;46(5):423–429. doi:10.1016/0895-4356(93)90018-V PMID: 8501467 Original source
- 06 Fleiss JL, Cohen J. The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability. Educational and Psychological Measurement. 1973;33(3):613–619. doi:10.1177/001316447303300309 Complementary
- 07 Gamer M, Lemon J, Fellows I, Singh P. irr: Various Coefficients of Interrater Reliability and Agreement. R package version 0.85. CRAN; 2026. Complementary
- 08 Sim J, Wright CC. The kappa statistic in reliability studies: use, interpretation, and sample size requirements. Physical Therapy. 2005;85(3):257–268. doi:10.1093/ptj/85.3.257 PMID: 15733050 Didactic reading
- 09 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading