← Open Biostatistics index
Instant calculation in the browser · verifiable with R Active

Inputs

3 × 3 table read

Copy the table of counts in your spreadsheet and paste it here: one table row per line, cells separated by tabs, commas, semicolons or spaces, so a table copied from a PDF works too. Rows are the categories assigned by rater A and columns those of rater B, in the same order. If you paste the category names as a first row or a first column, they are dropped for you. A note on thousands: in a space-separated table a space always separates cells, so "1 234" would read as two cells; write thousands with no separator (1234), or separate the cells with tabs, commas or semicolons. The table must be square, from 2 × 2 up to 10 × 10.

Unweighted, every disagreement counts the same: that is the right choice for nominal categories. Linear and quadratic weights give partial credit to neighbouring categories and only make sense if the categories are ordered.

Example loaded

Illustrative example: two clinicians independently classify 100 dengue cases as "no warning signs", "with warning signs" and "severe" (fictitious data, not real).

Illustrative data, not real.

Results

Cases classified (n)

100

Observed agreement (p₀)

76.0%

Unweighted

Agreement expected by chance (pₑ)

38.7%

Unweighted

Cohen's kappa (κ)

0.61

0.47 to 0.74

95% CI · Fleiss, Cohen and Everitt

Standard error of κ

0.0690

Fleiss, Cohen and Everitt

z against κ = 0

8.14

variance under κ = 0

p against κ = 0

< 0.001

variance under κ = 0

Largest κ these marginals allow

0.95

limit imposed by the marginals

PABAK

—

two categories only

Prevalence index

—

two categories only

Bias index

—

two categories only

Interpretation

The raters agreed on 76.0% of the 100 cases; the agreement expected by chance was 38.7%. κ = 0.61 (95% CI: 0.47 to 0.74): substantial agreement beyond chance.

On the Landis and Koch bands, κ = 0.61 is substantial agreement (0.61 to 0.80). The interval 0.47 to 0.74 shows how far the uncertainty reaches.

Unweighted κ: every disagreement counts the same, whether between neighbouring categories or between the extremes. That is the right choice when the 3 categories have no order.

z = 8.14, p = < 0.001 against κ = 0. The test says whether agreement exceeds chance, not whether agreement is good enough: with large samples a small kappa also comes out significant.

With these marginals the largest possible κ is 0.95: no arrangement of the same row and column totals would give more. The further it falls below 1, the more the result is limited by the two raters using the categories with different frequencies.

PABAK and the prevalence and bias indices are defined only with two categories; this table has 3.

Observed agreement (76.0%) and κ (0.61) point the same way: there is no sign of the kappa paradox, which appears when one category holds almost every case and chance agreement shoots up.

  • The Landis and Koch bands (poor, slight, fair, moderate, substantial, almost perfect) are conventional cut-offs, not derived from theory, and they mark no threshold of validity. Interpret κ together with its confidence interval and with the observed agreement.
Cohen's kappa with its confidence interval, against chance agreementCohen's kappa (κ): 0.61 (0.47 to 0.74); Observed agreement (p₀): 76.0%; Agreement expected by chance (pₑ): 38.7%Cohen's kappa (κ)0.000.250.500.751.00
Cohen's kappa with its confidence interval, against chance agreement

Explanation

Two people classify the same cases and agree on 90% of them. Is that good agreement? It depends on how much they would have agreed without looking: if 95% of the cases belong to a single category, two raters answering at random with those same frequencies would already agree almost always. Cohen's kappa answers exactly that: it compares observed agreement with the agreement expected by chance and expresses how much of the remaining room was actually covered. It is 1 for perfect agreement, 0 when agreement is what chance would produce, and below 0 when it is worse than chance.

When the categories are ordered (absent, mild, moderate, severe) not every disagreement is the same: confusing "mild" with "moderate" is not the same as confusing "mild" with "severe". Weighted kappa captures that idea by giving partial credit to the cells near the diagonal. With linear weights the credit falls in proportion to the distance between categories; with quadratic weights large disagreements count far more than small ones, and with equally spaced categories that version coincides with the intraclass correlation coefficient. Weighting categories that have no order (germ types, say) means nothing: there the right kappa is the unweighted one.

Kappa has a quirk worth knowing before you report it, the so-called kappa paradox. When one category holds almost every case, the agreement expected by chance rises so much that kappa collapses even though the raters agree on nearly everything. Two indices from Byrt, Bishop and Carlin help diagnose it with two categories: the prevalence index measures how unequal the two agreement cells are, and the bias index, how much the two raters differ in how often they use each category. PABAK is the kappa you would get if both indices were zero, that is, the observed agreement rescaled, and it is reported alongside kappa, never instead of it.

The confidence interval uses the asymptotic standard error of Fleiss, Cohen and Everitt (1969) and is truncated to [−1, 1]. The test against κ = 0 uses a different variance, the one that corresponds to the null hypothesis of chance agreement, and it is the one irr::kappa2 reports in R; it answers only whether agreement exceeds chance, not whether agreement is good enough, so with large samples a small kappa also comes out significant. The Landis and Koch bands (poor, slight, fair, moderate, substantial and almost perfect) are conventions quoted since 1977, not thresholds derived from theory: always read them with the interval alongside.

Equations

wij=1 if i=j,  0 if i≠j;wijL=1−∣i−j∣k−1;wijQ=1−(i−jk−1)2w_{ij}=1\ \mathrm{if}\ i=j,\ \ 0\ \mathrm{if}\ i\neq j;\qquad w^{\mathrm{L}}_{ij}=1-\frac{|i-j|}{k-1};\qquad w^{\mathrm{Q}}_{ij}=1-\left(\frac{i-j}{k-1}\right)^{2}
kk
number of categories in the table (from 2 to 10)
wijw_{ij}
agreement weight of the cell in row i, column j
wijLw^{\mathrm{L}}_{ij}
linear weights
wijQw^{\mathrm{Q}}_{ij}
quadratic weights
Cohen's agreement weights (1968). With two categories the three schemes coincide with the identity, so the weighting changes nothing.
po=∑i,jwij pij,pe=∑i,jwij pi⋅ p⋅j,κ^w=po−pe1−pep_o=\sum_{i,j}w_{ij}\,p_{ij},\qquad p_e=\sum_{i,j}w_{ij}\,p_{i\cdot}\,p_{\cdot j},\qquad \hat\kappa_w=\frac{p_o-p_e}{1-p_e}
pijp_{ij}
proportion of cases in the cell in row i, column j
pi⋅p_{i\cdot}
marginal of row i (rater A)
p⋅jp_{\cdot j}
marginal of column j (rater B)
pop_o
observed (weighted) agreement
pep_e
agreement expected by chance (weighted)
Kappa is the fraction of the room left above chance that was actually covered.
SE⁡(κ^w)=∑i,jpij[wij−(wˉi⋅+wˉ⋅j)(1−κ^w)]2−[κ^w−pe(1−κ^w)]2(1−pe)n\se(\hat\kappa_w)=\frac{\sqrt{\sum_{i,j}p_{ij}\left[w_{ij}-(\bar w_{i\cdot}+\bar w_{\cdot j})(1-\hat\kappa_w)\right]^{2}-\left[\hat\kappa_w-p_e(1-\hat\kappa_w)\right]^{2}}}{(1-p_e)\sqrt{n}}
wˉi⋅\bar w_{i\cdot}
weighted mean of row i, using the column marginals as weights
wˉ⋅j\bar w_{\cdot j}
weighted mean of column j, using the row marginals as weights
nn
number of cases classified by both raters
Asymptotic standard error of Fleiss, Cohen and Everitt (1969), computed from the observed cells. It is 0 under perfect agreement.
CI1−α(κ^w)=κ^w±z1−α/2⋅SE⁡(κ^w),zH0=κ^wSE⁡0(κ^w)\CI_{1-\alpha}(\hat\kappa_w)=\hat\kappa_w\pm z_{1-\alpha/2}\cdot\se(\hat\kappa_w),\qquad z_{H_0}=\frac{\hat\kappa_w}{\se_{0}(\hat\kappa_w)}
z1−α/2z_{1-\alpha/2}
standard normal quantile (1.96 at 95%)
SE⁡0\se_{0}
standard error under the null hypothesis of chance agreement
The interval is truncated to [−1, 1]. The test against κ = 0 uses a different variance, the null one, which is what irr::kappa2 reports.
κmax⁡=po,max⁡−pe1−pe;po,max⁡=∑imin⁡(pi⋅, p⋅i)  (unweighted),po,max⁡=∑i,jwij mij  (weighted)\kappa_{\max}=\frac{p_{o,\max}-p_e}{1-p_e};\qquad p_{o,\max}=\sum_{i}\min\left(p_{i\cdot},\,p_{\cdot i}\right)\;\text{(unweighted)},\qquad p_{o,\max}=\sum_{i,j}w_{ij}\,m_{ij}\;\text{(weighted)}
κmax⁡\kappa_{\max}
highest kappa compatible with the observed marginals
po,max⁡p_{o,\max}
maximum agreement attainable with those marginals
mijm_{ij}
north-west corner allocation: the table with those same marginals that maximises weighted agreement
Unweighted, the maximum is reached by filling the diagonal as far as the marginals allow. With linear or quadratic weights the north-west corner rule reaches it, because those weights satisfy Monge's condition.
PABAK=2po−1,PI=∣a−d∣n,BI=∣b−c∣n\mathrm{PABAK}=2p_o-1,\qquad \mathrm{PI}=\frac{|a-d|}{n},\qquad \mathrm{BI}=\frac{|b-c|}{n}
a, da,\ d
agreement cells of a 2 × 2 table
b, cb,\ c
disagreement cells of a 2 × 2 table
PI\mathrm{PI}
prevalence index
BI\mathrm{BI}
bias index
Byrt, Bishop and Carlin (1993); they are defined only with two categories. PABAK is the kappa you would get if both indices were 0.

R code

# Cohen's kappa (unweighted and weighted) for a k x k agreement table
#   - Bioestadistica abierta, UDG-CA-1190
# Runs as is in R, RStudio or webR; prints the results as JSON at the end.
library(irr)
library(jsonlite)

# The pasted table, row by row: rows = rater A, columns = rater B,
# the same categories in the same order.
x <- matrix(c(40, 8, 2, 6, 25, 4, 1, 3, 11), nrow = 3, byrow = TRUE)
ponderacion <- "ninguna"          # "ninguna" | "lineal" | "cuadratica"
nivel <- 0.95
tipo <- switch(ponderacion, ninguna = "unweighted", lineal = "equal", cuadratica = "squared")

# A category nobody used (empty row AND empty column) is dropped first: it adds no
# agreement, and the weights depend on how many categories there are.
vacias <- rowSums(x) == 0 & colSums(x) == 0
x <- x[!vacias, !vacias, drop = FALSE]
k <- nrow(x); n <- sum(x)

# Agreement weights w_ij (Cohen 1968); with k = 2 the three schemes are the identity.
w <- switch(tipo, unweighted = diag(k), equal = 1 - abs(outer(1:k, 1:k, "-"))/(k - 1),
            squared = 1 - (outer(1:k, 1:k, "-")/(k - 1))^2)
p <- x/n; pi <- rowSums(p); pj <- colSums(p)
po <- sum(w*p); pe <- sum(w*outer(pi, pj)); kappa <- (po - pe)/(1 - pe)

# Asymptotic standard error of Fleiss, Cohen & Everitt (1969), from the OBSERVED cells,
# and its Wald interval truncated to [-1, 1]. max(0, .) only absorbs the rounding of
# perfect agreement, where the radicand is 0 up to one ulp.
wi <- as.vector(w %*% pj); wj <- as.vector(t(w) %*% pi)   # weighted row and column means
ee <- sqrt(max(0, sum(p*(w - outer(wi, wj, "+")*(1 - kappa))^2) - (kappa - pe*(1 - kappa))^2))/((1 - pe)*sqrt(n))
z <- qnorm(1 - (1 - nivel)/2)
kappa_ic <- c(kappa, max(-1, kappa - z*ee), min(1, kappa + z*ee))

# z and p against kappa = 0 exactly as irr::kappa2, whose variance is the one under the
# null hypothesis of chance agreement (EXPECTED cells), not the one of the interval.
# kappa2 needs the ratings pair by pair, so they are rebuilt from the table; the
# categories are numbered from 11 because kappa2 sorts its levels as text, and with
# 1 to 10 it would read them "1", "10", "2" and shift the weights of a 10 x 10 table.
niveles <- 10 + 1:k
ratings <- cbind(rep(niveles, rowSums(x)), unlist(lapply(1:k, function(i) rep(niveles, x[i, ]))))
k2 <- kappa2(ratings, weight = tipo)

# Largest kappa these marginals allow. Unweighted that maximum is sum(min(pi, pj));
# with linear or quadratic weights the north-west corner rule reaches it, because
# |i - j| and (i - j)^2 satisfy Monge's condition.
esquina <- function(a, b) {
  m <- matrix(0, length(a), length(b)); i <- 1; j <- 1
  while (i <= length(a) && j <= length(b)) {
    v <- min(a[i], b[j]); m[i, j] <- m[i, j] + v; a[i] <- a[i] - v; b[j] <- b[j] - v
    if (a[i] <= 0) i <- i + 1 else j <- j + 1
  }
  m
}
po_max <- if (tipo == "unweighted") sum(pmin(pi, pj)) else sum(w*esquina(pi, pj))
kappa_max <- (po_max - pe)/(1 - pe)

# Only with two categories (Byrt, Bishop & Carlin 1993): PABAK removes the influence
# of prevalence and bias, and the two indices measure them.
pabak <- if (k == 2) 2*po - 1 else NA_real_
indice_prevalencia <- if (k == 2) abs(x[1, 1] - x[2, 2])/n else NA_real_
indice_sesgo <- if (k == 2) abs(x[1, 2] - x[2, 1])/n else NA_real_

res <- list(n = n, po = po, pe = pe, kappa = kappa_ic, ee = ee,
            z_h0 = unname(k2$statistic), p_h0 = k2$p.value, kappa_max = kappa_max,
            pabak = pabak, indice_prevalencia = indice_prevalencia,
            indice_sesgo = indice_sesgo)
cat(toJSON(res, auto_unbox = TRUE, digits = NA))

# Equivalent in RStudio (not run in the browser):
# DescTools::CohenKappa(x, weights = "Unweighted", conf.level = nivel)   # "Equal-Spacing", "Fleiss-Cohen"

This is the very code that validates the calculator: copy it and run it in R or RStudio to reproduce the result.

In-browser verification with R will arrive in a forthcoming version; meanwhile, copy the code and run it in RStudio.

Methods for a manuscript

Interobserver agreement was assessed with Cohen's kappa coefficient [1] (unweighted) with a 95% CI based on the asymptotic standard error of Fleiss, Cohen and Everitt [3]; its magnitude was interpreted following Landis and Koch [4]; the largest κ compatible with the observed marginals is also reported. 100 cases were classified into 3 categories and κ = 0.61 (95% CI 0.47 to 0.74) was obtained, with an observed agreement of 76.0%. Calculations used the "Cohen's kappa" calculator of Bioestadística abierta (Research Group UDG-CA-1190, https://udgca1190.com.mx/en/herramientas/bioestadistica/kappa), verified against R (irr::kappa2).

A paragraph ready for the Methods section; bracketed numbers refer to the reference list.

References

  1. 01 Cohen J. A coefficient of agreement for nominal scales. Educational and Psychological Measurement. 1960;20(1):37–46. doi:10.1177/001316446002000104 Original source
  2. 02 Cohen J. Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin. 1968;70(4):213–220. doi:10.1037/h0026256 PMID: 19673146 Original source
  3. 03 Fleiss JL, Cohen J, Everitt BS. Large sample standard errors of kappa and weighted kappa. Psychological Bulletin. 1969;72(5):323–327. doi:10.1037/h0028106 Original source
  4. 04 Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–174. doi:10.2307/2529310 PMID: 843571 Original source
  5. 05 Byrt T, Bishop J, Carlin JB. Bias, prevalence and kappa. Journal of Clinical Epidemiology. 1993;46(5):423–429. doi:10.1016/0895-4356(93)90018-V PMID: 8501467 Original source
  6. 06 Fleiss JL, Cohen J. The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability. Educational and Psychological Measurement. 1973;33(3):613–619. doi:10.1177/001316447303300309 Complementary
  7. 07 Gamer M, Lemon J, Fellows I, Singh P. irr: Various Coefficients of Interrater Reliability and Agreement. R package version 0.85. CRAN; 2026. Complementary
  8. 08 Sim J, Wright CC. The kappa statistic in reliability studies: use, interpretation, and sample size requirements. Physical Therapy. 2005;85(3):257–268. doi:10.1093/ptj/85.3.257 PMID: 15733050 Didactic reading
  9. 09 Altman DG. Practical Statistics for Medical Research. London: Chapman & Hall; 1991. Didactic reading