Cohen's kappa vs Fleiss' kappa: worked example for 2 vs 3+ raters
Search for either statistic on its own and you'll find a clean worked example: Cohen's kappa for two raters, or Fleiss' kappa for a classic textbook panel of psychiatrists. What neither kind of page shows is the same panel of judgments run through both calculations side by side — which is exactly the moment an LLM evaluation team actually needs this, when a third annotator joins a two-rater pilot and someone asks "can't we just average the three pairwise Cohen's kappas instead of learning a new formula?" The numbers below answer that directly, and the answer is: not quite, and here's the arithmetic that proves it.
Why there are two different statistics at all
Cohen's kappa (1960) is defined for exactly two raters labelling the same items. It compares the raters' observed agreement rate to the agreement you'd expect from pure chance, given how often each rater individually favours each label. Fleiss' kappa (1971) generalises this to any fixed number of raters per item — three, five, twenty — without needing those raters to even be the same people across every item. The two formulas share the same shape, (observed − chance) / (1 − chance), but they estimate "observed" and "chance" from different quantities, which is why they don't automatically reduce to the same number once you have more than two raters.
The dataset: three annotators scoring 10 LLM outputs
Three annotators (A, B, C) independently label 10 model outputs as Correct (C), Partial (P), or Incorrect (I) — a standard human-evaluation setup for checking whether an LLM-as-judge pipeline's category labels are reliable before trusting them:
| Item | Rater A | Rater B | Rater C |
|---|---|---|---|
| 1 | C | C | C |
| 2 | C | C | P |
| 3 | P | P | P |
| 4 | I | I | I |
| 5 | C | P | C |
| 6 | P | I | P |
| 7 | C | C | C |
| 8 | I | P | I |
| 9 | P | P | C |
| 10 | C | C | P |
Step 1: Cohen's kappa for one pair (A and B)
Ignore rater C for a moment and treat this as a standard two-rater check. A and B agree on items 1, 2, 3, 4, 7, 9 and 10 — 7 out of 10 — so observed agreement is p₀ = 7/10 = 0.70.
Expected (chance) agreement needs each rater's own label-use rate. Across the 10 items, A used C 5 times, P 3 times, I 2 times; B used C 4 times, P 4 times, I 2 times. Chance agreement is the sum, over each category, of the two raters' individual rates of using it:
pᵉ = (5/10 × 4/10) + (3/10 × 4/10) + (2/10 × 2/10) = 0.20 + 0.12 + 0.04 = 0.36
κ(A,B) = (0.70 − 0.36) / (1 − 0.36) = 0.34 / 0.64 = 0.531
Step 2: all three pairwise kappas — and why you can't just average them
Run the same calculation for the other two pairs. A and C also agree on 7 of 10 items, and their individual label-use rates happen to produce the same chance figure, so κ(A,C) = 0.531 too. B and C are a different story: they agree on only 4 of 10 items (1, 3, 4, 7), giving p₀ = 0.40 against the same pᵉ = 0.36, so κ(B,C) = (0.40 − 0.36) / 0.64 = 0.063 — barely above chance.
| Pair | Observed agreement | Chance agreement | Cohen's κ |
|---|---|---|---|
| A & B | 0.70 | 0.36 | 0.531 |
| A & C | 0.70 | 0.36 | 0.531 |
| B & C | 0.40 | 0.36 | 0.063 |
Average those three: (0.531 + 0.531 + 0.063) / 3 = 0.375. That single number quietly erases the fact that rater C tracks rater A reasonably well but barely tracks rater B at all — a pattern a team would genuinely want to know about before trusting C's labels. Averaging pairwise kappas isn't wrong because it's hard; it's wrong because a mean of three numbers can't carry information about which specific pair is driving the disagreement, and because it recomputes a separate chance-agreement baseline for every pair instead of using one baseline grounded in how often all three raters, together, actually used each label.
Step 3: Fleiss' kappa, computed directly on all three raters
Fleiss' kappa never splits the panel into pairs. For each item, count how many of the n = 3 raters chose each category, then measure how far that item's distribution is from "every rater picked the same label" (perfect per-item agreement). For item i, with n_ij the number of raters who chose category j on item i:
P_i = (sum over j of n_ij² − n) / (n × (n−1))
Item 1 (3 raters, all Correct): n_C=3, n_P=0, n_I=0, so P_1 = (9 − 3)/6 = 1.00 (full agreement). Item 2 (2 Correct, 1 Partial): P_2 = (4+1−3)/6 = 0.333. Working through all 10 items the same way and averaging gives the overall observed-agreement rate:
P-bar = (1.00 + 0.333 + 1.00 + 1.00 + 0.333 + 0.333 + 1.00 + 0.333 + 0.333 + 0.333) / 10 = 0.600
Chance agreement now comes from one pooled source: how often each category was used across all 30 ratings (10 items × 3 raters), not from any single pair's two marginals. Totals: Correct 13, Partial 11, Incorrect 6, out of 30:
p_C = 13/30 = 0.433, p_P = 11/30 = 0.367, p_I = 6/30 = 0.200
P-bar-e = 0.433² + 0.367² + 0.200² = 0.188 + 0.134 + 0.040 = 0.362
kappa_Fleiss = (0.600 − 0.362) / (1 − 0.362) = 0.238 / 0.638 = 0.373
Reading the two results side by side
| Method | Result | What it used |
|---|---|---|
| Average of 3 pairwise Cohen's κ | 0.375 | Three separate po/pe calculations, one per pair |
| Fleiss' κ (direct) | 0.373 | One pooled po/pe calculation across all three raters at once |
In this dataset the two land within a rounding error of each other — close enough to tempt anyone into thinking the shortcut is harmless. That closeness is a property of this particular dataset, not a guarantee: the three pairwise kappas that went into the average (0.531, 0.531, 0.063) span a huge range, and nothing about Fleiss' single 0.373 is obtained by averaging them — it comes from the n_ij counts directly, which is also why Fleiss' kappa scales cleanly to five, ten or twenty raters, where the number of pairwise combinations would otherwise explode (20 raters means 190 pairs to average, each needing its own po/pe). Both results land in the same qualitative band on the standard Landis & Koch (1977) interpretation scale — 0.21–0.40 is "fair" agreement — which is the practical verdict here: these three annotators are not yet reliable enough to treat their category labels as ground truth without further calibration or a clearer labelling guide.
Common mistakes
- Reaching for Cohen's kappa the moment a third rater joins the panel. Cohen's original formula has no term for a third rater at all; "generalising" it by hand (as the pairwise average above) is a different, less informative calculation, not a shortcut to the same one.
- Assuming a close match between the two numbers in one dataset means the methods are interchangeable. They happened to agree here because this example's category proportions are similar across raters; change the data and the gap can widen substantially.
- Forgetting that Fleiss' kappa's chance term,
P-bar-e, uses the pooled category rate across every rating in the dataset — not any one rater's individual rate, and not a simple average of the pairwise chance terms either.
Practice distinguishing these two statistics, and the related evaluation metrics and benchmarks that also show up in LLM evaluation pipelines, with PassDrill's LLM evaluation, testing & safety practice questions.
Source: Cohen, J. (1960), "A Coefficient of Agreement for Nominal Scales," Educational and Psychological Measurement, 20(1), 37–46; Fleiss, J.L. (1971), "Measuring Nominal Scale Agreement Among Many Raters," Psychological Bulletin, 76(5), 378–382; Landis, J.R. & Koch, G.G. (1977), "The Measurement of Observer Agreement for Categorical Data," Biometrics, 33(1), 159–174, for the 0.21–0.40 "fair agreement" interpretation band. The worked dataset above is an original example built for this page, not drawn from any cited paper.