passdrill

Cohen's kappa vs Fleiss' kappa: worked example for 2 vs 3+ raters

Search for either statistic on its own and you'll find a clean worked example: Cohen's kappa for two raters, or Fleiss' kappa for a classic textbook panel of psychiatrists. What neither kind of page shows is the same panel of judgments run through both calculations side by side — which is exactly the moment an LLM evaluation team actually needs this, when a third annotator joins a two-rater pilot and someone asks "can't we just average the three pairwise Cohen's kappas instead of learning a new formula?" The numbers below answer that directly, and the answer is: not quite, and here's the arithmetic that proves it.

Why there are two different statistics at all

Cohen's kappa (1960) is defined for exactly two raters labelling the same items. It compares the raters' observed agreement rate to the agreement you'd expect from pure chance, given how often each rater individually favours each label. Fleiss' kappa (1971) generalises this to any fixed number of raters per item — three, five, twenty — without needing those raters to even be the same people across every item. The two formulas share the same shape, (observed − chance) / (1 − chance), but they estimate "observed" and "chance" from different quantities, which is why they don't automatically reduce to the same number once you have more than two raters.

The dataset: three annotators scoring 10 LLM outputs

Three annotators (A, B, C) independently label 10 model outputs as Correct (C), Partial (P), or Incorrect (I) — a standard human-evaluation setup for checking whether an LLM-as-judge pipeline's category labels are reliable before trusting them:

ItemRater ARater BRater C
1CCC
2CCP
3PPP
4III
5CPC
6PIP
7CCC
8IPI
9PPC
10CCP

Step 1: Cohen's kappa for one pair (A and B)

Ignore rater C for a moment and treat this as a standard two-rater check. A and B agree on items 1, 2, 3, 4, 7, 9 and 10 — 7 out of 10 — so observed agreement is p₀ = 7/10 = 0.70.

Expected (chance) agreement needs each rater's own label-use rate. Across the 10 items, A used C 5 times, P 3 times, I 2 times; B used C 4 times, P 4 times, I 2 times. Chance agreement is the sum, over each category, of the two raters' individual rates of using it:

pᵉ = (5/10 × 4/10) + (3/10 × 4/10) + (2/10 × 2/10) = 0.20 + 0.12 + 0.04 = 0.36

κ(A,B) = (0.70 − 0.36) / (1 − 0.36) = 0.34 / 0.64 = 0.531

Step 2: all three pairwise kappas — and why you can't just average them

Run the same calculation for the other two pairs. A and C also agree on 7 of 10 items, and their individual label-use rates happen to produce the same chance figure, so κ(A,C) = 0.531 too. B and C are a different story: they agree on only 4 of 10 items (1, 3, 4, 7), giving p₀ = 0.40 against the same pᵉ = 0.36, so κ(B,C) = (0.40 − 0.36) / 0.64 = 0.063 — barely above chance.

PairObserved agreementChance agreementCohen's κ
A & B0.700.360.531
A & C0.700.360.531
B & C0.400.360.063

Average those three: (0.531 + 0.531 + 0.063) / 3 = 0.375. That single number quietly erases the fact that rater C tracks rater A reasonably well but barely tracks rater B at all — a pattern a team would genuinely want to know about before trusting C's labels. Averaging pairwise kappas isn't wrong because it's hard; it's wrong because a mean of three numbers can't carry information about which specific pair is driving the disagreement, and because it recomputes a separate chance-agreement baseline for every pair instead of using one baseline grounded in how often all three raters, together, actually used each label.

Step 3: Fleiss' kappa, computed directly on all three raters

Fleiss' kappa never splits the panel into pairs. For each item, count how many of the n = 3 raters chose each category, then measure how far that item's distribution is from "every rater picked the same label" (perfect per-item agreement). For item i, with n_ij the number of raters who chose category j on item i:

P_i = (sum over j of n_ij² − n) / (n × (n−1))

Item 1 (3 raters, all Correct): n_C=3, n_P=0, n_I=0, so P_1 = (9 − 3)/6 = 1.00 (full agreement). Item 2 (2 Correct, 1 Partial): P_2 = (4+1−3)/6 = 0.333. Working through all 10 items the same way and averaging gives the overall observed-agreement rate:

P-bar = (1.00 + 0.333 + 1.00 + 1.00 + 0.333 + 0.333 + 1.00 + 0.333 + 0.333 + 0.333) / 10 = 0.600

Chance agreement now comes from one pooled source: how often each category was used across all 30 ratings (10 items × 3 raters), not from any single pair's two marginals. Totals: Correct 13, Partial 11, Incorrect 6, out of 30:

p_C = 13/30 = 0.433,   p_P = 11/30 = 0.367,   p_I = 6/30 = 0.200

P-bar-e = 0.433² + 0.367² + 0.200² = 0.188 + 0.134 + 0.040 = 0.362

kappa_Fleiss = (0.600 − 0.362) / (1 − 0.362) = 0.238 / 0.638 = 0.373

Reading the two results side by side

MethodResultWhat it used
Average of 3 pairwise Cohen's κ0.375Three separate po/pe calculations, one per pair
Fleiss' κ (direct)0.373One pooled po/pe calculation across all three raters at once

In this dataset the two land within a rounding error of each other — close enough to tempt anyone into thinking the shortcut is harmless. That closeness is a property of this particular dataset, not a guarantee: the three pairwise kappas that went into the average (0.531, 0.531, 0.063) span a huge range, and nothing about Fleiss' single 0.373 is obtained by averaging them — it comes from the n_ij counts directly, which is also why Fleiss' kappa scales cleanly to five, ten or twenty raters, where the number of pairwise combinations would otherwise explode (20 raters means 190 pairs to average, each needing its own po/pe). Both results land in the same qualitative band on the standard Landis & Koch (1977) interpretation scale — 0.21–0.40 is "fair" agreement — which is the practical verdict here: these three annotators are not yet reliable enough to treat their category labels as ground truth without further calibration or a clearer labelling guide.

Common mistakes

Practice distinguishing these two statistics, and the related evaluation metrics and benchmarks that also show up in LLM evaluation pipelines, with PassDrill's LLM evaluation, testing & safety practice questions.

Source: Cohen, J. (1960), "A Coefficient of Agreement for Nominal Scales," Educational and Psychological Measurement, 20(1), 37–46; Fleiss, J.L. (1971), "Measuring Nominal Scale Agreement Among Many Raters," Psychological Bulletin, 76(5), 378–382; Landis, J.R. & Koch, G.G. (1977), "The Measurement of Observer Agreement for Categorical Data," Biometrics, 33(1), 159–174, for the 0.21–0.40 "fair agreement" interpretation band. The worked dataset above is an original example built for this page, not drawn from any cited paper.

Drill Evaluation, Testing & Safety practice questions →