All posts

Data Annotation

Inter-Annotator Agreement: The Metric Your Dataset Is Quietly Failing

A 98% QA pass rate means almost nothing if two trained annotators label the same image differently. Here is how to measure agreement properly, pick the right statistic, and set gates that stop bad labels before they reach training.

Two rows of label cells for annotator A and annotator B with four disagreements highlighted, beside a kappa score of 0.61 against a target of 0.80 and an agreement-by-slice breakdown.

A few months ago I sat in a review call where the delivery team presented a dataset with a 98% QA pass rate. The client's ML lead looked pleased for about ten seconds, then asked the question that should always be asked: "What was the agreement between annotators on the hard classes?"

Nobody had that number. So we ran it on a sample of 500 items. On the majority classes, agreement was fine. On the two classes that actually mattered to the model, Cohen's kappa came back at 0.61. The dataset looked clean because QA was checking whether labels followed the guideline. It was not checking whether two competent people would have written the same label in the first place.

That is the gap this post is about. Pass rates measure compliance. Agreement measures whether your label definition is even stable enough to learn from.

The two numbers teams confuse

There are two different questions you can ask about a label:

  1. Did the annotator follow the instructions? This is what QA sampling answers, and it is what produces "98% pass rate" slides.
  2. Would a different qualified person produce the same label? This is inter-annotator agreement, or IAA.

The first is necessary. The second is the one that predicts whether your model can learn the task. A guideline can be perfectly followed by every annotator and still be ambiguous, because everyone reads the same ambiguous sentence and applies it consistently—to themselves.

If agreement is low, the problem is almost never "bad annotators." It is one of three things:

  • The label definition has an unresolved edge case (what do you call a partially visible object?)
  • Two labels overlap (is this a "crosswalk" or "road marking"?)
  • The task needs context the annotator doesn't have (only the client knows the policy for this transaction)

You cannot fix any of those with more QA sampling. You fix them by measuring disagreement and then changing the guideline.

Pick the statistic that fits the task

This is where most teams grab Cohen's kappa off the shelf and get a misleading answer. Different tasks need different measurements.

SituationUseWhy
Two annotators, mutually exclusive classesCohen's kappaBaseline, but punishes rare classes unfairly
Three or more annotators, any number can label each itemFleiss' kappa or Krippendorff's alphaHandles missing and unequal ratings
Very imbalanced labels (one class is 2% of data)Gwet's AC1 or prevalence-adjusted kappaKappa collapses to zero on rare-but-agreed classes
Span or entity extraction (text)Span-level F1 on matched spans, plus boundary agreementToken-level agreement hides boundary disputes
Bounding boxesIoU-thresholded agreement (e.g. matched at IoU ≥ 0.5)There is no "exact" box to compare
Ranking or preference (RLHF)Pairwise agreement + transitivity rateOrdering has its own failure modes

The prevalence problem deserves a sentence of its own. If 98% of your items are label A and 2% are label B, an annotator who labels everything A agrees with a careful annotator 98% of the time. Kappa tries to correct for chance agreement and can crash toward 0 even when people are doing a reasonable job. If you report kappa on a rare safety class and it looks catastrophic, check whether the class is just rare before you fire anyone.

For safety-critical work I tend to report more than one number: percent agreement (so stakeholders understand the raw picture), a chance-corrected statistic, and agreement broken down by slice rather than globally.

What "good" looks like, by risk

There is no universal threshold, but here is the calibration I use when scoping a project. The right move is to agree on these numbers before annotation starts, with the client, and write them into the QA plan.

Task typeAcceptable agreementAction if below
Subjective quality rating (tone, helpfulness)α ≥ 0.60Tighten rubric, add examples, accept more noise
Standard object detection (cars, people)κ ≥ 0.80Retrain annotators on disputed examples
Safety / harm classificationκ ≥ 0.80 and 100% adjudication on borderlineDo not ship below this
Named entity extraction (clear types)Span F1 ≥ 0.85Clarify boundary rules
Clinical or legal judgmentDepends on the gold standardAdd expert adjudication layer

The subjective row is uncomfortable for clients. It is also realistic. If the task is genuinely subjective—"is this reply polite?"—you will not get 0.9 agreement, and pretending otherwise creates pressure to game the numbers. Better to accept 0.65 on tone and design the pipeline around it than to fake consensus.

The calibration round: 90 minutes that save weeks

Before any large batch goes into production, run a calibration round. It is cheap and it finds the guideline holes while they are still cheap to fix.

  1. Sample 50–100 items that intentionally include edge cases, not just easy typical examples.
  2. Have 3 annotators label them independently, with no discussion.
  3. Compute agreement on the whole set and per class.
  4. Pull every disputed item into a live discussion. This is the important part. You are not looking for who was right; you are looking for why reasonable people diverged.
  5. Rewrite the guideline to close the gap, and add the disputed items as documented examples with the final decision.
  6. Re-run a fresh sample and confirm the numbers moved.

I have watched a single calibration round turn a 0.58 into a 0.82 on the same task, without changing a single annotator, purely by resolving eleven ambiguous cases in the guideline. That is usually the cheapest accuracy gain available to a project.

Document the disputed examples in an internal "label decisions" log with a date and the person who decided. Future annotators will hit the same case, and future you will not remember the reasoning.

Monitoring after launch

Agreement is not a one-time gate. It drifts.

  • Embed hidden gold items (pre-agreed labels, unknown to the annotator) into live batches at 3–5%. This catches individual drift without telling anyone they are being evaluated.
  • Re-run an overlap sample every batch. 2–5% double-annotated is enough to compute a live agreement number.
  • Track agreement by annotator, by class, and by slice. Global agreement can look healthy while one annotator quietly disagrees on one class. Slice it by time-of-day, by geography, by source dataset, by language.
  • Watch the guidance channel. If the same question comes up twice from two annotators, it is a guideline gap, not an annotator problem.

One detail that saves a lot of pain: define "the label" before you fight about it. For bounding boxes, agree on the occlusion rule (do you box a 30%-visible pedestrian?), the minimum size rule (do you box a 12-pixel object?), and the grouping rule (five grapes, one box or five?). Most detection disagreements I have debugged were never about annotator skill. They were about those three rules being unwritten.

When you should not chase agreement

There is one case where low agreement is the correct answer, and teams sometimes miss it: when the task is genuinely a judgment the business has not yet made. If two policy experts disagree about whether a given model output should be allowed, that disagreement is a product decision in disguise. Escalate it. Do not bury it in a kappa score and hand it to the training pipeline as noise.

Agreement is a measuring instrument. Its highest-value reading is often the one that tells you a decision needs to be made upstream.

A short checklist

  • Agree on the agreement statistic and the target before annotation begins
  • Run a calibration round with deliberate edge cases; rewrite the guideline from the disputes
  • Keep a dated log of label decisions and reuse them as examples
  • Report agreement by slice, not just globally; watch rare classes for the prevalence trap
  • Embed gold items at 3–5% and re-sample overlaps every batch
  • Treat persistent disagreement across annotators as a guideline or product problem, never as an annotator problem

Good labels are not the ones nobody complained about. They are the ones two qualified strangers would produce independently. If you cannot measure that, you do not yet know the quality of your data—you only know it was processed.

If you are assessing agreement on an existing dataset, or designing a QA plan with defined gates for a new one, our team at DATACLAP DIGITAL runs calibration rounds and multi-pass human-in-the-loop review as a standard part of delivery.