Data Annotation
How to Measure Data Annotation Quality: Inter-Annotator Agreement, Gold Sets, and Review (2026 Guide)
The three metrics that actually measure annotation quality, how to calculate inter-annotator agreement, when to use gold sets, and how to catch reviewer drift.

Data annotation quality is usually measured the wrong way, by spot-checking a sample for obvious mistakes. That catches careless errors but misses the larger problem: trained annotators disagreeing on hard items, which enters your training data as noise your model cannot learn out. This guide explains how to actually measure annotation quality, covering the three metrics that matter, how to calculate inter-annotator agreement, when to use gold sets, and how to catch reviewer drift.
What is data annotation quality?
Data annotation quality is the degree to which labels in a training dataset are accurate, consistent, and fit for training a model. High quality means two things at once: the labels are correct against some reference, and independent annotators applying the same guidelines produce the same labels. A dataset can be accurate on average and still be low quality if annotators disagree on hard cases, because that disagreement becomes noise the model inherits.
Most teams measure only the first half, correctness on a sample, and never measure the second half, consistency between annotators. The second half is where most real quality problems live.
The three metrics that actually measure annotation quality

There is no single number for annotation quality. There are three, and each answers a different question.
Inter-annotator agreement (IAA) answers "do our annotators agree with each other?" It measures whether two people, labeling the same item independently, produce the same label. Low agreement means the guideline is ambiguous or the task is genuinely hard, and it is the earliest and most overlooked signal of a quality problem.
Accuracy against a gold set answers "are our annotators right?" A gold set is a batch of items labeled by an expert whose answers are treated as correct. You measure annotator output against it with precision, recall, and F1. This catches careless work and annotators who have drifted from the guidelines.
Review reliability answers "is our quality-control layer still working?" Reviewers decay over long sessions and start approving whatever is in front of them. You catch this by seeding the review queue with known-answer items and tracking whether reviewers catch them.
Use all three. Agreement without accuracy means your annotators are consistently wrong together. Accuracy without agreement means your average looks fine while hard cases are a coin flip. Neither alone tells you whether your dataset is trainable.
What is inter-annotator agreement and how do you calculate it?
Inter-annotator agreement (IAA) is the rate at which two or more annotators, labeling the same items independently, produce the same label. It is the core consistency metric for annotation quality.
The naive version is raw percent agreement: the share of items where annotators matched. The problem is that some agreement happens by chance, especially with few label classes, so percent agreement overstates quality. The standard metrics correct for chance.
Cohen's kappa measures agreement between two annotators on categorical labels, corrected for the agreement expected by chance. It ranges from below 0 (worse than random) to 1 (perfect). A common reading: below 0.40 is poor, 0.40 to 0.60 is moderate, 0.60 to 0.80 is substantial, and above 0.80 is strong, though the right threshold depends on how costly an error is in your domain.
Fleiss' kappa extends the same idea to more than two annotators, which matches how production labeling actually runs.
Krippendorff's alpha is the most flexible: it handles any number of annotators, missing data, and different data types (nominal, ordinal, interval), which makes it the safest default for real annotation pipelines that are rarely clean.
You calculate IAA by having two or more annotators label the same sample independently, with no conversation between them, then applying the metric to the pairs. The independence matters: if annotators discuss items first, you measure their consensus process, not whether the guideline is decidable on its own.
Why do good annotators disagree?
Trained, attentive annotators disagree for reasons that have nothing to do with skill or effort. The three main causes:
- Ambiguous guidelines. The rules never decided the edge case. "Label vehicles" said nothing about a vehicle on a billboard, a reflection, or a toy, so each annotator improvises differently. This is the largest and most fixable cause.
- Genuinely hard data. For some items there is no single answer two careful people independently reach. Occlusion boundaries, sentiment on sarcastic text, and intent classification are classic examples.
- Guideline drift. Over a long project, annotators quietly diverge in how they interpret the same rule, and agreement decays even though nobody got worse at the task.
The fix for all three is the same in shape: make the label more decidable, not more detailed. Replace judgment calls with checkable questions. "Is this occluded?" is a judgment call. "Is this specific keypoint visible, yes or no?" is a fact two people can check and agree on. We cover this mechanism in depth in why disagreement, not carelessness, is the real quality problem.
What is a gold set and how do you use it?
A gold set (also called a gold standard or ground-truth set) is a collection of items labeled by an expert or by consensus, whose labels are treated as correct and used to measure everyone else against. It is the reference you need to answer "are the annotators right," as opposed to merely "do they agree."
You use a gold set in three ways:
- Onboarding: have new annotators label it and measure their accuracy before they touch production data.
- Ongoing accuracy: mix gold items into live work, hidden so they cannot be spotted, and track accuracy over time to catch drift.
- Review reliability: seed gold items into the reviewer's queue to check that the quality-control layer itself is still catching errors.
A gold set is only as good as its own labels, so it should be built by your most expert people and revisited when the guidelines change, since a gold set labeled under old rules silently measures annotators against a standard you no longer use.
How do you measure annotation quality over time?

Measuring quality once, at the start of a project, is the most common mistake. Quality is not a fixed property of a dataset; it decays as production data drifts away from what the guidelines anticipated. A workflow that measures continuously looks like this:
- Write the guideline as a decision procedure, not a description. Its job is to make labels decidable.
- Have two annotators label a sample independently at the start, with no conversation.
- Measure inter-annotator agreement and read every disagreement. Each one points at an ambiguous rule.
- Check output against a gold set to confirm the agreed labels are also correct.
- Feed failures back into the guideline, split the ambiguous labels, and remeasure.
- Keep measuring agreement during production, not just at onboarding.
That last step is the one that pays off most and is skipped most. A drop in agreement mid-project is a leading indicator: it tells you production data has hit a case the guideline never covered, and it moves weeks before model accuracy does. Measured continuously, it is free early warning. Measured once, it is a smoke detector unplugged after the first test.
Annotation quality metrics compared
| Metric | Answers | Best for | Watch out for |
|---|---|---|---|
| Percent agreement | Do annotators match? | A quick first look | Overstates quality; ignores chance |
| Cohen's kappa | Do two annotators agree, beyond chance? | Two-annotator categorical tasks | Only handles two annotators |
| Fleiss' kappa | Do several annotators agree? | Production labeling with many annotators | Categorical labels only |
| Krippendorff's alpha | Do annotators agree, any setup? | Messy real pipelines, missing data, ordinal labels | More complex to compute |
| Accuracy vs. gold (P/R/F1) | Are annotators right? | Catching careless and drifting work | Needs a well-built gold set |
| Seeded review reliability | Is the reviewer still catching errors? | Catching reviewer drift | Known items must be unspottable |
Does adding more reviewers improve annotation quality?
Adding reviewers helps less than teams expect, because a reviewer is a third opinion, not a source of truth. Review catches careless mistakes well. It does not resolve genuine disagreement, since the reviewer simply adds another judgment to the pile. And review layers decay: a reviewer late in a long session starts approving by default, so a second or third layer can cost money while catching almost nothing.
The higher-leverage fixes, in order: make the guideline more decidable so disagreement falls at the source, measure agreement continuously so you catch drift early, and verify the review layer itself with seeded gold items. Review is part of a quality system, not a substitute for one.
Frequently asked questions
What is a good inter-annotator agreement score? It depends on the cost of an error. As a rough guide, Cohen's or Fleiss' kappa above 0.80 is strong, 0.60 to 0.80 is usually acceptable, and below 0.60 signals an ambiguous guideline or a task too hard as defined. Safety-critical work needs higher thresholds than low-stakes tagging.
What is the difference between inter-annotator agreement and accuracy? Agreement measures whether annotators match each other. Accuracy measures whether they match a correct reference (a gold set). You can have high agreement and low accuracy if everyone is consistently wrong, so you need both.
Which inter-annotator agreement metric should I use? For two annotators and categorical labels, Cohen's kappa. For more than two annotators, Fleiss' kappa. For mixed data types, missing data, or messy real pipelines, Krippendorff's alpha is the safest default.
How large should a gold set be? Large enough to be statistically meaningful for your class balance and refreshed when guidelines change. A few hundred items covering your hard and edge cases is more useful than thousands of easy ones.
Can you automate annotation quality control? Partly. Models can pre-annotate and flag likely errors, which speeds review, but automated QA cannot resolve genuine disagreement or decide ambiguous edge cases, because those require a decision about what the label should mean. The decision is human; the flagging can be automated.
Getting annotation quality right
Measuring annotation quality means answering three questions, not one: do annotators agree, are they right, and is the review layer still working. Start by running a blind double-labeling pass on a real sample, measuring inter-annotator agreement, and reading every disagreement, because those disagreements are the map to every ambiguous rule in your guidelines.
If you want a second opinion on a sample, send it over and we will label it blind against yours and report the agreement and the gaps. First read is free.
Frequently asked questions
- What is a good inter-annotator agreement score?
- It depends on the cost of an error. As a rough guide, Cohen's or Fleiss' kappa above 0.80 is strong, 0.60 to 0.80 is usually acceptable, and below 0.60 signals an ambiguous guideline or a task too hard as defined. Safety-critical work needs higher thresholds than low-stakes tagging.
- What is the difference between inter-annotator agreement and accuracy?
- Agreement measures whether annotators match each other. Accuracy measures whether they match a correct reference (a gold set). You can have high agreement and low accuracy if everyone is consistently wrong, so you need both.
- Which inter-annotator agreement metric should I use?
- For two annotators and categorical labels, Cohen's kappa. For more than two annotators, Fleiss' kappa. For mixed data types, missing data, or messy real pipelines, Krippendorff's alpha is the safest default.
- How large should a gold set be?
- Large enough to be statistically meaningful for your class balance and refreshed when guidelines change. A few hundred items covering your hard and edge cases is more useful than thousands of easy ones.
- Can you automate annotation quality control?
- Partly. Models can pre-annotate and flag likely errors, which speeds review, but automated QA cannot resolve genuine disagreement or decide ambiguous edge cases, because those require a decision about what the label should mean. The decision is human; the flagging can be automated.

