All posts

Data Annotation

You Don't Have a Labeling Problem. You Have a Disagreement Problem.

Most data quality problems are not careless mistakes. They are disagreement between trained annotators, poured into your training set as noise. Here is the fix.

You don't have a labeling problem. You have a disagreement problem.

Here is a test you can run on any annotation vendor, including your own in-house team. Take fifty items you have already labeled and labeled well. Strip the labels. Hand the same fifty to two of your best, most experienced annotators, separately. Then compare what comes back.

They will not match. Not on all fifty. On the hard ones, the ones that actually decide whether your model works, two people who both know the guidelines cold will label the same item differently. That gap is the real thing you are buying when you buy annotation, and almost nobody measures it.

Most conversations about data quality are about catching mistakes: the stray wrong box, the missed object, the careless click at the end of a long shift. That is worth catching. But it is the thin sliver of the problem. The thick part is that for genuinely hard data, there often is no single right answer that two careful people will independently arrive at, and every disagreement they have gets poured into your training set as noise your model can never learn its way out of.

The definitions that matter

Three terms, defined sharply, because the rest of this post runs on them.

Agreement is the rate at which two people, labeling the same item independently, produce the same label. Not whether they are right. Whether they match.

A guideline is a decision procedure. Its job is not to describe what a good label looks like. Its job is to make the label decidable, so that two people following it land in the same place without talking to each other.

Ground truth is the fiction that there is one correct label waiting to be found. For clear data it is real. For hard data it is manufactured, and the quality of your dataset is the quality of that manufacturing process.

Hold those, and a lot of annotation folklore falls apart. "Hire better annotators" does not fix disagreement if the guideline leaves the call ambiguous. "Add more reviewers" does not fix it either, as we will see. The thing you tune is the decision procedure, and the number you watch is agreement.

How disagreement becomes model noise

One image, two annotators, two labels, both poured into the training set as noise the model cannot train out.

Walk the mechanism, because it is the whole argument.

A model learns a distinction only if its labels contain that distinction cleanly. Show it a thousand images where "occluded" means one thing to annotator A and something else to annotator B, and you have not taught it what occluded means. You have taught it that occluded is a coin flip. The model does exactly what you trained it to do: it reproduces the coin flip. Then it fails in production, and everyone blames the model.

This is why accuracy against a held-out test set can look fine while the model misbehaves on real inputs. If your test set carries the same disagreement as your training set, the model scores well on reproducing the confusion. You have measured it against its own blind spot.

The uncomfortable part: this damage is invisible on every dashboard you own. A dataset built from two annotators who agreed 70% of the time looks identical, in every count and every chart, to a dataset built from two who agreed 95% of the time. Same number of labels, same class balance, same everything a spreadsheet can see. The 25-point difference in signal is simply gone, absorbed silently, waiting to surface as a model problem you will spend three months debugging in the wrong place.

Where error actually lives

A stacked comparison: teams look at careless mistakes, but error actually lives in disagreement, ambiguous guidelines, and reviewer drift.

When a labeled dataset underperforms, teams reach for the explanation they can see: someone was careless. So they add a QA pass to catch careless work. It helps a little, because some work is careless. But it leaves the bulk of the error untouched, because the bulk of the error was never carelessness.

Real label error has three sources, and only the smallest one is what QA is built to catch.

The first is disagreement: two trained, attentive people genuinely see the item differently. No amount of reviewing fixes this, because the reviewer is a third person who also has an opinion. You cannot review your way to agreement. You can only define your way there.

The second is ambiguous guidelines: the edge case the rules never decided. The guideline said "label vehicles" and said nothing about whether a vehicle on a billboard counts, or a vehicle reflected in a window, or a child's toy car. Every annotator now improvises, and they improvise differently. This is the largest source for most projects and the most fixable, which is a rare and happy combination.

The third is reviewer drift: the quality layer itself decays. A reviewer approving their four-hundredth item of the day starts agreeing with whatever is in front of them. The second line of defense quietly becomes a rubber stamp that costs money and catches nothing. We will come back to this, because it is the failure mode that makes "just add more review" actively dangerous.

Careless mistakes are real. They are also the part you already know how to fix. The other three-quarters is where your dataset is actually bleeding.

The move: make labels decidable, not detailed

A vague label splits step by step into a checkable question, raising agreement at each step.

Here is the counterintuitive fix. When agreement is low, the instinct is to write a longer, more detailed guideline. More detail usually makes it worse, because more words mean more clauses to interpret differently. The goal is not a more detailed label. It is a more decidable one.

Take "occluded," a genuine troublemaker. As written, it is a judgment call, and judgment calls do not agree. Split it. "Is more than half the object hidden?" is better, and agreement climbs, though people still argue about "half." Split again. "Is this specific keypoint visible, yes or no?" is a checkable fact, and now two people who have never met produce the same answer.

Notice what happened across those steps. The label did not get richer. It got more decidable. Each split replaced a judgment with a fact that two people can check independently and land on together. That is the entire craft of guideline design, and it is measurable: you split, you remeasure agreement, you keep the splits that move the number.

This reframes what a guideline is for. It is not documentation written once and filed. It is a decision procedure you tune against a number, the same way you tune a model against a loss. Teams that treat guidelines as a one-time writing task and teams that treat them as a tunable system produce datasets that are not in the same league, and the difference does not show up until the model does.

Agreement is your earliest warning, and the only one you get for free

A timeline: inter-annotator agreement falls weeks before model error rises in production, opening a warning window.

Everything so far has been about building the dataset. This is about running it over time, and it is the part most teams never set up.

Production data drifts. The supplier changes packaging, the camera gets remounted, users start phrasing requests in ways nobody has seen. When that happens, your annotators hit items the guideline never anticipated, and the first thing that moves, before model accuracy budges at all, is agreement. People stop matching, because the world handed them a case the decision procedure does not cover.

That drop is a leading indicator. Model error in production is a lagging one: by the time it shows up on a metrics dashboard, the bad labels are already trained in and the damage is weeks old. Agreement falls the moment the new case appears, which means the window between "agreement drops" and "model error climbs" is free warning time, if you are measuring agreement continuously instead of checking it once at the start of a project.

Most teams measure agreement exactly once, during onboarding, to decide whether an annotator is any good. Then they never look again. They have a smoke detector and they unplug it after the first test. Measured continuously, falling agreement tells you a new edge case has arrived, points you at exactly which items caused it, and buys you the time to update the guideline before the confusion reaches your model.

What we had to admit, and what we built around it

We ran into the wall that makes all of this harder than it sounds: the reviewer is not a source of truth. They are just a better-paid third opinion. For a long stretch, the standard answer to low quality was to stack another review layer on top, and we did it too, until we measured what those layers actually caught and found reviewer drift eating the benefit. A tired reviewer does not resolve disagreement. They add a third data point to it and call it resolved.

We have not solved that cleanly, and as far as we can tell no one has. What we do is refuse to trust the review layer on faith. We seed every reviewer's queue with items whose correct answer we already know, mixed in where they cannot be spotted, and we track whether those known items get caught. Reviewer reliability stops being a virtue we assume and becomes a number we watch, the same way agreement is. When the number slips, we rotate or retrain before the drift reaches your data. It is containment, not a cure. If a vendor tells you their reviewers simply do not drift, ask how they would know.

Where to start

Do not start by hiring annotators or buying a tool. Start by finding out whether your people agree.

Take one slice of your real data. Have two qualified people label it independently, with no conversation between them. Measure how often they match, and read every case where they did not. Those disagreements are a map: they show you exactly where your guideline is ambiguous, which distinctions are too hard to call as written, and what your real error rate is, as opposed to the one your test set flatters you with.

Run the two-person independent pass on fifty items this week, and let the disagreements tell you what your guidelines never decided.

If you want a second pair of eyes on that slice, send it over. We will label it blind against yours and show you where the gap is. First read is free.

Frequently asked questions

The rate at which two people labeling the same item independently produce the same label; a measure of whether they match, not whether they are right.
For genuinely hard data there is often no single answer two careful people independently reach, because the guideline left the call ambiguous.
No; a reviewer is a third opinion, not a source of truth, and review layers decay into rubber-stamping. You fix disagreement by making the guideline more decidable, not by adding review.
Measure inter-annotator agreement continuously; a drop is the earliest warning that production data has drifted, well before model error appears.