RLHF
Preference Data Is Not a Ranking Problem
If you collect comparisons the way you collect labels, your reward model will learn your annotators' habits instead of your users' preferences. Here is what to change in the collection design before you touch the training loop.

The most common instruction in RLHF collection is "pick the better response." It sounds self-evident and it produces bad data, because it treats preference as a ranking task when it is really a measurement task. You are not asking an annotator to order two texts. You are asking them to reveal a preference your users hold, through a human proxy who has their own biases, fatigue curve, and interpretation of "better."
Get that wrong and the reward model does exactly what you asked: it learns to reproduce your annotators. Then you optimise a policy against it and discover you have trained a model that is verbose, agreeable, and confidently wrong—the fingerprints of whichever annotators wrote the most comparisons.
The five ways preference collection goes wrong
1. Position bias. Swapping A and B changes the answer more often than anyone expects, especially for near-ties. If you never randomise or test for it, an unknown fraction of your data is noise.
2. Length bias. Longer responses are preferred even when they add nothing, if length correlates with thoroughness anywhere in your data. Once the model learns this, it writes more—and users pay for tokens that carry no information.
3. Unresolved ties. Forcing a choice on a genuine tie manufactures a preference that doesn't exist. But so does allowing an unchecked "tie" option, which some annotators choose to finish faster. Ties need a definition and a rate check.
4. Annotator drift and idiosyncrasy. Preferences shift across a long session, and some annotators develop stable personal quirks ("polite openings are good"). With many annotators, the aggregate can be fine while individual contributors inject systematic signal about themselves.
5. Subtle self-consistency failures. Annotators contradict themselves on the same pair shown twice in different contexts, and on cycles where A > B, B > C, but C > A. Transitivity violations are a direct measure of how noisy your preference signal is.
Design the collection, not just the prompt
Fixes that actually change the quality of the resulting reward model:
Write a rubric, then test it. Give annotators explicit criteria—correctness, instruction adherence, factual grounding, safety—and a priority order between them. Then run a calibration round where several annotators label the same pairs and you measure agreement. If agreement on "which is better" is low, your rubric is not doing its job. Tighten it before collecting at volume.
Randomise position and test for bias. Present A/B in random order, and periodically show the same pair reversed. Measure how often the answer flips. If it flips often, your comparisons on that topic are mostly noise and should be weighted down or dropped.
Control for length deliberately. Either balance your pairs by response length, or explicitly include length as a criterion so annotators are aware of it. Report the correlation between your final preference labels and length difference. If it is high, you have already decided your model will be verbose.
Define ties and cap the tie rate. A tie should mean "neither is clearly better on the rubric," not "I can't be bothered." Track tie rate per annotator and overall. A sudden rise in ties usually means the annotators are fatigued or the pairs are too easy.
Include hard pairs, not just easy ones. If every pair is obviously different, the reward model learns a coarse boundary and fails near the margin—exactly where production model outputs cluster. Sample a deliberate share of near-ties: responses of similar quality where the rubric decides.
Add trap and gold pairs. Seed known pairs with a documented correct answer, including pairs that are genuinely close. These catch annotators who are rubber-stamping, and they give you a continuous quality signal without a separate review pass.
Collect the rationale, and use it. "A is better because B invents a statistic" is far more useful than the label alone—it lets you group failure reasons, and it lets you detect when annotators are optimising for something you didn't intend. It also improves the annotator's own consistency.
The pairing strategy matters more than the volume
A million comparisons collected from one prompt distribution teaches a reward model about that distribution and little else. Coverage should come first:
- Model-output diversity. Pairs from a single model version overfit the reward model to that model's style. Include outputs from multiple models and from human writers.
- Prompt diversity. Instructions, questions, coding tasks, refusals, edge cases, multilingual, adversarial. If production will see code, your preference data cannot be all prose.
- Difficulty spread. Easy pairs establish the skeleton; hard pairs decide the boundary. You need both, in a deliberate ratio.
- Failure-mode coverage. Deliberately include pairs where the failure is subtle: a plausible hallucination, a right answer buried in a wrong format, a safe refusal that is unhelpful.
Evaluate the reward model against humans, not against itself
The loop that catches most problems: hold out a human-labelled preference set, and measure how often the reward model agrees with the humans. That number—not training loss—is your reward model's quality.
Then go further:
- Adversarial pairs. Construct pairs where the worse answer is longer, more confident, or better formatted. A reward model that gets these wrong has learned the surface cues.
- Preference distribution across axes. Measure whether the reward model tracks correctness, safety, style, and format separately or has collapsed them into one fuzzy axis.
- Disagreement analysis. Where the reward model and humans diverge, read the examples. Divergence is usually informative, not just an error rate.
A caution about annotator-level analysis: stripping individual annotators out changes what you're measuring. In small datasets, one prolific annotator can be most of your signal. Track per-annotator contribution and agreement, and if you weight, do it deliberately and document it.
Symptoms that your policy is being optimised against a broken reward model
Watch for these in the trained model, not just the reward model:
- Answers grew longer and less specific
- Unconditional agreement with the user's premise, including when the premise is false
- Format compliance at the cost of content
- Playbook phrases that appear in your annotation examples
- Sharp drops in performance on the slices your preference data barely covered
The last one is the tell. Reward hacking usually looks like competence in the region you measured and incompetence exactly where your data was thin.
A practical order of operations
- Write the rubric and define the axes explicitly, including the priority between them
- Run a calibration round and measure inter-annotator agreement before scaling
- Randomise positions, include reversed pairs, and measure the flip rate
- Define and cap ties; track per-annotator rate
- Balance length; report the label-length correlation
- Guarantee coverage across models, prompt types, difficulty, and failure modes
- Hold out a human preference set and evaluate the reward model against it
- Watch the policy for the classic hacking signatures on every iteration
Preference data is a measurement instrument. Treat collection the way you would treat a survey with a real sampling design, and the reward model becomes a faithful proxy for your users. Treat it as a sorting exercise, and you have built a machine for amplifying your annotators.
We design and run RLHF and preference-data programmes—rubrics, calibration, coverage strategy, and reward-model evaluation—alongside AI evaluation and red teaming for the models they shape.


