All posts

Human-in-the-Loop

Putting Humans in the Live Loop: Designing Review That Fits Inside a Real SLA

Reviewing training data offline is one job. Putting a person between an AI's prediction and a real-world action is another. Here is how to design live human review that improves outcomes without becoming the bottleneck it was meant to remove.

A four-stage production pipeline from model output through a confidence gate and human review to action, with a feedback loop labelled every correction becomes training data.

Most teams meet human-in-the-loop during training. Someone reviews annotations, QA catches errors, the dataset improves. That is a batch process, and it is well understood.

A different problem shows up later: the model is in production, and some share of its outputs cannot be allowed to execute unchecked. A claim gets approved or denied. A transaction gets blocked. A clinical summary gets filed into a patient record. A support agent sends a response in your brand's voice.

Now you need a person in the loop at inference time, and the design constraints are nothing like batch review. You have a latency budget, real users waiting, a queue that can back up, and a reviewer who gets tired. Get it wrong and you have built the bottleneck you were trying to avoid.

Offline review and live review are different systems

Worth stating plainly, because teams conflate them and then wonder why the live version underperforms.

Dataset / QA reviewLive production review
TriggerBatch schedulePer prediction, in the request path
Latency budgetHours to daysMilliseconds to minutes
Throughput neededWhatever the batch needsMatches live traffic
Cost of a missDoesn't shipReal consequence, now
Failure modeSilent data debtBlocked users, SLA breach

Live review is an operational system. It needs routing logic, queues, staffing models, SLAs, and monitoring—the things you would expect from a support desk, not from a labelling pipeline.

Rule 1: Decide per action, not per model

The first design question is not "how do we review outputs?" It is "which actions genuinely need a human before they happen?" That depends on three things per action type:

  • Reversibility. Can this be undone? A draft can; a sent email, a denied claim, or a filed record usually can't.
  • Cost of error. A mislabelled internal priority is cheap. A wrong medication or an incorrectly blocked payment is not.
  • Confidence and novelty. Is the model's output high-confidence and within the distribution it was trained on?

The useful output of this exercise is a routing policy per action type, not a single global setting. Most production systems end up with three lanes:

  1. Auto-execute. High confidence, in-distribution, low consequence. No human.
  2. Human confirm. Medium confidence, or a high-consequence action. A human approves, edits, or rejects.
  3. Human decide. Low confidence, novel input, or an action the policy marks as never-automatic. The model's output becomes a suggestion to a person making the call.

Rule 2: Route on calibrated confidence, not raw scores

If you route on the model's own confidence score, the routing is only as good as that score's calibration. Raw scores from a classifier or a generation model are not probabilities. Calibrate them against a labelled sample first, so that "0.9" means roughly 90% correct for that action type—then set thresholds that reflect the cost of error.

A pattern that works well in practice: set the review threshold per action type so that the expected cost of an unreviewed error equals the cost of reviewing one more item. That sounds abstract, but it forces the right conversation—with the business owner—instead of a developer picking 0.8 because it felt reasonable.

Also route on signals that aren't confidence:

  • Input novelty: is this far from the training distribution?
  • Business rules: amount over a limit, a new customer, a flagged account
  • Model disagreement: two model versions or two prompts disagree
  • Temporal: first N days after a model change, or during an incident

Rule 3: Make the reviewer fast, or the queue becomes the product

Reviewer throughput is not a soft concern; it is the capacity of your system. A few design choices dominate:

  • Put everything in one place. The input, the model's output, the evidence it used, and the decision buttons. A reviewer who has to open three tabs and scroll a PDF costs multiples per item.
  • Show the model's reasoning and its sources, not just its answer. Reviewers who can see why catch more errors and decide faster.
  • Pre-fill the edit. When a reviewer corrects an output, give them the model's version as an editable starting point, not a blank box.
  • Keyboard-first. Review is a high-volume repetitive task. Every mouse movement compounds.
  • Batch by type. Reviewers working on a single action type make fewer errors than those context-switching every item.
  • Right-size the reviewer skill. A high-consequence clinical decision needs a clinician. A formatting confirmation does not. Mismatched skill level is expensive in both directions.

Rule 4: Define the SLA and staff for the tail, not the mean

Live review SLAs are usually expressed as a percentile, not an average. "95% of items reviewed within 5 minutes" is a design constraint. Design for it:

  • Queue depth alarms, not just latency alarms—depth tells you about the future
  • Escalation tiers with different SLAs (tier 1 handles routine, tier 2 handles ambiguous, tier 3 gets domain experts)
  • Overflow behaviour that is written down: when the queue is full, does the system hold, degrade to human-decide, or fall back to a safe default?
  • Coverage across time zones if the product is global. A review desk that sleeps is not an SLA.

That last point matters more than teams expect. If you are serving users worldwide and your reviewers are in one time zone, your overnight SLA is really "the model decides alone." Either staff it or write the policy honestly.

Rule 5: Instrument the humans, not just the model

The failure mode of live review is rubber-stamping: reviewers under time pressure approve whatever they see. Approval rates creeping toward 99% is a warning sign, not a success metric.

Measure:

  • Override rate (how often humans change the model's output) and its trend
  • Agreement between reviewers on overlapping items—the same calibration logic as annotation QA
  • Time per item, and its distribution. Suspiciously fast items deserve a look.
  • Seeded gold items inserted into the queue at a low rate, with known correct answers
  • Post-hoc audits: sample auto-executed items and check them later, to measure what your routing let through

Without the audit of the auto-execute lane, you have no idea what your thresholds are actually costing you. That lane is where the risk accumulates silently.

Rule 6: Close the loop back into training

Every review is a labelled example: the input, the model's prediction, the human's decision, the confidence, and the reason for the override. That is the highest-quality supervision you will ever get, because it comes from the exact distribution your model faces and from the people accountable for the outcome.

Feed it back deliberately:

  • Override clusters reveal a systematic model weakness → new training data or a prompt fix
  • A specific input type that always needs review → either improve the model there or accept a permanent review lane and budget for it
  • Repeated escalations to tier 3 → your tier-1 guideline has a gap

Systems that capture this become better and cheaper over time. Systems that don't keep paying the same review cost forever.

A short design checklist

  • A routing policy defined per action type, agreed with the business owner
  • Confidence calibrated on labelled data before thresholds are set
  • Thresholds derived from the cost of error, reviewed periodically
  • Reviewer interface optimised for speed, with evidence and pre-filled edits
  • SLA expressed as a percentile, with queue-depth alerting and a written overflow policy
  • Time-zone coverage that matches where your users are, or an honest policy if it doesn't
  • Human-side metrics: override rate, inter-reviewer agreement, seeded gold items
  • An audit of the auto-execute lane
  • A feedback path from every correction back into training

Putting humans in the live loop is not a downgrade of an AI system and it is not a temporary scaffold you remove later. For consequential decisions it is the control that makes automation defensible. Design it as an operational system—routing, queue, SLA, audit, feedback—and it earns its cost. Bolt it on as an afterthought and it becomes a queue that users learn to hate.

Our production human-in-the-loop platform lets teams insert expert review into live AI workflows with configurable routing, audit logs, and a feedback path into training—delivered inside an ISO 27001 certified environment. For dataset-side review, see our HITL practice.