TECHNIQUE
Evaluation
Online evaluation is observed as live/operational scoring, A/B-test readouts, and human-alignment checks rather than a single standard pattern.
Use operational signals in evaluation: sampled production traffic, A/B-test rollout traffic, or recent peer-review outcomes.
3 of 3 operators with online-evaluation evidence in this pool.Continuously sample live production traffic and score it with the same metrics and logic used in offline suites.
1 of 3 operators with online-evaluation evidence in this pool.Combine online production scoring with pre-merge and staging evaluation gates.
1 of 3 operators with online-evaluation evidence in this pool.A/B test new models during rollout and report advertiser outcome uplift from phased deployment.
1 of 3 operators with online-evaluation evidence in this pool.Validate LLM workflow outputs against human analyst conclusions during peer review over a recent operating window.
1 of 3 operators with online-evaluation evidence in this pool.Every cited operator evaluates against real operating behavior, but the measured signal differs by product: production traffic scores, A/B-test traffic, or human peer-review alignment.
What online signal is used as the evaluation unit.
APPROACH 01
Sampled live production traffic is scored with the same metrics and logic as offline suites.
APPROACH 02
New model behavior is evaluated through A/B testing and phased rollout impact on ROAS.
APPROACH 03
LLM decisions are validated against human analyst conclusions during peer review.
Who or what judges online behavior.
APPROACH 01
Automated scorers and judge-model style checks are used for answer quality, citation support, formatting, and tone.
APPROACH 02
Business-performance readouts from A/B testing and phased rollout are used.
APPROACH 03
Human analyst conclusions in peer review are the comparison point.
Live or changed traffic can expose quality/calibration failures: Dropbox reports pipeline tweaks can turn a prior good answer into a hallucination, while Criteo reports A/B testing new models can acquire different traffic on which the models are not calibrated.
For insufficient-context LLM cases, Agoda reports over-escalating and routing to human review rather than relying on the model verdict.
| Name | Kind | When | Maturity |
|---|---|---|---|
| A/B testing with guardrail metrics | pattern | model or prompt changes ship behind controlled rollouts | established |
| Implicit feedback capture | pattern | thumbs, regenerations, and abandonment proxy quality at zero annotation cost | commodity |
| Langfuse online scoring | library | production traces scored continuously against eval criteria | emerging |