How to Evaluate a Post-Trained Model Without Fooling Yourself
A layered evaluation system for post-trained models that controls prompt effects, judge bias, contamination, and the gap between benchmark gains and product quality.
A post-training run produces a new checkpoint and an immediate temptation: put two scores next to each other and declare progress. The danger is that language-model evaluation is not a thermometer. It is a measurement system assembled from prompts, chat templates, sampling settings, parsers, judges, tools, and datasets. Change any one of them and the score can move even when the underlying model has not.
The goal is not to find one perfect benchmark. It is to build an evaluation process that makes self-deception expensive.
Freeze the inference contract first
Before comparing models, write down exactly how they are being run. That contract should include the system prompt, chat template, tool definitions, decoding settings, token budget, stop conditions, and answer parser. For an agent, it must also include the harness: how context is compacted, how tools report errors, what files are available, and when the loop terminates.
Prompt formatting is part of the test, not harmless packaging. A model trained to end a math answer with a particular marker may look weak when the evaluator expects a different format. A coding model with too small a token budget may fail because it cannot finish. If model A receives a carefully tuned prompt while model B receives a generic one, the comparison is about integration quality as much as model quality.
There are two honest ways to proceed. Use an identical, neutral contract to compare raw compatibility, or optimize the contract for each model to compare the best product you can build with it. Both are useful. Mixing them is not.
Measure four different layers
I separate evaluation into four layers because one aggregate score hides the tradeoffs post-training creates.
Capability
Can the model perform the underlying task? Use questions with verifiable outcomes where possible: executable code, retrieved facts with evidence, calculations, structured transformations, or tasks graded against a clear rubric. Capability tests should cover both the target domain and unrelated skills that might regress.
Behavior
Does the model follow instructions in the way users need? Measure format adherence, appropriate concision, calibrated uncertainty, citation behavior, and recovery after correction. Preference tuning often changes these traits without changing factual capability. A model can know the answer and still be a poor assistant.
Safety
Test both unsafe compliance and unnecessary refusal. A model that refuses every ambiguous request may score well on a narrow safety set while failing legitimate users. Include adversarial prompts, benign prompts containing alarming vocabulary, multi-turn escalation, and tests of whether the model preserves boundaries when tools are available.
Product outcomes
Finally, measure the deployed system: task completion, latency, inference cost, retries, user edits, escalation to humans, and abandonment. These metrics are noisy and context-dependent, but they are the closest representation of value. A benchmark improvement that makes answers twice as long and materially slower may be a product regression.
Do not force these layers into one number too early. A dashboard that exposes tradeoffs is more useful than a composite score whose weights quietly encode the conclusion.
Treat the judge as another model under test
LLM judges are useful because they make open-ended evaluation scalable. They are also susceptible to presentation effects. A judge may prefer the first answer, the longer answer, the more polished style, or language resembling its own. A detailed explanation can win over a shorter correct response even when brevity is part of the requirement.
Reduce these risks with procedure. Randomize answer order and score both orderings. Hide model names. Give the judge a task-specific rubric with separable criteria rather than a request to choose the response with the best overall “quality.” Use deterministic checks whenever the outcome is objectively testable. On a stratified sample, compare judge decisions with expert human ratings and examine disagreements rather than reporting correlation alone.
For important decisions, use more than one measurement channel. A judge preference, a factuality checker, and a human reviewer can fail differently. Agreement is useful evidence; disagreement is a map of where the evaluation is fragile.
Assume contamination is possible
Public benchmarks are public data. Their questions, solutions, paraphrases, and discussion pages can enter pretraining, instruction data, preference data, or synthetic-data pipelines. A high score therefore does not prove generalization.
Maintain a provenance record for every training source and evaluation set. Search for exact and approximate overlap before training, including long shared substrings and prompt templates. Keep a genuinely held-out set with restricted access, add newly written items after the training cutoff, and create perturbation tests that preserve the skill while changing names, numbers, framing, or output format.
Perturbation is diagnostic, not a courtroom test. A model may be sensitive to format for reasons other than memorization. The useful question is whether performance survives reasonable changes that a real user would make.
Build an evaluation flywheel
The best evaluation suite grows from observed failures. I use a simple loop:
- Define the behavior the training run is intended to change.
- Collect representative failures and assign them to a clear taxonomy.
- Reproduce each failure under a frozen inference contract.
- Add a small development test and a separate held-out test.
- Train, then run the full suite—not only the metric targeted by training.
- Inspect regressions and judge disagreements manually.
- Shadow-test or A/B-test the candidate in the product before promotion.
- Feed new production failures back into the next evaluation cycle.
Keep two kinds of tests: a stable core for tracking progress over time and a rotating frontier that is harder to hill-climb. The stable set gives comparability. The rotating set reduces the chance that the team merely learns the benchmark.
Consider a coding assistant trained to write more complete patches. Its code benchmark may improve, yet product telemetry may show more reverted changes. Investigation could reveal that the model edits unrelated files. The next evaluation should not be “more coding questions”; it should include repository-scoped tasks with checks for unnecessary diffs. Evaluation becomes valuable when it turns vague disappointment into a reproducible training signal.
The final discipline is to preserve uncertainty. Repeat sampled evaluations, report variation, and treat small score differences cautiously. Keep examples next to aggregates. A result you cannot explain is not yet an insight.
Post-training is powerful precisely because it can optimize whatever signal we provide. Evaluation must therefore do more than celebrate movement. It must tell us whether we moved the model, the metric, or merely the way we asked the question.