Helpful Is Not the Same as Safe: What MentalHealthBench Reveals About Evaluation Design
MentalHealthBench shows why high-stakes conversational evaluation needs expert guardrails, user perspectives, behavioral decomposition, and restraint about leaderboards.
Evaluating AI in a mental-health conversation is not a matter of checking one correct answer. A response can be accurate but poorly calibrated to urgency, empathetic while reinforcing an unsupported belief, or safe in form yet so cautious that ordinary distress feels like an emergency.
MentalHealthBench treats these tensions as the object of measurement. It does not establish that a model is a therapist—the authors explicitly say ChatGPT is not a substitute for professional care—but offers an expert-informed diagnostic across realistic situations.
A benchmark built around contextual rubrics
MentalHealthBench contains 1,215 synthetic conversation prefixes and 5,262 expert-authored rubric criteria. The scenarios range from everyday well-being to high-acuity situations and emergencies. They include adults, teenagers, caregivers, and clinicians across multiple languages and cultural contexts.
More than 80 licensed psychologists and psychiatrists from 22 countries contributed. They collectively speak 19 languages and represent nearly 20 mental-health subspecialties. Each conversation was reviewed by at least three experts. A criterion was retained only when at least two agreed and a third did not contradict it. Criteria receive signed weights from -10 to +10, allowing the benchmark to reward a desirable behavior or penalize a harmful one according to its importance in that specific context.
The dataset is deliberately broad: 53.5% of examples are non-acute, 18.2% are high-acuity, and 28.3% are emergencies. Adults account for 68.1% of the user profiles, teenagers 21.2%, clinicians 5.8%, and caregivers 4.9%. This composition should not be read as the prevalence of these topics in real ChatGPT use; it is an evaluation mixture designed to expose different failure modes.
The official results show meaningful but incomplete progress. On the reported task-clipped score, GPT-6 Astra reaches 57.3%, GPT-6 Sol 53.9%, and Claude Opus 5.5 52.4%. Even the leading score leaves substantial space between current behavior and the full expert rubric. More importantly, models with similar totals can reach them through different combinations of context seeking, actionable guidance, clinical accuracy, empathy, reality testing, urgency calibration, harm avoidance, agency, and communication.
That decomposition is more informative than a single ranking. In a high-stakes domain, the location of an error matters. Failing to ask for necessary context is different from offering an impractical next step, and both are different from missing an emergency.
Users and experts provide different signals
The project separately asks 44 adults who had used AI for emotional or mental-health support to evaluate non-acute examples. Agreement within the expert group is 63.4%, agreement within the user group is 62.0%, and expert-user agreement is 51.5%.
When the researchers compare independently written rubrics, only 25.7% of their absolute weight is aligned. Another 39.1% is expert-only and 34.2% user-only, while just 1.0% is directly contradictory. This is a subtle result. Users and clinicians are usually not issuing opposite instructions; they are noticing different dimensions. Users emphasize tone, interpretation, and practical next steps. Experts more often emphasize gathering context, preserving uncertainty, and avoiding unsafe assumptions.
My interpretation is that “helpful” and “safe” are neither identical nor simple tradeoffs. They are partially overlapping objectives with different blind spots. Optimizing only for user approval risks losing clinical guardrails. Optimizing only for expert criteria risks producing a technically cautious response that does not help someone decide what to do next. A well-designed post-training system needs both signals, with an explicit policy for resolving disagreement rather than averaging everything into one reward.
The surprising lesson from expert answers
One result deserves special attention. Responses generated with the rubric visible score 99.0%, while expert-written responses score only 38.5%. The paper argues that clinicians often respond briefly, as they might in a live conversation: one careful question or one restrained observation. The rubric-aware model produces a more exhaustive answer that covers many scorable criteria.
This does not make models better clinicians. It exposes a familiar artifact: completeness can earn points even when professional practice values pacing. A response may optimize one-turn coverage while undermining a relationship that unfolds over time.
This is why MentalHealthBench should be used as the authors describe it: an auditable diagnostic, not a definitive leaderboard. The paper’s behavioral axes and signed penalties can reveal where a system fails. The aggregate score alone cannot establish that the system behaves appropriately across an ongoing conversation.
Important limits
The conversations are synthetic, even though their themes are derived using privacy-preserving analysis of real usage patterns. The benchmark supplies a multi-turn prefix but evaluates the next model response; it is not a true adaptive rollout in which both participants continue for many turns. User input was collected only for non-acute cases for ethical reasons. Multilingual coverage is descriptive, and differences in language cannot be separated cleanly from topic, acuity, culture, or persona.
The benchmark also relies on GPT-5.6 Sol as an automated grader. Teenage users’ ages are supplied explicitly, which may not reproduce product-level detection and safeguards. The authors also ask people not to publish examples, reducing contamination risk while acknowledging that open benchmarks can become training data.
The broader lesson reaches beyond mental health. High-stakes evaluation needs multiple constituencies, behavior-specific rubrics, negative criteria, and explicit analysis of tensions. It also needs humility about what the measurement rewards. The goal is not a model that says the maximum number of acceptable things. It is a model that responds with the right degree of care, context, usefulness, and restraint for the person and moment in front of it.