OpenAI has released MentalHealthBench, an open benchmark designed to measure how AI systems respond in realistic mental health conversations. It was developed with 80-plus licensed mental health experts spanning 22 countries. The benchmark evaluates model capabilities across safety, context-seeking, user agency preservation, and actionable guidance.
The release lands as more than one billion people use ChatGPT each week, many of them turning to it for conversations about relationships, stress, caregiving, and difficult life situations. OpenAI argues that existing evaluations have leaned mostly on emergency scenarios and broad criteria, leaving the broader range of mental health interactions largely unmeasured.
What the benchmark covers
MentalHealthBench uses privacy-preserving techniques to build synthetic conversations that mirror real-world usage patterns. Scenarios span three tiers: non-acute everyday conversations, high-acuity situations that signal serious distress, and emergencies where immediate real-world support is needed. Personas cover adults, teens aged 13-17, caregivers, and clinicians, across multiple languages and regions.
At least three experts reviewed each conversation and wrote weighted rubric criteria, scored from -10 to +10. A criterion made the final rubric only if at least two experts endorsed it and no third expert disagreed. An automated grader — GPT-5.6 Sol — then assesses model responses against those expert-written criteria.
Model performance results
OpenAI evaluated a wide range of models on the benchmark. Results show steady improvement in how well AI systems handle mental health conversations, with newer frontier models scoring higher overall. Those overall scores break down into ten dimensions of model behavior defined by mental health experts. Models with similar totals can still differ in where they are strong — context-seeking, for instance, improves with more advanced models.
OpenAI stresses that ChatGPT cannot stand in for therapy or a licensed professional. Findings drawn from the expert panel help track progress toward models that respond with empathy, support well-being, and point people toward real-world help, such as local crisis hotlines and trusted contacts.
User perspectives vs expert guidance
Alongside the benchmark, OpenAI ran a separate analysis comparing expert guidance with what users actually find helpful. The company worked with 44 adults across 16 countries who had used AI for mental health support. Users valued practical next steps and tone — qualities that expert guidance emphasizes less. Experts, in turn, placed more weight on gathering context and reading ambiguous situations carefully.
Confirmed
- MentalHealthBench released as an open benchmark for evaluating AI in mental health conversations, available for community examination and extension
- Co-created with 80+ licensed mental health experts across 22 countries, 19 languages, and nearly 20 subspecialties
- Scenarios cover non-acute, high-acuity, and emergency conversations across adult, teen (ages 13-17), caregiver, and clinician personas, with rubric criteria weighted -10 to +10 and at least three experts reviewing each conversation
- Automated grading runs through GPT-5.6 Sol against expert-written criteria; overall scores decompose into 10 behavioral dimensions
- Separate user study (44 adults, 16 countries) found users prioritize practical steps and tone, while experts emphasize context-gathering
Unknown
- Specific model names and numerical scores from the evaluation (charts referenced but values not provided in text)
- Exact GA date or access method for the benchmark dataset and grading code
- Whether the teen persona system message approach captures all provider-specific safeguards
- Independent replication status of the automated grading methodology
Our take
MentalHealthBench addresses a real evaluation gap by moving beyond emergency-only testing into the messy middle of everyday mental health conversations. The expert consensus rubric and multi-dimensional scoring are thoughtful design choices. What matters next is whether the open release enables meaningful third-party replication — and whether the benchmark's synthetic conversations truly capture the cultural and linguistic nuance of real help-seeking behavior.