
People do not always announce that they are having a mental health conversation with a chatbot. A question about a difficult friend can slowly turn into a discussion of isolation, grief or immediate danger. An AI system has to notice that change and respond with more judgment than a generic reassuring sentence.
That is the problem OpenAI is trying to measure with MentalHealthBench, an open evaluation introduced on September 23. More than 80 licensed mental health professionals from 22 countries helped develop it. Their task was to judge what a helpful response should do in realistic but synthetic conversations, from everyday stress to emergencies.
The benchmark matters because many existing safety tests focus on extreme cases. Those cases are essential, but millions of conversations sit in the middle. A chatbot may not need to issue a crisis warning, yet it still needs to avoid overconfident advice, ask for context and leave the person in charge of their own decisions.
OpenAI says the scenarios cover adults, teenagers, caregivers and clinicians across languages and regions. Experts set criteria for each conversation, rewarding helpful behavior and penalizing responses that could do harm. At least three experts reviewed each synthetic conversation, with the final criteria reflecting areas of agreement.
The examples show how subtle this can be. In a conversation about a strained friendship, a useful answer might acknowledge the person’s uncertainty and ask what outcome they want. An answer that claims to know the friend’s motives or tells the user what they have already decided would score poorly. A model can sound warm and still take too much authority over someone else’s life.
OpenAI also asked 44 people who had used AI for emotional support to evaluate non-emergency examples. Those users emphasized practical next steps and tone, while clinicians placed more weight on gathering relevant context and handling ambiguity. That difference is important. A clinically cautious answer may feel cold; a pleasing answer may skip the questions that make it safe.
This work arrives as chatbot companies face closer scrutiny over the emotional role their products play. TechBooky has examined the stronger safeguards OpenAI introduced for teenagers. A common, expert-informed test could help show whether those safeguards work across the many situations people actually bring to AI, rather than only in a handful of scripted demonstrations.
There are limits. The benchmark uses synthetic conversations, and a score is not evidence that a chatbot will respond well to every real person in distress. OpenAI’s own model is used as an automated grader against expert-written criteria, a methodological choice independent researchers should be able to examine now that the work is open.
Nor does a better score make ChatGPT a therapist. OpenAI says it is not a substitute for therapy or professional care. The clearest value of MentalHealthBench may be more modest: it gives developers a way to find weak responses before a vulnerable user encounters them. The human support available outside the chat window remains the part no benchmark can supply.







