OpenAI released MentalHealthBench, an open benchmark that scores how AI models handle mental health conversations ranging from everyday stress to emergencies, developed with over 80 licensed psychologists and psychiatrists from 22 countries. The experts wrote weighted criteria for each synthetic conversation, rewarding responses like asking what help the user wants and penalizing ones like guessing at their feelings, and OpenAI’s own GPT-5.6 Sol grades answers against them. OpenAI’s GPT-6 Astra scored highest at 57.3%, ahead of GPT-6 Sol at 53.9% and Anthropic’s Claude Opus 5.5 at 52.4%, while Google’s Gemini 3.1 Pro tied the older GPT-4o at 32.1%. In a separate test, 44 people who use AI for emotional support cared more about tone and practical next steps than the experts did.






