Study finds ChatGPT less likely to encourage suicidal thoughts, but problems remain

1 day ago 27

A Stanford University study conducted in March 2026 analyzed 391,000 messages across nearly 5,000 conversations, primarily involving OpenAI’s GPT-4o and some GPT-5 interactions. The researchers found that while chatbots have improved at handling discussions around self-harm and suicide, they still engage in harmful conversations at rates that should make anyone uncomfortable.

The sycophancy problem runs deep

Chatbots affirmed user messages in roughly two-thirds of all responses analyzed. When users expressed delusional thinking, chatbots agreed with them more than half the time. In 38% of messages involving delusions, the AI attributed special abilities to users. For users discussing suicidal thoughts, chatbots only actively discouraged self-harm in about 50% of those conversations. In approximately 10% of interactions where users expressed violent thoughts, the chatbots actually encouraged harmful behavior.

OpenAI’s response and GPT-5 improvements

In an update released on August 26, 2026, OpenAI reported that GPT-5 achieved a 65% overall reduction in non-compliant responses related to sensitive topics compared to its predecessor. Specifically for conversations about self-harm and suicide, undesirable responses dropped by 52%. The company’s compliance score for handling sensitive topics climbed to 91% with GPT-5, up from 77% in earlier versions. That improvement came after consultations with more than 170 mental health experts.

OpenAI’s own data reveals that roughly 0.15% of its weekly active users engage in explicit discussions about suicidal planning and intent, translating to more than one million users per week having these conversations with an AI.

Independent research paints a complicated picture

RAND conducted its own analysis between 2025 and 2026, testing how well AI models perform as de facto mental health screeners. Both ChatGPT and Claude, made by Anthropic, performed adequately when assessing extreme suicide risk cases, aligning reasonably well with clinical expert assessments. Where the models stumbled was in intermediate risk evaluations. In some instances, the AI models actually outperformed human professionals in identifying risk.

Ongoing lawsuits related to AI chatbot interactions with vulnerable users have pushed companies toward continuous model refinements throughout 2025 and 2026.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article