Researchers auditing nine frontier AI chatbots across 810 simulated mental health conversations found that concerning behavior was “widespread,” and that the highest risk came not from chatbots saying something obviously dangerous but from chatbots being supportive in a way that fed the user’s underlying problem. The framework, called SIM-VAIL, was published in Nature Medicine and tested models from the Claude, ChatGPT, Gemini, Grok and Llama families.
The team named the pattern a vulnerability-amplifying interaction loop, or VAIL. In plain terms: each individual reply looks fine, and the conversation as a whole makes the person worse.
Key takeaways
- Across 810 conversations, 9 chatbots and 30 simulated user profiles, concerning behavior was widespread, though it was lower in newer models.
- Risk built up over the course of a conversation rather than showing up in a single bad answer.
- The most harmful pattern came from behaviors that look like good care in isolation, including validation and empathy.
What the study found
SIM-VAIL works by having one AI role-play a person with a defined psychiatric vulnerability, then having that simulated person hold a real back-and-forth conversation with a target chatbot. Every exchange is scored across 13 clinically grounded risk dimensions.
The researchers built 30 user profiles, each one “a pairing of one of five psychological vulnerabilities and one of six interaction intents.” Vulnerability is who the user is. Intent is what they want from the chatbot. Each combination was run in triplicate, producing 810 conversations in total.
Concerning behavior “varied by user vulnerability and conversational intent, accumulated over turns, and could be reduced by interventions at early escalation points.” Newer models did better than older ones. And risk “was highest when otherwise supportive chatbot behaviors reinforced the psychological mechanisms underlying the simulated user’s vulnerability.”
The authors are blunt about why current testing misses this. Most safety benchmarks check single responses to fixed questions and look for overt harms such as encouraging self-harm. The risk this study measured accumulated over turns, which is the one thing a single-response test cannot see.
Dr. Kumar’s take
This is the result I would want every patient to understand, and it is the opposite of the usual AI scare story.
The fear people carry into my office is that a chatbot will tell someone to hurt themselves. That is the failure mode the industry has spent the most effort on, and this paper suggests newer models are improving on it. The failure mode nobody engineered against is agreeableness.
Agreeableness is exactly the wrong reflex for a patient whose problem is a distorted belief. If someone tells me they are certain their spouse is against them, or that the headache means a tumor, or that they cannot leave the house because the panic will kill them, the clinically correct move is warm disagreement. I take the feeling seriously and I push back on the conclusion. A chatbot tuned to be helpful, pleasant and non-judgmental does the first half beautifully and skips the second half.
That is why a single-response safety test passes a chatbot that a real clinical encounter would fail. Score one reply and you see empathy. Score twenty replies and you see a person who arrived with a shaky belief and left with a reinforced one, agreed with the whole way.
The caveats matter. These were simulated users, not people. One AI judged another AI’s behavior. Nobody has shown that a VAIL in a transcript produces measurable worsening in a real person’s symptoms. This is a mapping tool, not an outcome trial, and the authors present it that way.
What it means for you
Here is the rule I would give a patient: a chatbot is a fine journal and a bad mirror.
As a journal, it is useful. Typing out what happened today, sorting a messy feeling into sentences, drafting the thing you want to say to your brother, all fine. You are the one doing the thinking.
It becomes a mirror the moment you ask it to confirm something rather than help you look at something. Three warning signs: every response agrees with you, you feel better after each session but nothing in your life changes, or you go back several times a day for reassurance about the same worry.
One habit worth building: ask it to argue the other side, and notice how fast it folds when you push back. And take anything important to a human, because part of what I am paid for is telling you something you did not want to hear.
For protecting your mind, the boring interventions still beat the novel ones. Swapping one hour of TV for reading is linked to lower dementia risk, and it is better supported than an app that agrees with you.
