Cookies on this website

We use cookies to ensure that we give you the best experience on our website. If you click 'Accept all cookies' we'll assume that you are happy to receive all cookies and you won't see this message again. If you click 'Reject all non-essential cookies' only necessary cookies providing core functionality such as security, network management, and accessibility will be enabled. Click 'Find out more' for information on how to change your cookie settings.

Scientists have developed a framework for stress-testing how AI chatbots respond to vulnerable users, helping researchers identify weaknesses and test ways of making mental-health interactions safer.

A person's hands can be seen typing on a keyboard with an AI chatbot screen superimposed over the top © Shutterstock

Published in Nature Medicine, the study introduces SIM-VAIL, a clinically validated framework for auditing AI chatbots in mental-health conversations. The work was led by researchers at the University of Oxford, University College London (UCL), and the UK AI Security Institute.

SIM-VAIL simulates users with specific psychological vulnerabilities, such as depression, mania, psychosis, obsessive-compulsive disorder, or insecure attachment, and with a wide range of intentions, for example trying to get the chatbot to agree with them, downplay their difficulties or endorse risky actions. It then engages chatbots in multi-turn conversations and scores each exchange across clinically grounded risk dimensions.

The researchers used the framework to investigate 810 conversations with nine frontier AI models (including Claude, ChatGPT, Gemini, Grok and Llama models), spanning 30 simulated user profiles and more than 90,000 clinical ratings.

They found that concerning behaviour in target chatbots was widespread, although this was significantly reduced in newer models. Safety depended strongly on the user’s psychological context and how the conversation developed.

Risks often emerged gradually when apparently supportive responses inadvertently reinforced the psychological processes underlying a user’s vulnerability. The researchers termed this pattern a “Vulnerability-Amplifying Interaction Loop”, or VAIL.

The study also identified a route towards improvement: replacing a single concerning response early in an interaction led to safer subsequent exchanges. This suggests that targeted interventions at early points of escalation could improve the trajectory of an entire conversation.

SIM-VAIL’s automated assessments showed substantial agreement with clinicians evaluating the same interactions, supporting its use as a scalable tool for identifying conversational weaknesses and testing new safeguards.

Senior author Dr Matthew Nour, Senior Clinical Researcher at the University of Oxford, said:

We developed SIM-VAIL to address a gap in the way chatbots are currently assessed for mental health risks. Many existing benchmarks assess how a chatbot responds to a single message, whereas important risks may emerge gradually over the course of a conversation. SIM-VAIL allows us to examine these multi-turn trajectories systematically and at scale, and provides a foundation for targeted, context-sensitive safety improvements.”

Rather than labelling a chatbot as simply “safe” or “unsafe” in mental-health contexts, SIM-VAIL provides a platform for identifying specific weaknesses and testing whether new safeguards address them.

The researchers have also released the open SIM-VAIL Explorer, which allows researchers, AI developers and the public to examine the study’s conversations, follow how risks developed turn by turn, and compare results across models, psychological vulnerabilities and conversational goals.

Lead author Dr Veith Weilnhammer, Fellow at the Max Planck UCL Centre for Computational Psychiatry and Ageing Research, said:

Millions of people already use general-purpose AI chatbots to discuss emotional and mental-health concerns. Our findings show why safety cannot be captured by a single overall score: it depends on who the user is and how the conversation develops. By making the framework, data and Explorer openly available, we hope to help researchers and developers identify specific weaknesses and measure whether new systems are genuinely improving.”

Because SIM-VAIL deliberately stress-tests chatbots under adversarial conditions, the findings should not be interpreted as estimates of how frequently concerning interactions occur in ordinary use. Instead, the framework offers a controlled and scalable way to uncover risks and test approaches for addressing them.