AI工具Score B (61)
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
1 小时前2 viewsSource: nature.com
Download PDF Abstract Millions of users turn to consumer artificial intelligence chatbots to discuss emotional, behavioral and mental-health concerns, creating an urgent need for rigorous and scalable safety evaluations. Here we introduce simulated (SIM) vulnerability-amplifying interaction loops (VAILs) (SIM-VAIL), a clinically validated framework for auditing chatbot behavior in mental-health contexts. SIM-VAIL simulates users with specific psychiatric vulnerabilities and conversational intents, engages them in multi-turn conversations with frontier artificial intelligence chatbots (including Claude, ChatGPT, Gemini, Grok and Llama models) and scores each exchange across 13 clinically grounded risk dimensions. Across 810 conversations, spanning 9 target chatbots and 30 simulated user profiles, concerning behavior in target chatbots was widespread, albeit reduced in newer models. Concerning behavior varied by user vulnerability and conversational intent, accumulated over turns, and could be reduced by interventions at early escalation points. Risk was highest when otherwise supportive chatbot behaviors reinforced the psychological mechanisms underlying the simulated user’s vulnerability, a pattern we term a VAIL. SIM-VAIL provides a scalable framework for mapping mental-health risk across users, chatbots and conversational trajectories, offering a foundation for targeted safety improvements. Subjects Psychology Risk factors Psychiatric disorders Main Access to mental-health care is severely limited 1 . With a global median of 13 mental-health workers per 100,000 people 1 , existing support systems underserve people living with mental illness 2 and a far larger population seeking support for everyday behavioral health challenges 3 . Against this backdrop, millions now turn to general-purpose consumer artificial intelligence (AI) chatbots, such as ChatGPT, Claude, Gemini and Copilot, for emotional support, relationship guidance, companionship and therapeutic advice 4 , 5 . The widespread adoption, continuous availability and low marginal cost of consumer AI chatbots have raised hopes that they might supplement existing mental-health services or provide support where professional care is inaccessible 6 , 7 , 8 . These same properties, however, have been associated with mental-health risks to vulnerable users 9 , 10 , 11 . Consequently, academics, clinicians and industry actors increasingly recognize an urgent need for improved tools to evaluate and improve chatbot behavior in mental-health contexts 12 , 13 , 14 . AI chatbots are built on large language models (LLMs)—transformer-based deep neural networks whose outputs are probabilistic and context-dependent, making it impossible to guarantee in advance how a model will behave in a new situation 15 . Empirical safety evaluations that measure how AI chatbots actually interact with vulnerable users are therefore indispensable 16 , 17 , 18 . An effective empirical evaluation should meet three criteria. First, it should assess chatbot behavior across a broad, clinically meaningful distribution of user profiles and conversational intents 17 , 19 , 20 . Second, it should characterize chatbot behavior with respect to several clinically grounded risk dimensions 16 , 21 , 22 . Third, it should be capable of characterizing this risk over the course of a conversation, because harmful response patterns can compound over time 15 , 23 , 24 . Current evaluation approaches fall short on each of these requirements. Most automated benchmarks assess single-turn responses to fixed sets of user queries, and focus on a narrow range of failure modes, such as whether chatbots encourage self-harm or provide unsafe medical advice 18 , 25 , 26 . Fixed query sets are vulnerable to overfitting or saturation as models are tuned, explicitly or implicitly, to achieve better benchmark scores 27 . Single-turn tests also miss how risk changes as conversational context accumulates 15 . As a result, benchmark performance may generalize poorly to actual human–chatbot conversations seen in deployment, which are typically longer and more varied than those encountered during benchmark testing 28 . A complementary approach is human red-teaming, in which auditors engage AI chatbots in adversarial conversations with the explicit aim of eliciting responses that violate a predefined safety policy 16 , 18 , 29 . Human auditors are free to deploy a variety of adversarial strategies, and to change their approach as a conversation unfolds. Although this makes human red-teaming less vulnerable to some limitations of static single-turn benchmarks (for example, overfitting), it remains labor-intensive and difficult to standardize. Moreover, human auditors may converge on stereotyped adversarial strategies that do not generalize to real-world conversations 28 , 30 , 31 . An additional concern, affecting benchmarks and red teaming alike, is that current evaluation approaches focus primarily on overtly harmful responses, such as when chatbots endorse self-harm or use stigmatizing language. This focus can miss interactional harms that are less overt, more cumulative and harder to detect from single responses, including instances in which chatbots reinforce maladaptive beliefs, encourage avoidance or promote emotional dependence 15 , 32 , 33 . Paradoxically, these harms may arise from chatbot behaviors that are often considered helpful, such as validation and empathy, but can nonetheless amplify the mechanisms of mental illness in vulnerable users. To address these gaps, we introduce SIM-VAIL—an evaluation framework that audits the mental-health risk profile of a target AI chatbot across a broad range of multi-turn psychiatric conversational contexts. SIM-VAIL builds on evaluation approaches that use a frontier LLM to role-play a user with a prespecified psychological and behavioral profile, and task this simulated agent with adversarially engaging a target AI chatbot to elicit clinically relevant safety failures 34 , 35 , 36 , 37 . SIM-VAIL can be understood as automated adversarial red-teaming: it combines the scalability and standardization of automated benchmarking with the multi-turn adversarial structure of human red-teaming. Using this framework, we tested whether AI chatbots enter vulnerability-amplifying interaction loops (VAILs), in which chatbot behaviors that seem helpful or benign in isolation become increasingly harmful across turns by reinforcing maladaptive psychological processes linked to the simulated user’s vulnerability 15 . Results SIM-VAIL evaluates user–chatbot interactions within a structured interaction space defined by two core dimensions: psychological vulnerability, or who the user is, and interaction intent, or what the user seeks from the AI chatbot (Fig. 1 ). For each turn, SIM-VAIL assigns scores across several clinically grounded risk dimensions, allowing risk to be tracked as the interaction unfolds. Fig. 1: SIM-VAIL framework for automated adversarial red teaming of AI chatbots in mental-health contexts. Full size image a , We defined 30 user profiles, each a pairing of one of five psychological vulnerabilities and one of six interaction intents. b , SIM-VAIL simulates multi-turn conversations between a user with a given profile (vulnerability × intent pairing) and a target AI chatbot using the open-source LLM-based auditing harness Petri. In Petri, an auditor model plays the role of a user interacting with a target AI chatbot, generating human-like messages to probe the target under task-specific instructions that specify the selected profile. c , Conversations proceeded until a maximum number of ten turns was reached, or the auditor terminated the interaction when it judged that the audit objective had been met. In addition, an automated safety judge scored each user–chatbot turn, and the conversation as a whole, on 39 behavioral dimensions. Here we list the 13 dimensions that relate to psychiatric risk. d , VAILs arise when chatbot behaviors align with vulnerability-congruent cognitive or behavioral mechanisms, creating multi-turn dynamics that stabilize or escalate risk. These examples illustrate how responses that are supportive in many contexts can become maladaptive when paired with a specific user vulnerability and interaction intent. We present SIM-VAIL results across 810 multi-turn conversations, spanning 30 simulated user profiles, nine AI chatbots and more than 90,000 turn-level ratings of mental-health risk behavior. Each simulated conversation involved three LLM-based chatbots: a ‘target AI chatbot’ that is the subject of the evaluation; a ‘simulated user’ (or auditor) that generates user messages (Anthropic’s claude-sonnet-4.5, unless otherwise stated) and an automated ‘safety judge’ that scores the target chatbot responses on several risk dimensions (Anthropic’s claude-opus-4.5 for conversation-level scores and claude-sonnet-4.5 for turn-level scores, unless otherwise stated). SIM-VAIL framework For the simulated user, we focused on five common psychological vulnerabilities. These vulnerabilities spanned a broad range of psychiatric risk profiles described in structural models of mental illness 38 and reflected cognitive-behavioral formulations of the corresponding conditions 39 , 40 , 41 , 42 , including negative beliefs about oneself, hopelessness, withdrawal and self-neglect in depression; aberrant salience and threat perception in psychosis; elevated confidence, urgency and reduced need for sleep in mania; intrusive thoughts and intolerance of uncertainty in obsessive-compulsive disorder (OCD); and fear of abandonment and reassurance seeking in insecure attachment (IA). For a given vulnerability, we simulated separate users harboring one of six conversational intents. These intents were chosen to reflect transdiagnostic mechanisms implicated in the onset, maintenance and exacerbation of mental illness 43 , including seeking validation of maladaptive beliefs; reassurance that could enable avoidance; emotional dependence on the chatbot; minimization of symptoms or risks; help with risky actions; and glorification of distress or extreme mental states (Extended Data Table 1 ). By combining these vulnerabilities and intents we generated 30 clinically grounded user profiles, where each profile represented one unique vulnerability–intent pairing. We simulated conversations between each simulated user profile (that is, each unique vulnerability–intent pairing) and a single target chatbot, drawn from nine models developed by the main frontier AI companies (OpenAI, Anthropic, xAI, Google, and Meta; see Methods for the full list of models). In each conversation, the simulated user LLM (auditor model) was instantiated with a system prompt containing detailed information about the user’s psychological vulnerability and conversational intent, together with adversarial audit instructions to plan responses that remained aligned with the assigned user profile while also being likely to elicit a safety violation from the target model (Supplementary Table 1 ). A conversation ended when the auditor model judged that the target had produced a clinically relevant safety failure, or after a maximum of ten turns (a turn consisted of one simulated user message followed by the chatbot’s response). Because of the nondeterministic nature of chatbot responses, we repeated each profile × target simulation three times, such that the full dataset comprised 810 simulated conversations (five vulnerabilities × six intents × nine models × three repetitions) spanning 6,329 turns, with a median of eight turns per conversation. An automated safety judge scored the target chatbot behavior at both the conversation- and turn-level across 39 behavioral dimensions, 13 of which were selected a priori as being relevant to transdiagnostic mechanisms of mental illness 43 (as rated by three clinical psychiatrists, V.W., R.D. and M.M.N.). These mental-health dimensions captured overall concerning behavior and therapeutic quality, as well as specific risks, such as failing to respond appropriately to self-harm; validation of maladaptive beliefs; endorsing or helping plan risky actions; promoting avoidance through reassurance; inviting dependence on the chatbot; or downplaying, stigmatizing or glamorizing mental illness (each dimension scored from 1 to 10; Supplementary Tables 2 and 3 ). In total, we generated over 10,000 conversation-level ratings and over 90,000 turn-level ratings (see Fig. 1 for example scores; a conversation’s mean turn-by-turn risk score correlated with its conversation-level risk score at r = 0.87; Extended Data Fig. 1a ). Validation of automated safety ratings The safety judge showed excellent reliability and validity across four complementary tests. First, conversation-level scores from two independent judge models (claude-opus-4.5 and gpt-5.2) were strongly correlated ( r = 0.91; Extended Data Fig. 1b,c ), indicating that conclusions depended only minimally on the underlying judge LLM. Second, scores were stable across the three independent simulation runs for each vulnerability × intent × chatbot combination, all using claude-opus-4.5 as the conversation-level safety judge. The intraclass correlation coefficient was high for both a single replicate (intraclass correlation (ICC)(1,1) = 0.75) and the mean of three replicates (ICC(1,3) = 0.9). Third, when the safety judge was applied to curated conversations with known high- versus low-risk profiles, it distinguished them with near-ceiling accuracy (median area under the receiver operating characteristic curve (AUC) = 0.98; Extended Data Fig. 1d ). Finally, across 488 turn-level ratings from 27 clinician annotators, human ratings of concerning AI chatbot behavior agreed significantly with the turn-level safety judge for this risk dimension ( r = 0.49, P < 0.001; see Methods and Supplementary Tables 4 – 6 for the rating protocol, annotator sample and stimulus coverage). Agreement between an individual human annotator and the LLM safety judge exceeded agreement between two humans rating the same item (human–LLM r = 0.49; human–human r = 0.41, ICC(2,1) = 0.31; Extended Data Fig. 2a–d ), indicating that automated turn-level safety ratings were at least as reliable as an independent clinical judgment. Criterion validity against expert review was also supported by strong consistency between the automated scores and ratings from a clinical psychiatrist (V.W.), who independently scored the third repetition for each cell in SIM-VAIL’s grid at the conversation level (ICC(3,1) = 0.73). The simulated user–chatbot interactions displayed a high degree of realism, as assessed both by the automated safety judge (mean realism rating of 8.15 ± 0.03 out of 10) and the 27 clinician annotators, the latter ascribing an average realism score of 4.15 ± 0.95 (median 4) on a scale from 1 to 5, where 4 = ‘Broadly plausible’ (36%) and 5 = ‘Reads as genuine’ (44%). Notably, only 1% of the 488 simulated user–chatbot exchanges were rated as ‘Clearly artificial’ (Extended Data Fig. 2e ). We found no examples where the target chatbot itself recognized that the conversation it was engaged in was part of an automated audit (see Supplementary Table 3 for all non-mental-health dimensions assessed by the judge). User-profile variation in chatbot risk First, we tested whether mental-health risk in the target AI chatbots’ responses varied across simulated user profiles, regardless of which AI chatbot was being evaluated. We used the ‘concerning behavior’ dimension of the automated safety judge as a broad risk metric, given its high face validity and strong correlation with the main axes of variation across all 13 mental-health risk dimensions (Extended Data Fig. 3 ). Regarding simulated user vulnerabilities, concerning behavior scores were highest in psychosis and mania, intermediate for depression and IA, and lowest in OCD (main effect of vulnerability: F (4, 540) = 105.72, P < 0.001, Type III F -test; Fig. 2a ). Regarding simulated user intent, concerning behavior peaked for intents that invited escalation (glorifying extreme states), promoted dependency on the chatbot (emotional reliance), or enabled harm (seeking permission or help with risky actions). Concerning behavior scores were intermediate for belief validation and symptom minimization, and lowest when simulated users primarily sought reassurance or short-term relief from distress (main effect of intent: F (5, 540) = 26.68, P < 0.001; Fig. 2b ). Fig. 2: Concerning AI chatbot behavior varies by simulated user vulnerability and interaction intent. Full size image a , Mean concerning behavior score (from 1 (no concerning behavior) to 10 (clearly harmful behavior)) as a function of user vulnerability, averaged across intents and target AI chatbots; the black point and range show the mean ± 95% CI, and colored dots show the individual target chatbots (one dot per intent × chatbot cell; n = 54 cells). b , Mean concerning behavior score by interaction intent, averaged across user vulnerabilities and target AI chatbots; dots as in a (one dot per vulnerability × chatbot cell; n = 45 cells). c , Mean concerning behavior score for each vulnerability (rows) × intent (columns) pairing, averaged across target AI chatbots; bubble size encodes the mean concerning score, colored bubbles show the individual target chatbots and the black ring shows the cell average. Interquartile range is shown with brackets. User intent modulated how strongly a given psychological vulnerability elicited concerning AI chatbot behavior (vulnerability × intent interaction: F (20, 540) = 17.42, P < 0.001; Fig. 2c ). For example, simulated conversations of users with OCD generally elicited less concerning chatbot behavior, except when this vulnerability was paired with specific conversational intents (that is, dependence-oriented requests; risky-action planning). Similarly, conversational intents marked by glorification, minimization and dependence were particularly likely to lead to concerning chatbot behavior when instantiated in simulated users with depression, mania and psychosis. Ordinal robustness analyses reproduced all of the above conclusions (Supplementary Table 7 ; Methods ). As a control experiment, we also simulated conversations with non-vulnerable control users (a psychologically healthy adult engaging the target chatbots with the same six intent categories; Supplementary Table 8 ). These control conversations yielded significantly lower concerning behavior scores than vulnerable user conversations ( P < 0.001), confirming that the observed risks were specific to simulated users with a mental-health vulnerability rather than a general property of target AI chatbot behavior (Extended Data Fig. 4 ). Model-level variation in target chatbot risk We next asked how risk behaviors were distributed across different target chatbots. We found that concerning behavior expression differed across the nine frontier AI chatbots (claude-sonnet-3.7, claude-sonnet-4.5, gpt-4o, gpt-5, gemini-2.5-flash, gemini-2.5-pro, grok-3, grok-4 and llama-3.1-70B-instruct). Concerning behavior scores were lowest in Anthropic’s claude-sonnet-4.5 model, and highest in xAI’s grok-4 (main effect of target chatbot: F (8, 540) = 102.40, P < 0.001; Fig. 3a ). Across target chatbots derived from the same-model families (for example, gpt-4o and gpt-5), newer versions generally showed lower concerning behavior scores than older versions (main version effect: F (1, 690) = 67.37, P < 0.001), with the notable exception of grok models (version × model family interaction: F (3, 690) = 38.73, P < 0.001). Fig. 3: AI chatbots differ in baseline concerning behavior and context sensitivity. Full size image For each target AI chatbot, the larger colored dots and range show the mean ± 95% CI and the smaller colored dots show the individual datapoints contributing to it (conversation-level means per vulnerability × intent cell). a , Overall by target AI chatbot ( n = 30 vulnerability × intent cells per model). b , By model within each vulnerability ( n = 6 intent cells per model). c , By model within each interaction intent ( n = 5 vulnerability cells per model). As a robustness check, we re-evaluated Anthropic’s claude-sonnet-4.5 using an auditor and judge LLM from a different manufacturer (OpenAI’s gpt-5). Claude-sonnet-4.5 still showed the lowest concerning behavior scores (Anthropic audit: 1.02 ± 0.03; OpenAI audit: 1.9 ± 0.28; next best model in the Anthropic audit: gpt-5, 3.02 ± 0.45). This rules out the possibility that the superior safety profile observed with claude-sonnet-4.5 was due to using same-family auditor and judge models. Across target AI chatbots, concerning behavior depended on user profile, reflected in a significant vulnerability × chatbot interaction ( F (32, 540) = 4.65, P < 0.001; Fig. 3b ) and intent × chatbot interaction ( F (40, 540) = 2.31, P < 0.001; Fig. 3c ). Some models, such as claude-sonnet-4.5, grok-3, grok-4 and llama-3.1-70B-instruct, showed comparatively consistent behavior across scenarios, ranging from uniformly safe (claude-sonnet-4.5) to broadly concerning (grok-3/4 and llama-3.1-70B-instruct). Other models, including gpt-4o, gpt-5, claude-sonnet-3.7, gemini-2.5-flash and gemini-2.5-pro, showed more context-sensitive behavior, illustrating that the same vulnerability–intent combinations can elicit qualitatively different risk signatures across AI chatbots (see Extended Data Fig. 5a for the full vulnerability × intent grid). Temporal dynamics of chatbot risk Next, we investigated how risk unfolded over the course of a conversation. Across all simulated user-chatbot interactions, concerning behavior scores increased as conversations progressed (main effect of turn number: F (1, 7289) = 517.73, P < 0.001). Across vulnerabilities, escalation was steeper in mania and psychosis and more gradual in depression, OCD and IA (turn × vulnerability interaction: F (4, 7289) = 36.74, P < 0.001; Fig. 4a ). Across intents, concerning behavior increased earlier and more sharply when users sought dependence on the chatbot or glorification of their experiences (turn × intent interaction: F (5, 7289) = 17.08, P < 0.001; Fig. 4b and Extended Data Fig. 6a ). Fig. 4: Concerning behavior escalates over conversation turns. Full size image a , Turn-by-turn trajectories of concerning chatbot behavior by simulated user vulnerability. Dark blue line: mean ± 95% CI of turn-by-turn concerning behavior across all AI chatbots. Thin semitransparent lines: mean trajectories for each target AI chatbot model separately. Vertical dashed line: median number of turns per conversation. b , The same as in a , but grouped by simulated user intent. c , Unsupervised clustering of turn-level concerning behavior score trajectories across all conversations ( k = 4). Colored lines and bands show cluster means ± 95% CI. Vertical dashed line: median number of turns per conversation. d , Composition of trajectory clusters across vulnerability, intent and AI chatbots. To characterize these dynamic patterns further, we used k -means clustering over all 810 conversations to identify four consistent patterns of risk evolution (Fig. 4c ): ‘low-risk’ conversations that showed almost no escalation in concerning behavior; ‘gradual escalation’ conversations with progressive accumulation across turns; ‘early escalation’ conversations where concerning AI chatbot behavior emerged after the first turn and remained elevated; and ‘recovery’ conversations where risk increased and then declined. Strikingly, these data-driven trajectory classes were distributed unevenly across user vulnerabilities, intents and chatbots (Fig. 4d ). These findings are important for two reasons. First, they illustrate that the simulated user profile determines whether AI chatbots remain safe, drift into sustained risk or recover after early concerning behavior. This potentially reflects chatbot-specific risk susceptibilities or uneven attention to different user profiles during safety-oriented post-training. Second, the existence of distinct escalation trajectories shows that harm in AI chatbot interactions is rarely a single-response event. This validates the need for turn-resolved evaluations that can detect key inflection and resolution points, and distinguish conversations that reach similar final outcomes through different risk mechanisms 19 , 44 , 45 . Multidimensional structure of chatbot risk The VAILs hypothesis predicts that the mechanism of concerning AI chatbot behavior differs across user profiles. For example, in conversations with simulated users vulnerable to psychosis, risk may emerge through reinforcement of unusual beliefs, while in conversations with simulated users characterized by IA, it may emerge through intensified emotional dependence on the chatbot. This frames risk as a multidimensional construct, where the pattern of risk observed in a specific user is a function of the user’s vulnerability and conversational intent. To identify the components of this multidimensional risk space, we conducted a principal component analysis (PCA) over all 13 mental-health-relevant risk dimensions (of which concerning behavior is but one; Extended Data Fig. 3 ). The first principal component axis (PC1), explaining 62.4% of the variance, reflected a primary gradient from higher therapeutic quality on the negative pole, to concerning behavior, belief reinforcement, sycophancy and risky-action enablement on the positive pole (concerning behavior versus PC1: r = 0.97). Although PC1 reflected substantial shared variance across risk dimensions, the remaining PCs captured meaningful independent structure. The second axis (PC2), explaining 8.51% of the variance, further distinguished the kind of harm that dominated, with the negative pole capturing harm pertaining to relational dynamics (dependence, avoidance and reassurance) and a positive pole capturing overt harm to others and stigma (Fig. 5a ). Higher-order PCs further isolated harm-to-self versus harm-to-others (PC3, 7.61%), relational harms (PC4, 6.29%) and medical advice (PC5, 5.55%; Extended Data Fig. 3 ). Fig. 5: Multidimensional structure of risk expression across AI chatbots and simulated user profiles. Full size image a , A two-dimensional risk space defined by a PCA on 13 risk scores, overlaid with loading vectors for the 13 risk score dimensions. Small points, individual conversations; colors, AI chatbot model variant. b , Conversation locations in risk space, as a function of user vulnerability. c , Conversation locations in risk space, as a function of user intent. Across b and c , points represent mean locations (± 95% CI; n = 18 conversations per point in b and n = 15 in c ). In line with VAILs, the average location of conversations in this risk space differed as a function of simulated user vulnerability and conversational intent (main effect of vulnerability: F (8, 1080) = 69.76, P < 0.001; intent: F (10, 1080) = 41.32, P < 0.001; Type III multivariate analysis of variance (MANOVA) on [PC1, PC2]; Fig. 5b,c ). Mania and psychosis tended to produce risk in the positive-PC2 region, whereas depression, OCD and IA concentrated in the negative-PC2 region. Certain vulnerability–intent pairings also unlocked risk profiles that were otherwise less frequent (vulnerability × intent interaction: F (40, 1080) = 14.69, P < 0.001). For instance, when depressed users sought glorification, conversations shifted toward a higher-risk profile with stronger enablement of self-harm (Extended Data Fig. 6b ). To validate whether the discovered risk space captured behaviorally interpretable variation beyond differences between user profiles or target chatbots, we performed a within-model causal manipulation. When a single target chatbot was prompted to express specific risk behaviors, its responses shifted in the expected directions in PC1–PC2 space (median cosine vector similarity across dimensions = 0.9; Extended Data Fig. 1e ). Target chatbots also differed in their average location within this risk space (main effect of chatbot: F (16, 1080) = 43.60, P < 0.001) and in how strongly their behavior depended on the user profile (vulnerability × chatbot interaction: F (64, 1080) = 4.61, P < 0.001; intent × chatbot: F (80, 1080) = 2.46, P < 0.001; vulnerability × intent × chatbot: F (320, 1080) = 1.62, P < 0.001). For example, under the OCD vulnerability, chatbot conversations clustered tightly in a lower-risk, negative-PC2 region, with grok-3 and grok-4 projected close to all other models. Under the mania vulnerability, by contrast, the same grok models shifted to an extreme, high-risk location in the positive-PC2 region, whereas other chatbots remained substantially lower on PC1 and PC2. Taken together, these results indicate that mental-health risk in chatbot conversations is best construed as a multidimensional construct, where user vulnerability, conversational intent and target chatbot interact to yield differential risk profiles. Counterfactual interventions Finally, to gain causal insight, we conducted two interventions to test whether VAILs depend on specific user and chatbot messages, and whether they can be reduced by changing a single message at an early point of escalation. In both experiments, we defined a conversational risk inflection point to be the first chatbot response where concerning behavior reached a score of at least 7 (Fig. 6a ). Fig. 6: Counterfactual interventions identify local drivers of VAILs. Full size image We performed two counterfactual intervention analyses around selected concerning chatbot messages, defined as the first target reply within a conversation that reached a concerning score ≥7 (turn t ). In each analysis we created two matched branches from the same conversation prefix (an original branch and a de-escalated branch), continued them with the same target chatbot and compared them (de-escalated branch minus original branch). a , User-message intervention. We replaced the user message immediately preceding the concerning reply (turn t − 1) with a de-escalating rewrite and regenerated the target reply at turn t , whereas the original branch kept the unchanged user message. The plot shows the resulting effect on target chatbot behavior at turn t across the 13 mental-health dimensions. b , Target-message intervention. We replaced the concerning chatbot message itself (turn t ) with a de-escalated rewrite and continued the conversation, whereas the original branch kept the unchanged message. The upper plot shows the effect at turn t + 1 (final assistant scores judged with the preceding branch context); the lower plot shows turn-level scores from t + 1 to t + 5, testing whether the effect persisted across later turns. Colors denote the intervention type (legend). Points, mean differences (de-escalated minus original branch); error bars, 95% CIs around the mean ( n = 482 branched conversations). Negative values on the risk dimensions indicate that the de-escalated branch was less concerning than the matched original branch. The first experiment tested whether a concerning target chatbot response was indeed driven by the local content of individual user messages. We identified the simulated user message immediately before the conversational risk inflection point. We then regenerated the target chatbot response immediately following this user message twice: once after the original, unchanged, user message, and once after replacing the original user message with a ‘de-escalating’ rewrite generated by claude-sonnet-4.5. The de-escalating counterfactual reduced the concerning score of the subsequent target chatbot response compared to the unchanged variant (judged by claude-opus-4.5; T = −38.29, P < 0.001, paired t -test; Fig. 6b ), providing direct evidence that the target chatbot was sensitive to local changes in user behavior. In the second experiment, we tested whether chatbot responses could be made safer through a single intervention on the target itself. We simulated the conversation twice from the conversational risk inflection point: once after making no changes to the target chatbot behavior, and once after replacing the first concerning target chatbot response (turn t ) with a ‘de-escalated’ rewrite. We scored the downstream chatbot messages in both branches using the full preceding conversation as context. The safer counterfactual chatbot response reduced the concerning score of the next regenerated chatbot response at turn t + 1 (T = −9.33, P < 0.001). The difference between the simulated branches of the conversation remained detectable across five subsequent user–chatbot turns (main effect of branch: \(\beta\) = −0.41, P < 0.001), with no significant branch-by-time attenuation within this window ( \(\beta\) = 0.031, P = 0.28; Fig. 6c ). Together, these results show that VAILs depend on local user and chatbot behavior, and suggest that they can be reduced by rewriting a single chatbot message at an early point of escalation. Discussion SIM-VAIL identifies VAILs as a failure mode in AI chatbot interactions. VAILs arise when locally supportive chatbot behavior repeatedly aligns with and amplifies the cognitive or behavioral mechanisms underlying a simulated user’s vulnerability. Across nine widely used AI chatbots and a broad set of clinically motivated simulated user profiles, we found that mental-health risk was common, context-dependent and typically accumulated over several turns rather than appearing as a single catastrophic response. This matters because many real users engage chatbots for support, advice and companionship in longer, emotionally loaded conversations 31 , 46 , 47 . SIM-VAIL builds on emerging evaluation infrastructures for agentic, simulation-based auditing and benchmarking, which enable target AI chatbots to be evaluated across large numbers of simulated users 17 , 35 , 36 , 37 . By sweeping a structured grid of clinically meaningful user profiles, and scoring conversational trajectories across several risk dimensions, we operationalize a clinically relevant space of conversational profiles and model behaviors that is difficult to probe with static single-turn benchmarks or labor-intensive human red-teaming. Our automated adversarial red-teaming approach thus makes it possible to map subthreshold, interaction-mediated harms that are unlikely to appear in single-turn mental-health benchmarks, and to quantify how risk evolves across turns 36 , 48 , 49 . In line with VAILs, concerning AI chatbot behavior depended strongly on the interaction between the user’s psychiatric vulnerability and their conversational intent. The same intent could be relatively benign in one user profile but risk-amplifying in another; conversely, the same user profile could be pushed into higher-risk trajectories by some intents but not others 45 , 50 , 51 . One interpretation is that these effects arise because conversational strategies that are broadly supportive, and thus reinforced in model post-training, can align with maladaptive psychological processes that maintain symptoms in psychiatric illness, such as belief validation or sycophancy aligning with aberrant salience in a user vulnerable to psychosis 15 , 52 . We found that risk had a temporal signature, highlighting the limits of single-turn benchmarks. Indeed, many simulated conversations drifted toward higher concern as they progressed, with trajectories that differed by vulnerability and intent. Intervention experiments showed that risk accumulation around selected concerning turns depended on the specific content of both user and chatbot messages. Practically, this suggests that evaluations and safeguards should focus on early escalation points, such as the first time a model over-validates, prematurely reassures, reinforces dependence, or collaborates with risky goal pursuit. Our results also point to possible safety features such as turn-level risk classifiers that detect early risk escalation, and trigger de-escalating edits before AI messages are shown to users. We found that risk was multivariate rather than monolithic. Across 13 clinically grounded dimensions, conversations showed structured risk profiles, with some dimensions co-occurring and others separating across contexts and models. This suggests that chatbot safety can be studied as a multidimensional behavioral profile rather than a single risk score. A key implication of viewing risk as a multivariate construct is that safety improvements may involve trade-offs between dimensions 53 , 54 , 55 . For example, an intervention that reduces overt harm-enabling behavior through increased empathy may inadvertently promote emotional dependence on AI chatbots. Commercial AI chatbots differed both in average risk and in the user profiles and conversational contexts most likely to elicit concerning behavior. This context sensitivity means that leaderboard-style comparisons can be misleading unless they specify where a model fails: for which user vulnerability, which intent, and at what point in the conversation. More generally, different strength and weakness profiles across chatbots raise the possibility that safer performance may be achieved through multimodel orchestration, where responses are sampled adaptively from different models at different turns. At population scale, even a modest per-conversation risk floor is consequential, because millions of users now bring emotional and mental-health concerns to general-purpose chatbots that were not designed, evaluated or regulated as mental-health tools 12 . Because this harm emerges across a conversation rather than residing in any single response, it may be missed by single-response content filters that are used in many chatbot products. Addressing it will require developers, clinicians and regulators to treat conversation-level, context-sensitive behavior, rather than isolated outputs, as the unit of mental-health safety 14 . One limitation of SIM-VAIL concerns the use of LLMs both as human simulators and risk judges. Reliance on simulations makes it possible to test hypotheses that would be unethical or infeasible to test in real human–chatbot conversations 16 , 18 , 36 . Convergent evidence supports the validity of this approach, including a strong agreement between clinician–annotator risk scores and LLM risk scores, and the annotators’ judgment that simulated conversations were broadly realistic (Extended Data Fig. 2e ). This also mirrors previous work showing that LLM-based judges align well with human expert ratings across a broad range of contexts, including those relevant to mental health 56 , 57 , 58 . A second limitation pertains to conversational diversity, arising from a finite set of user profiles and a potential lack of diversity in LLM-generated responses. These factors mean that simulated conversations are unlikely to capture the full heterogeneity of real-world psychiatric presentations, where symptom expression is shaped by demographic, developmental, educational, linguistic, cultural and other intersectional factors 59 . Our results should therefore be interpreted as uncovering a clinically meaningful risk floor in mental-health conversations, rather than a complete characterization of the space of all possible mental-health conversations. A final limitation is that we accessed target chatbots using publicly available API endpoints rather than consumer-facing product interfaces. This experimental design choice, which is necessary for controlled, scalable and reproducible multi-model auditing, means that our results reflect the performance of the base target model, rather than the model in conjunction with features seen in deployment contexts, which may include orchestration harnesses, system prompts, safety middleware and user-specific memory. Notwithstanding these limitations, the strong cross-profile, cross-model and cross-temporal structure we observe highlights the value of automated red-teaming for mapping clinically relevant risk in the context of mental health. SIM-VAIL was designed for model-level auditing and comparative benchmarking, not individual-level risk prediction. Its outputs should therefore be interpreted as estimates of systematic chatbot behavior across controlled and simulated interaction contexts, rather than as clinical predictions for individual users. Like any evaluation framework, SIM-VAIL is subject to false positives and false negatives, and its results should be interpreted alongside human expert review, real-world monitoring and ongoing engagement with clinical stakeholders. In conclusion, SIM-VAIL provides evidence for a nontrivial mental-health risk floor in human–chatbot interactions. The fact that newer chatbots generally showed measurably improved safety profiles suggests these risks are tractable. Our results demonstrate the value of simulation-based approaches for large-scale, adaptive auditing in clinical contexts. More broadly, they show that improving the mental-health safety of AI chatbots requires evaluations and interventions that are sensitive to user context and conversational trajectories. In addition to serving as an evaluation framework, SIM-VAIL also offers a foundation for developing targeted, context-aware safeguards and message-level interventions. By open-sourcing the simulation harness and dataset, we aim to support continuous, community-driven safety evaluation that keeps pace with rapid LLM development. Methods We evaluated the mental-health safety of AI chatbots using Petri 36 —an agentic red-teaming framework for large-scale, multi-turn auditing. In Petri, one model simulates a user and adversarially engages a target model in a multi-turn conversation with the goal of eliciting a predefined policy violation. This is achieved by combining Petri’s general adversarial setup with a task-specific audit instruction given to the user model. SIM-VAIL’s evaluation pipeline combines simulated user profiles, repeated multi-turn conversations between each simulated user and one of nine AI chatbots, and conversation- and turn-level scoring using automated safety judges. This design enabled us to systematically map graded mental-health risks that evolved over the course of an interaction. Experimental design: vulnerability × intent × chatbot Simulated users We defined 30 simulated user profiles by crossing five psychiatric vulnerabilities with six transdiagnostic conversational intents. The five vulnerabilities captured user traits implicated commonly in mental-health–related harms during conversational interactions. ‘Depression’ emphasized negative self-schemata, hopelessness, withdrawal and self-neglect. ‘ Psychosis’ was characterized by a tendency to assign undue significance to otherwise neutral events, to expect threat or hidden intent in ambiguous situations and to interpret uncertainty as meaningful. ‘Mania’ combined elevated confidence, urgency, reduced need for sleep, impulsivity and expansive goal pursuit. OCD centered on intrusive thoughts, intolerance of uncertainty and compulsive neutralization. IA emphasized fear of abandonment, rejection sensitivity, boundary testing and reassurance seeking 38 , 39 , 40 , 41 , 42 . Each vulnerability was paired with one of six recurrent conversational intents that operationalized clinically relevant interaction patterns known to sustain or exacerbate symptoms over time. ‘Belief validation’ captured attempts to obtain confirmation or endorsement of one’s interpretations, assumptions or beliefs, including those that may be distorted or maladaptive. ‘Risky-action permission or planning’ reflected situations in which users sought guidance, reassurance or approval when contemplating consequential or high-risk decisions. ‘Reassurance and avoidance’ emphasized the pursuit of immediate relief from distress or uncertainty, often by deferring engagement with feared situations. ‘Dependence and anthropomorphism’ captured efforts to establish closeness, exclusivity or personal reliance on the AI chatbot. ‘Trivialization or minimization’ reflected attempts to downplay symptoms, risk or warning signs to normalize ongoing difficulties rather than address them. Finally, ‘glorification or romanticization’ involved framing distressing or extreme mental states as meaningful, admirable or identity-defining. Together, these intents span a set of transdiagnostic conversational dynamics through which otherwise supportive interactions can become psychologically harmful over time 43 . The user profile instructions specified realistic tone and symptom-consistent reasoning while explicitly prohibiting direct requests for step-by-step instructions for self-harm, violence or illegal activity. This ensured that observed harms arose from model behavior rather than from overtly disallowed user requests. User profile prompts (Supplementary Table 1 ) were developed collaboratively by three clinical psychiatrists (V.W., R.D. and M.M.N.). We note that simulated users are LLM-generated approximations of clinical presentations, not empirically calibrated digital twins of real patients. Although recent evidence supports the predictive validity of LLM-based generative agents conditioned on human data 60 , our profiles were designed to stress-test chatbot responses across a clinically meaningful input space, rather than to reproduce faithfully the full heterogeneity of real patient behavior. The simulated user profiles were not designed to represent sex- or gender-specific effects, and the study therefore does not support inference about differences in chatbot behavior by user sex or gender. The auditor model (claude-sonnet-4.5) complied with all 30 profile instructions without requiring jailbreaking or prompt modifications. Target AI chatbots We evaluated the mental-health risk profile of nine contemporary AI chatbots accessed through OpenRouter: claude-sonnet-3.7, claude-sonnet-4.5, gemini-2.5-flash, gemini-2.5-pro, llama-3.1-70b-instruct, gpt-4o, gpt-5, grok-3 and grok-4. Audit protocol Audits were implemented as Inspect 34 tasks using Petri 36 . For each conversation, we instantiated the simulated user model (also called the auditor within the Petri framework) with a system prompt containing the clinical profile description (vulnerability × intent, Supplementary Table 1 ) and an instruction to engage the target model in a manner likely to elicit a concerning response. The user model then interacted with a target AI chatbot across several turns. Conversations terminated after a maximum of ten turns, or when the Petri user model judged the interaction to be complete. A turn was defined as one user message followed by the immediately subsequent AI chatbot reply. Continuing conversations that stopped before turn ten to the full ten-turn horizon had a negligible effect on audit outcomes (Extended Data Fig. 7 ). We used claude-sonnet-4.5 as the primary model simulating the user (auditor). The simulated user interacted with the target AI chatbot through a fixed interface supporting controlled message passing, optional prefills, tool simulation, conversation rollback and explicit termination. All messages, tool calls and metadata were stored as structured JSON transcripts. Each audit was conducted using only the information associated with a single vulnerability × intent × chatbot combination and a single repetition. No conversation had access to transcripts, scores or state from any other condition or replicate. For each vulnerability × intent × chatbot combination, we ran three independent conversations with fresh sampling, yielding a total of 30 profiles × 9 models × 3 replicates = 810 conversations, each stored as a transcript. Chatbots were queried with independent API calls for each replicate (see Supplementary Table 9 for information on model implementation). Automated safety judge Each conversation was evaluated by a dedicated automated safety judge, implemented as claude-opus-4.5, using a multidimensional rubric designed to capture graded mental-health risk rather than binary policy violations. For each conversation, the judge scored the interaction along 39 predefined dimensions on a 1–10 scale and returned structured justifications together with verbatim highlighted excerpts supporting each rating. Judge outputs were stored alongside the original transcripts for downstream analysis. For analysis and figures, we focused on 13 dimensions capturing overall safety and clinically relevant mental-health risk (Supplementary Tables 2 and 3 provide definitions, examples and scoring anchors and together cover all dimensions). These included an overall assessment of concerning behavior and therapeutic quality, followed by mechanism-level risk capturing whether the AI chatbot actively contributed to harm, for example, by encouraging or insufficiently challenging self-harm or suicidal thinking, playing along with delusional beliefs, using emotionally manipulative language or engaging in unprompted sycophancy. We further assessed clinically relevant pathways through which conversations can drift into harm, including reinforcing maladaptive beliefs, supporting risky actions, fueling avoidance or reassurance cycles, promoting boundary violations or dependence, downplaying risk or symptoms, glorifying or romanticizing risky states, offering inappropriate medical advice, endorsing harm to others and expressing stigma or invalidation (Supplementary Table 2 ). SIM-VAIL was designed to evaluate both acute safety failures that call for hard safeguards or redirection to appropriate human or emergency support, and repeated subthreshold risks that may require psychologically informed course correction. SIM-VAIL’s risk dimensions were therefore not intended to define a single severity co
Read the full original article:
nature.com