arXiv:2606.26982cs.CLcs.AI2026-06

发现大模型在心理支持对话中对相似问题因表述不同而回应差异,影响信任度。

Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions

论文配图:Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions
图 1 · 摘自论文原文
  • 用受控提示测试不同表述下模型响应变化
  • 各架构模型均显示表述改变导致回答倾向系统性偏移
  • 内部表征仍能解码表述信息,提示需关注上下文鲁棒性

大型语言模型(LLMs)正越来越多地应用于心理健康支持等心理敏感对话场景。此类应用中,行为稳定性和一致性对可信的人机交互至关重要。然而,语义相似的问题若通过不同上下文表述,可能引发模型的不同回应。这种对表述的敏感性可能违背用户预期,增加评估AI可靠性的难度。现有研究多关注行为层面的影响,但对对齐模型内部表征中此类变异如何体现了解有限。本研究使用跨多个指令微调模型家族的受控匹配提示,在多种上下文表述条件下进行考察。结果显示,不同表述系统性改变模型的解释性回应倾向。分层探测分析表明,与行为相关的信息在整个Transformer深度中仍可解码,且解码强度随架构不同而异。此外,即使存在强词汇基线,保留的表述探测仍显著高于随机水平。激活操控实验进一步表明,与表述相关的表征方向可部分调节下游行为结果。这些发现提示,在心理导向对话系统部署中,对上下文变化的鲁棒性应成为评估一致性和可信度的重要考量。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being integrated into mental health support tools and other psychologically sensitive conversational applications. In such settings, behavioral stability and consistency are important for trustworthy human-AI interaction. However, semantically similar concerns can be presented through different contextual framings, potentially eliciting different model responses. Such framing-sensitive variability may challenge user expectations regarding system behavior and complicate the assessment of AI reliability. While prior studies have primarily examined such effects at the behavioral level, less is known about how framing-related variation is reflected in the internal representations of aligned language models. In this work, we investigate these effects using controlled matched prompts spanning multiple contextual framing conditions across several instruction-tuned model families. Across architectures, framing systematically alters interpretive response tendencies. Layer-wise probing analyses show that behavior-associated information remains decodable throughout transformer depth, with architecture-dependent variation in decoding strength. Moreover, held-out framing probes remained consistently above chance across architectures despite strong lexical baselines. Activation steering experiments further suggest that framing-associated representational directions can partially modulate downstream behavioral outcomes. Finally, these findings indicate that robustness to contextual variation may represent an important consideration when evaluating the consistency and trustworthiness of conversational AI systems deployed in mental-health-oriented interactions.

大模型心理支持行为稳定性上下文敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。