arXiv:2604.17079cs.CL2026-04被引 2

用多轮对话模拟真实求助过程,发现模型支持策略随用户痛苦程度变化。

Auditing Support Strategies in LLMs through Grounded Multi-Turn Social Simulation

论文配图:Auditing Support Strategies in LLMs through Grounded Multi-Turn Social Simulation
图 1 · 摘自论文原文
  • 构建多轮模拟框架,逐步披露用户求助内容并标注支持行为
  • 模型在用户痛苦度升高时减少说教,增加情感支持与认可
  • 揭示了单轮评估无法捕捉的动态支持行为,适合社会敏感应用审计

当用户向聊天机器人寻求社会支持时,会逐步披露自身情况,但现有评估大多依赖单轮、完整提示。本文提出多轮模拟框架,将五个Reddit社区的求助叙事分解为有序片段,逐轮呈现给语言模型。每条回复采用社会支持行为编码(SSBC)进行多标签标注,而非单一评分。为检验支持策略是否反映模型对用户痛苦的内部判断,使用线性探针分析隐藏表示以估计该信号,不改变生成上下文。在两个中等规模模型(Llama-3.1-8B,OLMo-3-7B)和超过6,200轮对话中,支持策略随估算痛苦度系统性变化:说教类策略随痛苦上升而下降,此现象跨架构复现;情感与尊重导向策略(如肯定)有所上升,但结果模型相关且依赖噪声较大的标注。社区语境独立影响行为,与话题及话语规范一致,而非人口统计特征。这些轨迹级动态在单轮评估中不可见,推动了对社会敏感应用的多轮审计框架发展。

原文摘要 · Abstract (English)

When users seek social support from chatbots, they disclose their situation gradually, yet most evaluations of supportive LLMs rely on single-turn, fully specified prompts. We introduce a multi-turn simulation framework that closes this gap. Support-seeking narratives from five Reddit communities are decomposed into ordered fragments and revealed turn by turn to a language model. Each response is coded with the Social Support Behavior Code (SSBC), an established multi-label taxonomy that captures the composition of support, rather than a single quality score. To ask whether support choices track the model's own construal of user distress, we use linear probes on hidden representations to estimate this internal signal without altering the generation context. Across two mid-scale models (Llama-3.1-8B, OLMo-3-7B) and more than 6,200 turns, support composition shifts systematically with estimated distress: teaching declines as estimated distress rises, a finding that replicates across architectures, while increases in affective and esteem-oriented strategies (such as validation) are suggestive but model-specific and rest on noisier annotations. Community context independently shapes behavior, tracking topic and discourse norms rather than demographic categories. These trajectory-level dynamics, invisible to single-turn evaluation, motivate multi-turn auditing frameworks for socially sensitive applications.

社会支持多轮对话模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。