测试大模型在长时间对话中保持角色一致性、听从指令和安全性的能力
Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions
- 设计超百轮对话评测协议,模拟真实长期交互场景
- 发现角色一致性随对话轮次下降,尤其在目标导向对话中更明显
- 揭示角色保持与指令遵循的权衡,适合评估对话系统可靠性
角色设定的大语言模型广泛应用于教育、医疗及社会模拟等领域,但现有评估多局限于短时单轮对话,无法反映真实使用场景。本文提出一种结合超百轮角色对话与数据集的评测协议,构建条件化对话基准,以稳健衡量长上下文影响。我们评估了七种先进开源与闭源大模型在长对话中的角色一致性、指令遵循性与安全性表现。结果显示,角色一致性随对话轮次逐渐退化,尤其在目标导向对话中更为显著;模型需同时维持角色一致性和指令遵循,导致性能下降。初始阶段,非角色基线模型表现优于角色模型;随着对话推进,角色响应趋同于基线,表明角色设定在长程交互中存在脆弱性。本研究为系统评估此类失败提供了可复现的协议。
原文摘要 · Abstract (English)
Persona-assigned large language models (LLMs) are used in domains such as education, healthcare, and sociodemographic simulation. Yet, they are typically evaluated only in short, single-round settings that do not reflect real-world usage. We introduce an evaluation protocol that combines long persona dialogues (over 100 rounds) and evaluation datasets to create dialogue-conditioned benchmarks that can robustly measure long-context effects. We then investigate the effects of dialogue length on persona fidelity, instruction-following, and safety of seven state-of-the-art open- and closed-weight LLMs. We find that persona fidelity degrades over the course of dialogues, especially in goal-oriented conversations, where models must sustain both persona fidelity and instruction following. We identify a trade-off between fidelity and instruction following, with non-persona baselines initially outperforming persona-assigned models; as dialogues progress and fidelity fades, persona responses become increasingly similar to baseline responses. Our findings highlight the fragility of persona applications in extended interactions and our work provides a protocol to systematically measure such failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。