arXiv:2510.16829cs.CLcs.AI2025-10被引 1

通过模拟用户角色提升对话AI评估的现实性

Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation

  • 基于真实社区数据构建患者/照护者/从业者三类角色
  • 生成1.5万条带角色特征的问题,验证可信度与真实性
  • 发现不同角色引发模型回应差异:弱势角色获更多支持

语言模型用户常在问题中隐含个人与社会背景,提问者的角色决定了对回应的具体需求。然而多数评估忽视提问者身份,这一缺失在阿片类药物使用障碍(OUD)等敏感领域尤为关键。本文提出CoRUS框架,基于角色理论和r/OpiatesRecovery社区帖子,构建患者、照护者、从业者三类角色分类体系,并据此生成15,321条嵌入角色目标、行为与经验的模拟问题。评估显示这些问题高度可信且接近真实数据。在对五款大模型的测试中,同一问题因角色不同导致响应显著差异:患者与照护者角色触发更支持性回复(+17%)、知识含量更低(-19%),而从业者则获得更专业回应。研究揭示了角色隐含信息对模型输出的影响,提供了角色感知型对话AI评估方法。

原文摘要 · Abstract (English)

Language model users often embed personal and social context in their questions. The asker's role -- implicit in how the question is framed -- creates specific needs for an appropriate response. However, most evaluations, while capturing the model's capability to respond, often ignore who is asking. This gap is especially critical in stigmatized domains such as opioid use disorder (OUD), where accounting for users' contexts is essential to provide accessible, stigma-free responses. We propose CoRUS (COmmunity-driven Roles for User-centric Question Simulation), a framework for simulating role-based questions. Drawing on role theory and posts from an online OUD recovery community (r/OpiatesRecovery), we first build a taxonomy of asker roles -- patients, caregivers, practitioners. Next, we use it to simulate 15,321 questions that embed each role's goals, behaviors, and experiences. Our evaluations show that these questions are both highly believable and comparable to real-world data. When used to evaluate five LLMs, for the same question but differing roles, we find systematic differences: vulnerable roles, such as patients and caregivers, elicit more supportive responses (+17%) and reduced knowledge content (-19%) in comparison to practitioners. Our work demonstrates how implicitly signaling a user's role shapes model responses, and provides a methodology for role-informed evaluation of conversational AI.

对话评估角色建模OUDLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。