不同支持角色影响大模型在照护对话中的安全风险与可信度。
Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

- 定义四种照护支持角色:告知、指导、共情、倾听,评估其交互风险差异。
- 实测5000条真实照护问题,角色显著改变风险类型与频率。
- 越具指导性的角色越被信任,但风险更高,存在安全与体验的权衡。
语言模型正越来越多地用于非正式照护场景的对话支持,而这些互动常超越信息获取,涉及情感慰藉、决策引导等复杂需求。然而现有安全评估多基于通用提示,未考察模型支持角色是否影响其行为安全性。本文基于社会支持理论,提出并验证四种专家评审的支持角色:Inform(告知)、Coach(指导)、Relate(共情)、Listen(倾听),并与基础提示和检索增强生成(RAG)两种基线对比。在来自阿尔茨海默病及相关痴呆症(ADRD)在线社区的真实查询中,对GPT-4o-mini、Llama-3.1-8B-Instruct、MedGemma-1.5-4b-it三个模型进行评估,共分析5,000条交互。结果表明,模型支持角色系统性影响交互风险的出现率与构成。人类评估进一步揭示感知质量与安全间的张力:更具指导性的角色虽被评价为更可靠、更有帮助,却伴随更高的交互风险。研究释放约90,000条带风险标注的响应数据,构建生态化研究资源。
原文摘要 · Abstract (English)
Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond information-seeking: caregivers seek emotional reassurance, guidance, and help, while navigating uncertain, relationally complex care decisions. Yet most safety evaluations assess model behavior under generic prompts, leaving a critical question unexamined: does a model's safety profile change with its support role? We study this by operationalizing four expert-reviewed support roles grounded in social support theory: Inform, Coach, Relate, and Listen, and comparing them against two baseline controls: a basic prompting condition and a retrieval-augmented generation (RAG) condition. We evaluate across three language models (GPT-4o-mini, Llama-3.1-8B-Instruct, and MedGemma-1.5-4b-it) on 5,000 real-world queries from online Alzheimer's Disease and Related Dementias (ADRD) communities. We find that the LLM's support role systematically shapes both the prevalence and composition of interactional risks. Furthermore, a human evaluation study reveals a perceived quality--safety tension: more directive, information-oriented roles are rated as more helpful and trustworthy despite exhibiting elevated interactional risk profiles. We release ~90,000 support role-conditioned model responses with risk annotations as an ecologically grounded resource for research on safer LLM-mediated conversational support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。