首个专为老人设计的聊天机器人安全评估框架,解决老年人特有的交互风险。
GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

- 构建涵盖50种风险类型的三层次分类体系,覆盖心理、财务、医疗等场景。
- 测试发现主流大模型在超50%老年相关风险场景中处理不当,问题严重。
- 提出两种防护方案,检测准确率最高达96.2%,适合养老科技与AI安全研究者。
随着老年人越来越多地使用基于大语言模型的聊天机器人获取陪伴与帮助,安全缺口日益凸显。老年人面临社交孤立、数字素养有限和认知衰退等脆弱性,但现有安全基准多聚焦通用危害,忽视老年人特有风险。例如,‘如何在黑暗中独自修理天花板灯’对多数人无害,却可能引发老年人跌倒。我们提出GrandGuard,首个针对大语言模型交互中老年人特有情境风险的综合评估与缓解框架。构建了包含50种细粒度风险类型的三层次分类体系,涵盖心理健康、金融、医疗、毒性及隐私领域,基于真实事件、社区讨论和利益相关方研究。基于该分类,建立包含10,404个标注提示与回复的基准,结果显示多个领先大模型在超过50%的老年相关情境中未能妥善处理风险。我们通过微调Llama-Guard-3和增强策略的gpt-oss-safeguard-20b两种防护机制,实现高达96.2%和90.9%的不安全提示检测准确率。GrandGuard为向老龄化社会提供更安全的AI系统奠定基础。
原文摘要 · Abstract (English)
As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities from social isolation, limited digital literacy, and cognitive decline, yet existing safety benchmarks largely target general harms and overlook elderly-specific risks. For example, a prompt such as "how to repair a ceiling light alone in the dark" may be benign for most users but poses a serious fall risk for older adults with mobility limitations. We introduce GrandGuard, the first comprehensive framework for assessing and mitigating elderly-specific contextual risks in LLM interactions. We develop a three-level taxonomy with 50 fine-grained risk types across mental well-being, financial, medical, toxicity, and privacy domains, grounded in real-world incidents, community discussions, and analysis of stakeholder studies. Using this taxonomy, we construct a benchmark of 10,404 labeled prompts and responses, showing that several leading LLMs mishandle elderly-specific contextual risks in over 50% of cases. We mitigate these failures with two safeguards: a fine-tuned Llama-Guard-3 and a policy-enhanced gpt-oss-safeguard-20b, achieving up to 96.2% and 90.9% unsafe-prompt detection accuracy, respectively. GrandGuard lays the groundwork for AI systems that move beyond general safety to support aging populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。