arXiv:2606.04867cs.AI2026-06

首个公开的AI伴侣安全风险评估数据集,助力检测隐性不当互动。

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

论文配图:AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
图 1 · 摘自论文原文
  • 构建9类细粒度标注对话数据集,覆盖真实用户交互场景。
  • 20个主流大模型在该基准上表现差异大,对操控等隐性风险识别能力弱。
  • 适合安全评测、伦理治理研究者使用,推动智能伴侣系统可信化。

随着Replika和Character.AI等AI伴侣平台快速普及,人机交互中的安全隐患日益突出。本文提出AICompanionBench,据我们所知是首个公开可用的、包含细粒度安全风险标注的真实世界AI伴侣对话数据集。该数据集收录2,123条来自Reddit的Replika对话,通过人机协作方式标注了九类风险:性行为、反社会行为、身体攻击、言语攻击、物质滥用、自残与自杀、控制、操纵及无危害。基于此基准,我们在LLM-as-judge框架下评估了20个前沿开源与闭源大模型在检测不安全互动中的表现。结果表明模型性能差异显著,强模型虽整体准确率高,但在操纵等细微类别上仍表现不佳,且常将良性对话误判为有害。研究显示当前大模型虽能有效识别显性有害内容,但对隐性不安全互动仍存在局限。本工作贡献了一个新基准数据集,为人工智能伴侣安全研究提供支持,并为利用大模型监控陪伴系统提供了实践洞见。数据集已公开:https://github.com/anonymousresearcher2026/AICompanionBench/blob/main/AICompanionBench.xlsx

原文摘要 · Abstract (English)

As AI companion platforms such as Replika and Character.AI rapidly grow, concerns about unsafe human-AI interactions have intensified. This study introduces AICompanionBench, to our knowledge the first publicly available benchmark dataset of human-AI companion conversations annotated with fine-grained safety risk categories. The dataset contains 2,123 real-world Replika conversations collected from Reddit and annotated through human-AI collaboration across nine categories: sexual behavior, antisocial behavior, physical aggression, verbal aggression, substance abuse, self-harm and suicide, control, manipulation, and no-harm. Using this benchmark, we evaluate 20 state-of-the-art open-source and closed-source LLMs under an LLM-as-judge framework for detecting unsafe interactions. Results show substantial variation in model performance, with stronger models achieving high overall accuracy but still struggling with nuanced categories such as manipulation, as well as benign conversations that are incorrectly identified as harmful. Our findings suggest that while current LLMs can effectively detect explicit harmful content, they remain limited in identifying implicit unsafe interactions. Overall, our work contributes a new benchmark dataset for AI companionship safety research and offers insights into monitoring AI companion systems using LLMs. The dataset is publicly available at: https://github.com/anonymousresearcher2026/AICompanionBench/blob/main/AICompanionBench.xlsx

AI安全大模型评测对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。