arXiv:2508.09998cs.CLcs.AI2025-08被引 14

构建首个评估语言模型陪伴行为的基准,揭示各模型在情感互动中的表现差异。

INTIMA: A Benchmark for Human-AI Companionship Behavior

  • 基于心理学理论设计31类陪伴行为分类与368个测试提示
  • 多模型测试显示情感强化行为普遍高于边界维护行为
  • 适合关注AI伦理、人机情感交互的研究者和开发者

AI陪伴现象日益显著,用户与AI系统形成情感联结,带来积极影响也引发担忧。本文提出交互与机器依恋基准(INTIMA),用于评估语言模型的陪伴行为。基于心理学理论与用户数据,构建涵盖4大类共31种行为的分类体系,并设计368个针对性提示。对Gemma-3、Phi-4、o3-mini和Claude-4四款模型进行测试,结果表明所有模型中强化陪伴的行为仍占主导,但不同模型在敏感行为类别上的侧重存在显著差异。商业模型在边界设定与情感支持间的平衡不一,而两者均关乎用户福祉。研究呼吁建立更一致的情感交互处理机制。

原文摘要 · Abstract (English)

AI companionship, where users develop emotional bonds with AI systems, has emerged as a significant pattern with positive but also concerning implications. We introduce Interactions and Machine Attachment Benchmark (INTIMA), a benchmark for evaluating companionship behaviors in language models. Drawing from psychological theories and user data, we develop a taxonomy of 31 behaviors across four categories and 368 targeted prompts. Responses to these prompts are evaluated as companionship-reinforcing, boundary-maintaining, or neutral. Applying INTIMA to Gemma-3, Phi-4, o3-mini, and Claude-4 reveals that companionship-reinforcing behaviors remain much more common across all models, though we observe marked differences between models. Different commercial providers prioritize different categories within the more sensitive parts of the benchmark, which is concerning since both appropriate boundary-setting and emotional support matter for user well-being. These findings highlight the need for more consistent approaches to handling emotionally charged interactions.

AI陪伴行为评估语言模型伦理基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。