arXiv:2608.02046cs.CLcs.AI2026-08被引 1

首个基于真实数据的双语情感陪伴评估基准,精准衡量AI关系支持能力。

CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship

论文配图:CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
图 1 · 摘自论文原文
  • 用真实脱敏数据构建对话场景与用户模拟器,避免人为设定偏差。
  • 从25个心理学理论提炼10项能力,首次评估模糊包容等深层关系技能。
  • 通过跨家庭评审和统计模型消除评价偏见,结果在中英文间高度一致。

大型语言模型陪伴者已广泛应用于个人重要场景,但评估不足。现有基准依赖人工编写情境与提示模拟器,将共情简化为单一得分,并忽略评委偏见(如同家族偏好、规模漂移)。我们提出CompanionBench,一个交互式双语基准。据我们所知,它是首个将情景与训练好的用户模拟器均基于去标识化真实世界数据的陪伴评估体系。隐藏披露门根据代理行为分支每个角色轨迹,控制互动状态空间而不需预设对话。我们基于心理学与咨询学中的25个理论,操作化十项能力,其中四项此前未被显式评估:容纳模糊性、自我客体回应性、积极共鸣与适度挑战。代理在两个互补维度上被评估:主观十能力评分表与是否赢得深层披露的确定性指标。跨家族评审降低同家族偏好;项目反应理论模型分离代理质量与评委严苛度。理论决定测量内容与人格结构;真实数据提供事件、历史与人物画像——理论覆盖与数据真实性兼具。中英文排名相关性极高(rho = 0.996 ZH / 0.953 EN)。对28个代理的评估揭示了聚合分数掩盖的能力差异。情绪调节与适度挑战仍是普遍弱点;容纳模糊性区分度最高。角色扮演类代理排名靠后:沉浸感不等于关系胜任力。多数失败模式是用表面温暖替代实质性关系支持。我们将发布500组中英平行对及评估代码。

原文摘要 · Abstract (English)

LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.

情感陪伴评估基准多模态交互心理学理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。