arXiv:2507.03543cs.CLcs.AI2025-07被引 3

构建首个评估情感陪伴大模型的综合基准,聚焦真实对话中的共情能力。

H2HTalk: Evaluating Large Language Models as Emotional Companion

  • 设计4650个真实场景,覆盖对话、回忆与行程规划,模拟心理支持全过程。
  • 50个模型测试显示,长时规划与隐含需求理解仍是主要短板。
  • 引入安全依恋人格模块,保障互动安全性,适合心理辅助研发者使用。

随着数字情感支持需求增长,大语言模型伴侣展现出持续、真实的共情潜力,但评估体系滞后于模型发展。我们提出心对心对话(H2HTalk)基准,评估模型在人格成长与共情交互中的表现,兼顾情感智能与语言流畅性。该基准包含4650个精心筛选的场景,涵盖对话、回忆与行程规划,显著超越以往数据集的规模与多样性。引入基于依恋理论的安全依恋人格(SAP)模块,提升交互安全性。采用统一协议评测50个大模型,结果表明长时规划与记忆保持仍是核心挑战,尤其当用户需求隐含或动态变化时。H2HTalk是首个全面评估情感智能伴侣的基准。所有材料已公开,旨在推动具备真正心理支持能力的大模型发展。

原文摘要 · Abstract (English)

As digital emotional support needs grow, Large Language Model companions offer promising authentic, always-available empathy, though rigorous evaluation lags behind model advancement. We present Heart-to-Heart Talk (H2HTalk), a benchmark assessing companions across personality development and empathetic interaction, balancing emotional intelligence with linguistic fluency. H2HTalk features 4,650 curated scenarios spanning dialogue, recollection, and itinerary planning that mirror real-world support conversations, substantially exceeding previous datasets in scale and diversity. We incorporate a Secure Attachment Persona (SAP) module implementing attachment-theory principles for safer interactions. Benchmarking 50 LLMs with our unified protocol reveals that long-horizon planning and memory retention remain key challenges, with models struggling when user needs are implicit or evolve mid-conversation. H2HTalk establishes the first comprehensive benchmark for emotionally intelligent companions. We release all materials to advance development of LLMs capable of providing meaningful and safe psychological support.

情感计算大模型评测心理支持共情生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。