评测中文大模型的人性化智能,发现顶尖模型仅达专家标准的60%
HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
- 基于心理咨询场景构建五维评估体系,用评分规则量化情感文化能力
- 13个主流模型平均仅达专家理想分的60%,复杂情绪与伦理难题下性能骤降
- 适合关注AI人文对齐、社会情感计算的研究者与开发者
尽管大语言模型在认知与推理任务中表现卓越,但在处理社会、情感与伦理等人性化智能方面仍存在明显短板,尤其在中文语境下,因缺乏专用评估框架与高质量社情数据而进展受限。为此,我们提出HeartBench,一个面向中文大模型的情感、文化与伦理综合评估框架。该框架基于真实心理咨询场景,由临床专家参与设计,包含五个主维度与十五项次级能力,采用‘先推理后评分’的案例特定评分机制,将抽象的人类特质转化为可测量标准。对13个先进模型的评估显示,其性能上限显著:即使领先模型也仅达到专家定义理想分数的60%。进一步通过难度分层的‘Hard Set’分析发现,在涉及微妙情感隐含意义与复杂伦理权衡的情境中,模型表现出现明显下降。HeartBench为人性化智能评估提供了标准化指标,并为构建高质量、人类对齐的训练数据提供了方法论蓝图。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have achieved remarkable success in cognitive and reasoning benchmarks, they exhibit a persistent deficit in anthropomorphic intelligence-the capacity to navigate complex social, emotional, and ethical nuances. This gap is particularly acute in the Chinese linguistic and cultural context, where a lack of specialized evaluation frameworks and high-quality socio-emotional data impedes progress. To address these limitations, we present HeartBench, a framework designed to evaluate the integrated emotional, cultural, and ethical dimensions of Chinese LLMs. Grounded in authentic psychological counseling scenarios and developed in collaboration with clinical experts, the benchmark is structured around a theory-driven taxonomy comprising five primary dimensions and 15 secondary capabilities. We implement a case-specific, rubric-based methodology that translates abstract human-like traits into granular, measurable criteria through a ``reasoning-before-scoring'' evaluation protocol. Our assessment of 13 state-of-the-art LLMs indicates a substantial performance ceiling: even leading models achieve only 60% of the expert-defined ideal score. Furthermore, analysis using a difficulty-stratified ``Hard Set'' reveals a significant performance decay in scenarios involving subtle emotional subtexts and complex ethical trade-offs. HeartBench establishes a standardized metric for anthropomorphic AI evaluation and provides a methodological blueprint for constructing high-quality, human-aligned training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。