评测大模型隐含表达能力,发现其在情感上表现好但社会语言信号弱。
ExpressivityBench: Can LLMs Communicate Implicitly?
- 用信息论模型量化文本隐含传递属性的能力
- 九项任务中情感表达较好,社会语言信号显著落后于人
- 适合关注对话系统、心理支持等场景的研究者
人类交流常具隐含性,传递语气、身份和意图等超越字面意义的信息。尽管大语言模型在摘要、推理等显式任务上表现优异,其表达力(即隐含沟通能力)仍缺乏深入研究。本文提出「ExpressivityBench」,一个基于信息论通信模型的评估框架,用于衡量大模型生成文本在未明确提及目标属性时的传达效果,涵盖情绪、身份与语调共九项任务。为实现可扩展且可复现的评估,采用经人类判断验证的基于LLM的评分器。结果表明,模型在表达情感内容方面表现良好,但在社会语言信号方面明显落后于人类基准。该研究为评估类人隐含沟通能力提供了必要基础,对教育、心理健康支持及社交感知对话系统具有重要意义。代码与数据已随论文公开。
原文摘要 · Abstract (English)
Human communication is often implicit, conveying tone, identity, and intent beyond literal meanings. While large language models have achieved strong performance on explicit tasks such as summarization and reasoning, their capacity for expressivity, or implicit communication, remains underexplored. We introduce \textbf{ExpressivityBench}, a framework for evaluating the expressivity of LLMs using information-theoretic communication models. Our approach quantifies how well LLM-generated text communicates target properties without explicit mention, across nine tasks spanning emotion, identity, and tone. To enable scalable and reproducible evaluation, we employ LLM-based graders validated against human judgments. Our results reveal that while models are adept at expressing affective content, they struggle with sociolinguistic signals, lagging behind human baselines. This study provides a necessary step to evaluate human-like implicit communication, with implications for applications such as education, mental health support, and socially-aware dialogue systems. We provide code and data for our benchmark alongside our paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。