评测多模态大模型的共情能力,聚焦语音与文本融合的回应生成与判断。
AEQ-Bench: Measuring Empathy of Omni-Modal Large Models
- 构建双模态共情评测基准,评估模型理解语音+文本情感线索的能力。
- 有语音输出能力的模型表现优于仅支持文本输出的模型。
- 模型在粗粒度评价上接近人类,但在细微语气表达判断上仍不可靠。
尽管自动评估多模态大模型(OLMs)至关重要,但因其内在的情感属性,共情评估仍面临重大挑战。为此,我们提出AEQ-Bench(音频共情商基准),一个系统性评估OLMs两种核心共情能力的新基准:(i) 通过理解多模态输入(音频+文本)中的情感线索生成共情回应;(ii) 在无需文本转录的前提下判断音频回应的共情程度。相较于现有基准,AEQ-Bench引入两种新设置,分别在上下文特异性与语调变化上进行区分。通过语言与副语言指标的综合评估发现:(1) 具备音频输出能力的OLMs普遍优于仅支持文本输出的模型;(2) 虽然OLMs在粗粒度质量评估上与人类判断一致,但在细粒度副语言表达力评估上仍不可靠。
原文摘要 · Abstract (English)
While the automatic evaluation of omni-modal large models (OLMs) is essential, assessing empathy remains a significant challenge due to its inherent affectivity. To investigate this challenge, we introduce AEQ-Bench (Audio Empathy Quotient Benchmark), a novel benchmark to systematically assess two core empathetic capabilities of OLMs: (i) generating empathetic responses by comprehending affective cues from multi-modal inputs (audio + text), and (ii) judging the empathy of audio responses without relying on text transcription. Compared to existing benchmarks, AEQ-Bench incorporates two novel settings that vary in context specificity and speech tone. Comprehensive assessment across linguistic and paralinguistic metrics reveals that (1) OLMs trained with audio output capabilities generally outperformed models with text-only outputs, and (2) while OLMs align with human judgments for coarse-grained quality assessment, they remain unreliable for evaluating fine-grained paralinguistic expressiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。