提出新评估框架,精准判断长段情感语音描述的准确性
EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
- 将复杂描述拆解为独立感知单元,逐个验证
- 与人工评价正相关,传统方法反而因长度敏感呈负相关
- 构建标准化基准数据集,解决资源稀缺问题
近年来语音描述模型已能生成丰富、细粒度的情感语音描述,但其评估仍是关键瓶颈:传统N-gram指标无法捕捉语义细节,大语言模型评判常因推理不一致和上下文坍塌而失准。本文提出EmoSURA,一种从整体评分转向原子验证的新评估框架。该框架将复杂描述分解为自包含的原子感知单元(APUs),即关于音调或情绪属性的独立陈述,并通过音频对齐验证机制,将其与原始语音信号比对。同时,为缓解评估资源稀缺问题,我们构建了SURABench——一个精心平衡且分层采样的基准数据集。实验表明,EmoSURA与人工评价呈现正相关,相比传统指标在长文本评估中表现更可靠;后者因对描述长度敏感,反而出现负相关。
原文摘要 · Abstract (English)
Recent advancements in speech captioning models have enabled the generation of rich, fine-grained captions for emotional speech. However, the evaluation of such captions remains a critical bottleneck: traditional N-gram metrics fail to capture semantic nuances, while LLM judges often suffer from reasoning inconsistency and context-collapse when processing long-form descriptions. In this work, we propose EmoSURA, a novel evaluation framework that shifts the paradigm from holistic scoring to atomic verification. EmoSURA decomposes complex captions into Atomic Perceptual Units, which are self-contained statements regarding vocal or emotional attributes, and employs an audio-grounded verification mechanism to validate each unit against the raw speech signal. Furthermore, we address the scarcity of standardized evaluation resources by introducing SURABench, a carefully balanced and stratified benchmark. Our experiments show that EmoSURA achieves a positive correlation with human judgments, offering a more reliable assessment for long-form captions compared to traditional metrics, which demonstrated negative correlations due to their sensitivity to caption length.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。