用大模型生成多参考答案,让语音断句评估更准更省力
LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations

- 基于大模型生成多种合理断句方案,解决单参考的局限性
- 在1356条韩语标注上,与人工判断一致性显著提升
- 适合需要高效、多视角语音评估的研究与工业场景
可靠的短语断点标注评估至关重要,因语调边界细微差异直接影响语音清晰度与自然度。现有方法存在明显缺陷:单参考评估假设每句话仅有一种正确断法,但实际存在多种有效断法;人工评判虽灵活却耗时费力且难以扩展。为此,我们提出基于大模型的多参考评估(LMRE),通过少量示例生成多个合理断句方案,建模语调断法的一对多特性。在包含5种策略的1356条韩语测试数据上,LMRE在接纳行为和评分相关性上均优于单参考评估,证明其兼具可扩展性与多参考支持能力,彰显大模型在语音评估中的潜力。
原文摘要 · Abstract (English)
Reliable evaluation of phrase break annotations is crucial, as subtle variations in prosodic boundaries directly affect the clarity and naturalness of speech. However, existing approaches exhibit major limitations: single-reference evaluation assumes a unique gold phrasing for an utterance despite multiple valid phrasings, while human judgment, though flexible, is labor-intensive and unscalable. To address these, we propose LLM-based Multi-Reference Evaluation (LMRE) for phrase break annotations that models the one-to-many nature of prosodic phrasing and generates multiple valid phrasings from minimal demonstrations. On a Korean testbed of 1,356 annotations covering five strategies, LMRE shows stronger alignment with human judgment than single-reference evaluation in both acceptance behavior and score correlation. Our findings demonstrate that LMRE effectively achieves both scalability and multi-reference support, highlighting the potential of LLMs for evaluation in the speech domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。