arXiv:2604.17958eess.AScs.SD2026-04被引 4

首个多语言指令跟随语音生成评测基准,全面评估可控语音合成能力。

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

论文配图:MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
图 1 · 摘自论文原文
  • 构建多轴分类体系与分阶段数据流水线,支持多语言指令生成评测。
  • 跨十种语言测试显示主流系统仍存差距,中文场景下开源模型可超商用系统。
  • 聚焦复合指令与副语言控制难点,适合研究可控语音合成的开发者使用。

指令跟随文本转语音(TTS)已成为实现可控且富有表现力语音生成的重要能力,但其评估仍因基准覆盖有限、诊断粒度弱及多语言支持不足而发展滞后。本文提出MINT-Bench,一个全面的多语言指令跟随TTS评测基准。MINT-Bench基于层级多轴分类体系、可扩展的多阶段数据构建流程,以及分层混合评估协议,联合评估内容一致性、指令遵循度与感知质量。在十种语言上的实验表明,当前系统仍未解决:前沿商用系统整体领先,而领先开源模型在局部场景(如中文)已具备竞争力,甚至超越商用方案。该基准进一步揭示,复杂组合与副语言控制仍是现有系统的主要瓶颈。我们发布MINT-Bench及其数据构建与评估工具包,以支持未来可控、多语言及诊断性扎实的TTS评估研究。排行榜与演示地址:https://aslp-lab.github.io/MINT-Bench-Demo/

原文摘要 · Abstract (English)

Instruction-following text-to-speech (TTS) has emerged as an important capability for controllable and expressive speech generation, yet its evaluation remains underdeveloped due to limited benchmark coverage, weak diagnostic granularity, and insufficient multilingual support. We present \textbf{MINT-Bench}, a comprehensive multilingual benchmark for instruction-following TTS. MINT-Bench is built upon a hierarchical multi-axis taxonomy, a scalable multi-stage data construction pipeline, and a hierarchical hybrid evaluation protocol that jointly assesses content consistency, instruction following, and perceptual quality. Experiments across ten languages show that current systems remain far from solved: frontier commercial systems lead overall, while leading open-source models become highly competitive and can even outperform commercial counterparts in localized settings such as Chinese. The benchmark further reveals that harder compositional and paralinguistic controls remain major bottlenecks for current systems. We release MINT-Bench together with the data construction and evaluation toolkit to support future research on controllable, multilingual, and diagnostically grounded TTS evaluation. The leaderboard and demo are available at https://aslp-lab.github.io/MINT-Bench-Demo/

语音生成多语言指令跟随评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。