arXiv:2510.06927eess.AS2025-10中稿 · ICML被引 2

提出负责任评估框架,提升TTS技术评价的全面性与伦理导向

Position: Towards Responsible Evaluation for Text-to-Speech

  • 构建三层次评估体系:能力真实反映、标准可比、治理公平安全
  • 强调需改进主客观评分方法以更准确揭示模型优劣
  • 适合关注AI伦理、语音合成安全与政策制定的研究者

近年来文本转语音(TTS)技术进步显著,生成语音已接近人类水平,在无障碍、内容创作和人机交互中带来诸多益处。然而,当前评估方式难以全面捕捉现代TTS系统的能力、局限及社会影响。本文提出“负责任评估”概念,主张其对下一阶段TTS发展至关重要,包含三个递进层面:(1) 通过更稳健、区分度高、全面的客观与主观评分方法,真实反映模型能力与限制;(2) 借助标准化基准、透明报告和可迁移评估指标,实现可比性、标准化与可转移性;(3) 评估数据溯源、偏见、滥用、欺骗和可追溯性等治理、公平与安全问题。本文批判现有评估实践,识别系统性缺陷,并提出可操作建议,旨在推动更可靠且符合伦理、服务社会的TTS技术发展。

原文摘要 · Abstract (English)

Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, content creation, and human-computer interaction. However, current evaluation practices are increasingly inadequate for capturing the full range of capabilities, limitations, and societal impacts of modern TTS systems. This position paper introduces the concept of Responsible Evaluation and argues that it is essential and urgent for the next phase of TTS development, structured through three progressive levels: (1) ensuring the faithful and accurate reflection of a model's true capabilities and limitations, with more robust, discriminative, and comprehensive objective and subjective scoring methodologies; (2) enabling comparability, standardization, and transferability through standardized benchmarks, transparent reporting, and transferable evaluation metrics; and (3) assessing governance, fairness, and security concerns around data provenance, disparities, misuse, spoofing, and traceability. Through this concept, we critically examine current evaluation practices, identify systemic shortcomings, and propose actionable recommendations. We hope this concept will not only foster more reliable TTS technology but also guide its development toward ethically sound and societally beneficial applications.

TTS评估人工智能伦理语音合成负责任AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。