构建首个播客音频生成评估框架,解决长文本生成无标准评价难题
PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
- 分文本、语音、音频三维度设计多模态评估策略
- 基于真实播客数据集,实现内容与格式双重点评估
- 开源框架支持自研模型与商业系统对比评测
近年来,越来越多的多模态(文本与音频)基准测试出现,主要聚焦于模型理解能力的评估。然而,针对生成能力的探索仍有限,尤其在开放式长时内容生成方面。主要挑战包括缺乏参考标准答案、无统一评价指标以及难以控制的人类评判。本文以播客类音频生成为切入点,提出 PodEval——一个全面且精心设计的开源评估框架。该框架包含:1)构建涵盖多种主题的真实播客数据集,作为人类创作质量的参考基准;2)引入多模态评估策略,将复杂任务分解为文本、语音和音频三个维度,并对“内容”与“格式”分别设定不同评估重点;3)针对每种模态设计相应的评估方法,结合客观指标与主观听觉测试。我们在实验中使用代表性播客生成系统(包括开源、闭源及人工生成)进行验证。结果提供了对播客生成的深入分析与洞见,证明了 PodEval 在评估开放式长时音频生成中的有效性。本项目已开源,欢迎使用:https://github.com/yujxx/PodEval。
原文摘要 · Abstract (English)
Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remains limited, especially for open-ended long-form content generation. Significant challenges lie in no reference standard answer, no unified evaluation metrics and uncontrollable human judgments. In this work, we take podcast-like audio generation as a starting point and propose PodEval, a comprehensive and well-designed open-source evaluation framework. In this framework: 1) We construct a real-world podcast dataset spanning diverse topics, serving as a reference for human-level creative quality. 2) We introduce a multimodal evaluation strategy and decompose the complex task into three dimensions: text, speech and audio, with different evaluation emphasis on "Content" and "Format". 3) For each modality, we design corresponding evaluation methods, involving both objective metrics and subjective listening test. We leverage representative podcast generation systems (including open-source, close-source, and human-made) in our experiments. The results offer in-depth analysis and insights into podcast generation, demonstrating the effectiveness of PodEval in evaluating open-ended long-form audio. This project is open-source to facilitate public use: https://github.com/yujxx/PodEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。