arXiv:2601.14903cs.CL2026-01ACL被引 1

构建音频播客脚本生成的综合评测基准,评估模型对复杂指令的理解能力。

PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation

  • 设计包含800个样本的多说话人指令数据集,支持长达21K token输入。
  • 发现开源模型经显式推理训练后,在长文本和多人对话中表现更稳健。
  • 揭示高指令遵循度不等于内容丰富性,为评估提供新视角。

播客脚本生成要求大语言模型从多样化输入中合成结构化、上下文相关的对话,但该任务缺乏系统性评估资源。为此,我们提出PodBench,一个包含800个样本的基准,输入长度可达21K tokens,涵盖复杂的多说话人指令。我们设计了融合定量约束与大语言模型评估的多维度评价框架。大量实验表明,尽管专有模型整体表现更优,但配备显式推理的开源模型在处理长上下文和多说话人协调方面优于标准基线。然而分析发现,高指令遵循度并不保证内容充实性。PodBench为长时序、以音频为中心的生成任务提供了可复现的测试平台。

原文摘要 · Abstract (English)

Podcast script generation requires LLMs to synthesize structured, context-grounded dialogue from diverse inputs, yet systematic evaluation resources for this task remain limited. To bridge this gap, we introduce PodBench, a benchmark comprising 800 samples with inputs up to 21K tokens and complex multi-speaker instructions. We propose a multifaceted evaluation framework that integrates quantitative constraints with LLM-based quality assessment. Extensive experiments reveal that while proprietary models generally excel, open-source models equipped with explicit reasoning demonstrate superior robustness in handling long contexts and multi-speaker coordination compared to standard baselines. However, our analysis uncovers a persistent divergence where high instruction following does not guarantee high content substance. PodBench offers a reproducible testbed to address these challenges in long-form, audio-centric generation.

播客生成指令遵循长文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。