提升文档生成播客的忠实度,解决长对话中信息失真问题。
On Improving Faithfulness of Podcasts from Documents

- 提出逐轮评估框架,检测对话是否基于原文
- 实测顶级模型如GPT-4o仍频繁出现无依据内容
- 设计通用修复框架,保持流畅性同时修正错误
大型语言模型(LLMs)正被广泛用于从文本源生成长篇对话式内容,如播客。尽管输出流畅生动,但常引入无依据信息。本文首次系统研究文档驱动播客生成中的忠实度问题,要求在多角色、长对话中持续保持与原文一致。我们构建了涵盖五个领域的1500余份文档数据集,并用多个LLM生成播客转录文本。提出一种逐轮级的LLM作为裁判评估框架,判断每轮对话是否由源文档支持,并通过人工验证其可靠性。分析显示,即使最先进的模型如GPT-4o也常产生不忠实内容。为此,我们提出“catch-n-repair”框架,无需依赖特定模型,可检测并重写不忠实对话轮次,同时保持对话自然流畅。实验表明,在域内与域外设置下均显著提升忠实度。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to generate long-form conversational content such as podcasts from textual sources. While these systems produce fluent and engaging narratives, they often introduce ungrounded information. In this work, we present the first systematic study of faithfulness in document-grounded podcast generation, where grounding must be maintained across conversational turns in long-form, multi-speaker transcripts. We construct a dataset of over 1500 documents spanning five domains and generate podcast transcripts using multiple LLMs. We introduce a turn-level LLM-as-a-judge framework for evaluating whether conversational turns are supported by the source document, and validate its reliability through human studies. Our analysis shows that even state-of-the-art models, including GPT-4o, frequently generate ungrounded content. To mitigate this issue, we propose catch-n-repair, a model-agnostic framework that detects and rewrites unfaithful conversational turns while preserving conversational flow. Experiments demonstrate consistent improvements in faithfulness across both in-domain and out-of-domain settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。