用多轮思维链让大模型自动修正摘要事实错误
Multi-round, Chain-of-thought Post-editing for Unfaithful Summaries
- 用思维链提示定位并修正摘要与原文的事实偏差
- 多轮编辑使难以一次修正的错误逐步改善,成功率超前人工作
- 无需微调模型,效果媲美专门训练的后编辑系统
近期大型语言模型(LLMs)在自然语言理解与生成任务中表现出色。本文研究了利用 LLMs 评估新闻摘要忠实度的能力,发现其与人工判断具有强相关性。进一步探索了 LLMs 作为忠实度后编辑器的潜力,通过不同思维链提示来定位并修正生成摘要与源新闻文档之间的事实不一致,取得比以往工作更高的编辑成功率。我们进行了自动化与人工评估,结果表明:使用关于事实错误类型的思维链提示,是有效的忠实度后编辑策略,性能可媲美微调的后编辑模型。此外,我们首次证明多轮后编辑可行,能逐步提升那些单轮无法完全纠正的摘要忠实度。
原文摘要 · Abstract (English)
Recent large language models (LLMs) have demonstrated a remarkable ability to perform natural language understanding and generation tasks. In this work, we investigate the use of LLMs for evaluating faithfulness in news summarization, finding that it achieves a strong correlation with human judgments. We further investigate LLMs' capabilities as a faithfulness post-editor, experimenting with different chain-of-thought prompts for locating and correcting factual inconsistencies between a generated summary and the source news document and are able to achieve a higher editing success rate than was reported in prior work. We perform both automated and human evaluations of the post-edited summaries, finding that prompting LLMs using chain-of-thought reasoning about factual error types is an effective faithfulness post-editing strategy, performing comparably to fine-tuned post-editing models. We also demonstrate that multiple rounds of post-editing, which has not previously been explored, can be used to gradually improve the faithfulness of summaries whose errors cannot be fully corrected in a single round.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。