用外部搜索证据迭代修正新闻摘要中的幻觉,提升事实准确性。
Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge
- 通过多轮提问-验证-修正机制,利用搜索引擎获取真实证据
- 使用三个搜索引擎的片段,使摘要错误率显著降低
- 少量示例提示+搜索结果可有效对齐人工评估标准
尽管大型语言模型(LLMs)在生成连贯文本方面表现优异,但仍存在幻觉问题——即生成与事实不符的内容。针对此问题,自修正方法展现出良好前景:它们利用大模型的多轮交互能力,通过生成验证性问题、结合内部或外部知识回答,并据此修正原始输出。此类方法已在百科类生成中被探索,但在新闻摘要等更复杂领域研究较少。本文将两种先进的自修正系统应用于新闻摘要纠错任务,基于三个搜索引擎的证据进行测试。实验分析揭示了搜索片段和少量示例提示的有效性,同时发现G-Eval评分与人工评估高度一致,表明该方法具备可靠评估潜力。
原文摘要 · Abstract (English)
While large language models (LLMs) have shown remarkable capabilities to generate coherent text, they suffer from the issue of hallucinations -- factually inaccurate statements. Among numerous approaches to tackle hallucinations, especially promising are the self-correcting methods. They leverage the multi-turn nature of LLMs to iteratively generate verification questions inquiring additional evidence, answer them with internal or external knowledge, and use that to refine the original response with the new corrections. These methods have been explored for encyclopedic generation, but less so for domains like news summarization. In this work, we investigate two state-of-the-art self-correcting systems by applying them to correct hallucinated summaries using evidence from three search engines. We analyze the results and provide insights into systems' performance, revealing interesting practical findings on the benefits of search engine snippets and few-shot prompts, as well as high alignment of G-Eval and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。