arXiv:2505.09039cs.CL2025-05被引 3

无需外部模型,通过一致性信号提升长文本问答的准确性

Atomic Consistency Preference Optimization for Long-Form Question Answering

  • 利用多次生成结果中事实的一致性作为自我监督信号
  • 在LongFact和BioGen数据集上比有监督基线高1.95分
  • 适合缺乏外部知识源的长文本问答场景

大语言模型常产生看似合理但错误的事实幻觉。现有缓解方法依赖强模型(如GPT-4)或外部知识库判断事实正确性,但这些资源并不总可得。为此,我们提出原子一致性偏好优化(ACPO),一种完全自监督的偏好微调方法,无需外部监督即可提升事实准确性。ACPO通过多个随机生成结果中单个事实的一致性信号,识别高质量与低质量的数据对用于模型对齐。尽管完全自监督,ACPO在Phi-3和Llama3模型上于LongFact和BioGen数据集上的平均表现优于强监督基线1.95分,证明其在不依赖外部模型或知识库的情况下有效提升事实可靠性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often produce factoid hallucinations - plausible yet incorrect answers. A common mitigation strategy is model alignment, which improves factual accuracy by training on curated (factual, non-factual) pairs. However, this approach often relies on a stronger model (e.g., GPT-4) or an external knowledge base to assess factual correctness that may not always be accessible. Addressing this, we propose Atomic Consistency Preference Optimization (ACPO), a self-supervised preference-tuning method that enhances factual accuracy without external supervision. ACPO leverages atomic consistency signals (i.e., the agreement of individual facts across multiple stochastic responses) to identify high- and low-quality data pairs for model alignment. Despite being fully self-supervised, ACPO outperforms the strong supervised alignment baseline by 1.95 points averaged across Phi-3 and Llama3 on the LongFact and BioGen datasets, demonstrating its effectiveness in improving factual reliability without relying on external models or knowledge bases.

长文本问答事实准确性自监督偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。