arXiv:2505.20487cs.CLcs.AI2025-05EMNLP被引 4

让大模型生成更准确且信息量更高的事实性内容

InFact: Informativeness Alignment for Improved LLM Factuality

  • 引入信息量对齐机制,优先选择正确且信息丰富的回答
  • 在多个基准上同时提升事实准确率和信息量
  • 适合需要高可信度与细节的问答系统开发者

事实完整性指文本在事实正确的基础上,其详尽程度与信息丰富度。例如,“巴拉克·奥巴马出生于美国”虽正确,但信息量低于“巴拉克·奥巴马出生于夏威夷州火奴鲁鲁市,美国”。尽管大模型常产生错误内容,也可能生成正确但信息贫乏的表述。本文提出信息量对齐机制,利用近期事实性评测数据构建目标函数,使模型在保证正确性的同时优先生成更具信息量的回答。关键发现:训练模型以最大化该目标或优化偏好时,不仅能提升信息量,还能同步增强事实准确性。

原文摘要 · Abstract (English)

Factual completeness is a general term that captures how detailed and informative a factually correct text is. For instance, the factual sentence ``Barack Obama was born in the United States'' is factually correct, though less informative than the factual sentence ``Barack Obama was born in Honolulu, Hawaii, United States''. Despite the known fact that LLMs tend to hallucinate and generate factually incorrect text, they might also tend to choose to generate factual text that is indeed factually correct and yet less informative than other, more informative choices. In this work, we tackle this problem by proposing an informativeness alignment mechanism. This mechanism takes advantage of recent factual benchmarks to propose an informativeness alignment objective. This objective prioritizes answers that are both correct and informative. A key finding of our work is that when training a model to maximize this objective or optimize its preference, we can improve not just informativeness but also factuality.

大模型事实性信息量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。