发现大模型在简短问答和长篇叙述中对同一事实的回答常不一致,暴露其知识可靠性隐患。
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
- 设计对照实验框架SLAQ,比较模型在简单与复杂问题中的事实回答一致性
- 16个模型在600个问题上均出现系统性错位,长题答案准确率下降显著
- 揭示连续回答存在自我强化错误模式,机制相似性可预测一致性达78%
大型语言模型(LLMs)能正确回答‘爱因斯坦何时出生?’这类问题,但在撰写爱因斯坦生平的长篇内容时却可能给出错误日期,暴露出模型在任务复杂度变化下获取事实知识的内在不一致。尽管模型在简单问答基准上表现优异,但其在复杂查询中的可靠性差距仍不明确,削弱了可信度。本文提出短-长形式事实问答对齐评估框架SLAQ,通过对比16个模型在600个问题上的孤立提问(短)与嵌入复杂上下文提问(长)的回答,发现普遍存在的答案错位现象。进一步分析发现,答案准确性受位置影响,并存在自增强的正确或错误模式。机制分析表明,一致的事实会激活重叠的模型内部表征,基于机制相似性的指标可预测短-长答案对齐,最高准确率达78%。研究确立了跨任务复杂度的事实一致性作为模型可信度的关键维度,挑战了当前评价体系中‘简单任务优秀即代表复杂任务可靠’的隐含假设。
原文摘要 · Abstract (English)
Large language models (LLMs) can correctly answer "When was Einstein born?" yet fail to provide the same date when writing about Einstein's life revealing a fundamental inconsistency in how models access factual knowledge across task complexities. While models display impressive accuracy on factual question-answering benchmarks, the reliability gap between simple and complex queries remains poorly understood, eroding their trustworthiness. In this work, we introduce Short-Long Form Alignment for Factual Question Answering (SLAQ), a controlled evaluation framework that compares LLMs' answers to the same factual questions asked (a) in isolation (short) vs. (b) integrated into complex queries (long). Looking at 16 LLMs across 600 queries, we find a systematic misalignment of answers to the corresponding short and long queries. We further uncover position-dependent accuracy loss and momentum effects where consecutive correct or incorrect answers create self-reinforcing patterns. Through mechanistic analysis, we find that aligned facts activate overlapping model internals, and that metrics based on mechanistic similarity can predict short-long answer alignment with up to 78% accuracy. Our work establishes factual consistency over query complexity as an important aspect of LLMs' trustworthiness and challenges current evaluation practices, which implicitly assume that good performance for simple factual queries implies reliability in more complex knowledge-seeking tasks too.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。