arXiv:2510.16815cs.CLcs.AI2025-10Conference of the …被引 2

大模型比较实体时常走捷径,而非依赖真实知识。

Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities

  • 通过数值属性比较任务,发现模型受流行度、顺序等表面线索影响
  • 32B参数大模型能判断何时用真实知识,小模型则无此能力
  • 思维链提示可让所有模型转向使用真实数据,提升准确性

大型语言模型(LLMs)被广泛用于基于知识的推理任务,但其在何时依赖真实知识、何时依赖表面启发式仍不清晰。我们通过实体数值比较任务(如“多瑙河与尼罗河哪条更长?”)进行系统分析,这些任务有明确的真值。尽管模型具备足够的数值知识,却常做出违背该知识的判断。我们识别出三种强影响模型决策的启发式偏差:实体流行度、提及顺序和语义共现。对小模型(7–8B参数),仅用这些表面线索的简单逻辑回归,比模型自身数值预测更准确,表明启发式完全覆盖了理性推理。关键发现是,大模型(32B参数)会根据可靠性选择性地使用数值知识,而小模型无此区分能力,这解释了为何大模型表现更好,即使小模型拥有更准确的知识。思维链提示可引导所有模型使用数值特征,显著提升性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used for knowledge-based reasoning tasks, yet understanding when they rely on genuine knowledge versus superficial heuristics remains challenging. We investigate this question through entity comparison tasks by asking models to compare entities along numerical attributes (e.g., ``Which river is longer, the Danube or the Nile?''), which offer clear ground truth for systematic analysis. Despite having sufficient numerical knowledge to answer correctly, LLMs frequently make predictions that contradict this knowledge. We identify three heuristic biases that strongly influence model predictions: entity popularity, mention order, and semantic co-occurrence. For smaller models, a simple logistic regression using only these surface cues predicts model choices more accurately than the model's own numerical predictions, suggesting heuristics largely override principled reasoning. Crucially, we find that larger models (32B parameters) selectively rely on numerical knowledge when it is more reliable, while smaller models (7--8B parameters) show no such discrimination, which explains why larger models outperform smaller ones even when the smaller models possess more accurate knowledge. Chain-of-thought prompting steers all models towards using the numerical features across all model sizes.

大模型知识推理启发式偏差思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。