arXiv:2606.08129cs.AI2026-06

不同大模型推理时竟有相似内部模式,暗示它们可能被隐式优化到共同路径。

Cross-LLM Consistency in Inference: Evidence from Shared Interactions

论文配图:Cross-LLM Consistency in Inference: Evidence from Shared Interactions
图 1 · 摘自论文原文
  • 用交互解释法分析模型推理过程,发现跨模型共享相似预测模式。
  • 先进模型间共享交互更明显,且多为低阶、正负抵消弱的模式。
  • 适合关注大模型内在一致性的研究者,尤其关心可解释性与泛化机制。

大型语言模型(LLMs)在架构、训练数据和优化方式上存在差异,但仍可能发展出相似的内部推理模式。本文通过基于交互的解释方法检验该假设,发现当多个模型对同一提示预测相同目标词时,常表现出共享的交互模式。这种一致性在先进模型之间更为显著。共享的交互通常为低阶,且正负效应抵消较弱,而非共享交互则相反。结果表明,先进大模型可能被隐式优化至共有的推理路径,但驱动这种跨模型一致性的具体机制仍不明确。

原文摘要 · Abstract (English)

Large language models (LLMs) differ in architecture, training data, and optimization procedures, yet they may still develop similar internal inference patterns. In this paper, we examine this hypothesis using interaction-based explanations. We find that LLMs often share interaction patterns when predicting the same target token from the same prompt. This consistency is more pronounced among advanced LLMs. Shared interactions also tend to be lower-order and show weaker positive-negative cancellation than non-shared interactions. These results suggest that advanced LLMs may be implicitly optimized toward common inference patterns, even though the mechanisms that give rise to such cross-model consistency remain open.

大模型推理一致性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。