arXiv:2603.29396cs.CL2026-03被引 1

用词级困惑度检测大模型是否真懂语言规律

Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity

  • 通过对比关键词微调句的困惑度差异,检验模型是否依赖语言线索
  • 发现模型行为虽受语言重要词影响,却无法完全解释困惑度变化
  • 适合关注模型推理机制、避免误判正确结果的研究者

大语言模型的标准评估侧重任务表现,难以判断其正确行为是否基于恰当机制,易导致确认偏误。本文提出一种基于词级困惑度的简单、严谨可解释性框架,用于检验模型是否依赖语言相关线索。通过比较仅在少数‘关键’词上不同的最小句子对的困惑度分布,该方法实现精准、假设驱动的分析,无需依赖不稳定的特征归因技术。在多个开源大模型和受控语言基准上的实验表明,尽管语言重要词会影响模型行为,但从未能完全解释困惑度变化,揭示模型实际依赖的是非预期的语言启发式策略。

原文摘要 · Abstract (English)

Standard evaluations of Large language models (LLMs) focus on task performance, offering limited insight into whether correct behavior reflects appropriate underlying mechanisms and risking confirmation bias. We introduce a simple, principled interpretability framework based on token-level perplexity to test whether models rely on linguistically relevant cues. By comparing perplexity distributions over minimal sentence pairs differing in one or a few `pivotal' tokens, our method enables precise, hypothesis-driven analysis without relying on unstable feature-attribution techniques. Experiments on controlled linguistic benchmarks with several open-weight LLMs show that, while linguistically important tokens influence model behavior, they never fully explain perplexity shifts, revealing that models rely on heuristics other than the expected linguistic ones.

大模型分析可解释性困惑度语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。