arXiv:2509.20065cs.CL2025-09EMNLP被引 2

通过输入特征提前发现语言模型的理解盲区。

From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors

  • 仅用输入的词元级似然特征预测理解偏差。
  • 在五个语言挑战数据集上优于基准方法。
  • 轻量高效,适合大中小模型预判错误。

语言模型常因对习语、隐喻或上下文敏感输入的初始误解而表现不佳,而非输出错误。本文提出一种仅依赖输入的方法,利用受意外性与均匀信息密度假设启发的词元级似然特征,捕捉输入理解中的局部不确定性。该方法在五个语言挑战性数据集上表现优异,且局部特征对大模型更有效,全局模式则利于小模型。无需输出或隐藏激活,具备轻量与通用性,可实现生成前的错误预测。

原文摘要 · Abstract (English)

Language models often struggle with idiomatic, figurative, or context-sensitive inputs, not because they produce flawed outputs, but because they misinterpret the input from the outset. We propose an input-only method for anticipating such failures using token-level likelihood features inspired by surprisal and the Uniform Information Density hypothesis. These features capture localized uncertainty in input comprehension and outperform standard baselines across five linguistically challenging datasets. We show that span-localized features improve error detection for larger models, while smaller models benefit from global patterns. Our method requires no access to outputs or hidden activations, offering a lightweight and generalizable approach to pre-generation error prediction.

模型误差输入感知语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。