arXiv:2601.05009cs.AI2026-01被引 1

大模型在表格扭曲下易出错,需显式提示才可部分纠正。

An Empirical Investigation of Robustness in Large Language Models under Tabular Distortions

  • 通过专家构建的测试集评估模型对表格错误的修正能力。
  • 即使顶尖模型如GPT-5.2在扭曲下准确率最低下降22%。
  • 提醒研究者关注模型自主纠错机制的设计问题。

我们研究大型语言模型(LLMs)在表格数据受到语义与结构扭曲时的鲁棒性表现。结果表明,模型缺乏内在检测和纠正表格表示中细微错误的能力。仅当通过系统提示提供显式先验知识时,模型才能部分调整推理策略并修正部分错误,但效果不一致且不完整。为此,我们构建了一个小型、专家标注的数据集,专门用于评估模型在需要额外纠错步骤的表格问答任务中的表现。结果显示,不同模型在扭曲条件下表现出系统性差异,即使是当前最先进的模型如GPT-5.2,准确率也至少下降22%。这些发现引发重要思考:未来研究应探索模型何时以及如何像人类一样自主重对齐表格输入,而无需依赖显式提示或预处理。

原文摘要 · Abstract (English)

We investigate how large language models (LLMs) fail when tabular data in an otherwise canonical representation is subjected to semantic and structural distortions. Our findings reveal that LLMs lack an inherent ability to detect and correct subtle distortions in table representations. Only when provided with an explicit prior, via a system prompt, do models partially adjust their reasoning strategies and correct some distortions, though not consistently or completely. To study this phenomenon, we introduce a small, expert-curated dataset that explicitly evaluates LLMs on table question answering (TQA) tasks requiring an additional error-correction step prior to analysis. Our results reveal systematic differences in how LLMs ingest and interpret tabular information under distortion, with even SoTA models such as GPT-5.2 model exhibiting a drop of minimum 22% accuracy under distortion. These findings raise important questions for future research, particularly regarding when and how models should autonomously decide to realign tabular inputs, analogous to human behavior, without relying on explicit prompts or tabular data pre-processing.

大模型表格理解鲁棒性误差纠正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。