arXiv:2511.10667cs.CLcs.AI2025-11被引 1

用专业决策考试评估大模型是否真懂,而非仅猜对答案。

Evaluating LLM Understanding via Structured Tabular Decision Simulations

  • 设计模拟专家决策的表格化测试场景,考察模型真实理解力。
  • 多数模型在多领域表现不稳,准确但常与理由不符。
  • 适合关注大模型可信推理、评测方法研究者。

大型语言模型(LLMs)常表现出令人印象深刻的预测准确性,但正确性本身并不等同于真正理解。真正的理解应像人类专家一样,在多个实例和不同领域中做出一致且有依据的决策,并依赖相关且领域相关的决策因素。我们提出结构化表格决策模拟(STaDS),一套模拟专业人士进行结构化决策“考试”的评估环境。在此设定下,理解被定义为识别并依赖正确决策因素的能力,即决定结果的特征。STaDS从三个方面共同评估理解力:(i) 问题与指令理解,(ii) 基于知识的预测,(iii) 对相关决策因素的依赖。通过对15个多样化决策场景中9个前沿大模型的分析发现:(a) 多数模型难以在多领域保持稳定高准确率;(b) 模型可能准确但全局不忠实,其陈述理由与实际驱动预测的因素常不一致。研究结果凸显了对全局理解能力评估协议的需求,并呼吁超越准确率的新框架,以提升大模型的理解能力。

原文摘要 · Abstract (English)

Large language models (LLMs) often achieve impressive predictive accuracy, yet correctness alone does not imply genuine understanding. True LLM understanding, analogous to human expertise, requires making consistent, well-founded decisions across multiple instances and diverse domains, relying on relevant and domain-grounded decision factors. We introduce Structured Tabular Decision Simulations (STaDS), a suite of expert-like decision settings that evaluate LLMs as if they were professionals undertaking structured decision ``exams''. In this context, understanding is defined as the ability to identify and rely on the correct decision factors, features that determine outcomes within a domain. STaDS jointly assesses understanding through: (i) question and instruction comprehension, (ii) knowledge-based prediction, and (iii) reliance on relevant decision factors. By analyzing 9 frontier LLMs across 15 diverse decision settings, we find that (a) most models struggle to achieve consistently strong accuracy across diverse domains; (b) models can be accurate yet globally unfaithful, and there are frequent mismatches between stated rationales and factors driving predictions. Our findings highlight the need for global-level understanding evaluation protocols and advocate for novel frameworks that go beyond accuracy to enhance LLMs' understanding ability.

模型理解决策评估可信推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。