评测模型如何理解科学文档中的高亮等标记并据此推理。
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
- 构建五类任务分解评估标记驱动的表格理解能力。
- 强模型在标记与符号推理对齐时仍表现不稳定。
- 提供可复现基准,帮助定位视觉到执行链中的错误。
高亮、下划线和加粗等视觉标记在以表格为核心的文档中十分常见。尽管多模态大模型在文档理解方面取得显著进展,但其将此类提示视为显式逻辑指令的能力仍未被充分探索。更重要的是,现有评估无法区分模型是未能识别标记,还是无法基于标记进行推理。这导致对标记条件下的表格行为评估存在关键盲区。为此,我们提出HighlightBench,一个诊断性基准,用于标记驱动的表格理解,将评估分解为五类任务:标记定位、约束检索、局部关系、聚合与比较、一致性与缺失性。我们还提供参考流水线,使中间决策过程显式化,支持可复现基线及对感知到执行链中错误的细粒度归因。实验表明,即使强模型在结构化输出约束下,也难以保持标记与符号推理的一致性。
原文摘要 · Abstract (English)
Visual markups such as highlights, underlines, and bold text are common in table-centric documents. Although multimodal large language models (MLLMs) have made substantial progress in document understanding, their ability to treat such cues as explicit logical directives remains under-explored. More importantly, existing evaluations cannot distinguish whether a model fails to see the markup or fails to reason with it. This creates a key blind spot in assessing markup-conditioned behavior over tables. To address this gap, we introduce HighlightBench, a diagnostic benchmark for markup-driven table understanding that decomposes evaluation into five task families: Markup Grounding, Constrained Retrieval, Local Relations, Aggregation \& Comparison, and Consistency \& Missingness. We further provide a reference pipeline that makes intermediate decisions explicit, enabling reproducible baselines and finer-grained attribution of errors along the perception-to-execution chain. Experiments show that even strong models remain unstable when visual cues must be consistently aligned with symbolic reasoning under structured output constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。