arXiv:2508.15690cs.AIcs.LG2025-08被引 2

构建图文对齐的结构化推理基准,评测大模型看图答题能力。

GRAFT: GRaPH and Table Reasoning for Textual Alignment -- A Benchmark for Structured Instruction Following and Visual Reasoning

  • 用程序生成图表和合成表格,配多步分析题
  • 要求模型输出JSON/YAML等结构化结果
  • 涵盖比较、趋势、异常检测等10类推理操作

GRAFT是一个面向结构化指令遵循与视觉推理的多模态基准,旨在评估大语言模型在指令遵循、视觉推理及强图文对齐任务中的表现。数据集基于程序生成的图表与合成渲染的表格,每组图像均配有一道需完全依赖图像信息推断的多步骤分析问题。响应采用结构化格式(如JSON或YAML),支持对推理过程与输出规范符合度的细粒度评估。该基准引入涵盖比较、趋势识别、排名、聚合、比例估算与异常检测等推理操作的分类体系,实现对模型能力的全面评估。GRAFT提供了一个统一且可扩展的框架,为未来多模态大模型在视觉驱动的结构化推理任务中设立更严格的评测标准。

原文摘要 · Abstract (English)

GRAFT is a structured multimodal benchmark designed to probe how well LLMs handle instruction following, visual reasoning, and tasks requiring tight visual textual alignment. The dataset is built around programmatically generated charts and synthetically rendered tables, each paired with a carefully constructed, multi step analytical question that depends solely on what can be inferred from the image itself. Responses are formatted in structured outputs such as JSON or YAML, enabling consistent and fine grained evaluation of both reasoning processes and adherence to output specifications. The benchmark further introduces a taxonomy of reasoning operations ranging from comparison and trend identification to ranking, aggregation, proportional estimation, and anomaly detection to support a comprehensive assessment of model capabilities. Taken together, GRAFT provides a unified and scalable framework for evaluating multimodal LLMs on visually grounded, structured reasoning tasks, offering a more rigorous standard for future benchmarking efforts.

多模态推理评测结构化输出

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。