arXiv:2503.19498cs.CL2025-03

构建领域专用图表问答数据集,提升模型深层推理能力。

DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts

  • 按复杂度选图,多层级生成问题,专家审核确保质量。
  • 天文领域生成1690组问答,482张图表,暴露21个大模型短板。
  • 可拓展至生化、经济等多领域,适合评估专业图表理解能力。

图表问答(CQA)用于评估多模态大模型在图表数据上的视觉理解与推理能力。现有基准大多仅测试表面级解析,如读取标签和图例,忽视深层科学推理。我们提出DomainCQA框架,用于构建强调视觉理解与知识密集型推理的领域特定CQA基准。该框架包含复杂度感知的图表选择、多层级问答生成与专家验证。应用于天文学,生成AstroChart数据集,包含1690组问答对和482张图表,揭示21个MLLM在细粒度感知、数值推理与领域知识融合方面的持续缺陷。在AstroChart上微调可提升基础与高级任务表现。生物化学、经济学、医学和社会科学的初步问答集进一步证明DomainCQA的通用性。结果表明,DomainCQA是构建与扩充领域特定图表推理基准的统一流程。

原文摘要 · Abstract (English)

Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test surface-level parsing, such as reading labels and legends, while overlooking deeper scientific reasoning. We propose DomainCQA, a framework for constructing domain-specific CQA benchmarks that emphasize both visual comprehension and knowledge-intensive reasoning. It integrates complexity-aware chart selection, multitier QA generation, and expert validation. Applied to astronomy, DomainCQA yields AstroChart, a benchmark of 1,690 QA pairs over 482 charts, exposing persistent weaknesses in fine-grained perception, numerical reasoning, and domain knowledge integration across 21 MLLMs. Fine-tuning on AstroChart improves performance across fundamental and advanced tasks. Pilot QA sets in biochemistry, economics, medicine, and social science further demonstrate DomainCQA's generality. Together, our results establish DomainCQA as a unified pipeline for constructing and augmenting domain-specific chart reasoning benchmarks.

图表问答多模态领域知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。