arXiv:2604.23733cs.CL2026-04

让模型从科学图表中提出有深度的疑问,提升科研理解能力。

Multimodal QUD: Inquisitive Questions from Scientific Figures

论文配图:Multimodal QUD: Inquisitive Questions from Scientific Figures
图 1 · 摘自论文原文
  • 基于图文上下文,识别图表引发的未解问题。
  • 在1250个问题数据集上,模型学会提出与论文论点相关的疑问。
  • 适合科研助手、AI辅助写作等需要深度理解文献的场景。

复杂文档中的话语理解常涉及持续提出并解答讨论中的问题(QUD)。尽管现有QUD框架多聚焦文本,科学文献本质上是多模态的:图表传达的论述目标不同于文字,会引发周围文本所回答的隐含问题。在科学发现中,知道该问什么和如何回答同样重要,但当前模型缺乏此能力。本文将QUD扩展至科学文献的多模态话语,聚焦由图表引发的三类问题:(1) 带有探究性的,即前文未解决;(2) 显著的,与论文研究主张相关且后续被回应;(3) 基于视觉洞察的。为评估模型生成此类问题的能力,我们构建了MQUD数据集,包含来自56篇论文的1,250个图表引发的问题,其中708个由原作者标注。实验表明,开源视觉语言模型如Qwen 3.5主要提出仅凭图像即可回答的问题。在MQUD上微调后,模型能生成更具科学相关性、探究性的提问,聚焦图表在论证中的作用。

原文摘要 · Abstract (English)

Discourse comprehension in complex documents often involves continuously posing and resolving Questions Under Discussion (QUDs). While QUD frameworks have so far focused on text, scientific literature is inherently multimodal: figures convey discourse goals distinct from their textual counterparts, thus invoking implicit questions that the surrounding text answers. In scientific discovery, knowing the right questions to ask is as important as knowing how to answer them, yet this capability remains largely absent in current models. In this work, we extend QUD to multimodal discourse in scientific literature, targeting questions evoked by figures that are (1) inquisitive, i.e., not resolved in the prior context; (2) salient, i.e., relevant to the paper's research claims and addressed later in the paper; (3) grounded in visual insights. To benchmark model capability to generate such questions, we introduce MQUD, a dataset of 1,250 figure-evoked questions from 56 scientific papers, including 708 questions annotated by the original authors of these papers. Our experiments show that open-source VLMs such as Qwen 3.5 predominantly ask questions that can be answered by the figure alone. Fine-tuning on MQUD teaches the model to ask scientifically relevant, inquisitive questions that target the role a figure plays in the paper's argument.

多模态科学文献问答系统视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。