让AI根据问题精准选择图像观察区域,提升视觉推理准确率
Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning

- 基于问题生成条件化视觉观测点,动态规划查看位置
- 在19%图像区域下达0.833准确率,接近全图推理效果
- 适合需要精确定位证据的视觉问答任务
高分辨率像素和裁剪/缩放工具使多模态大语言模型具备图像细看能力,但缺乏可靠的任务驱动决策机制。Q-CueGraph 明确该决策过程:将问题与图像表示映射为受限的坐标级观察点。文本丰富的图像使用可复用的OCR/布局图;自然图像搜索则在相同选择、组合与预算接口下生成查询条件化的视觉节点。可选的效用精炼通过训练-答案正确性学习冻结阅读器可利用的候选裁剪区域,无需区域框监督。在冻结的 Qwen2.5-VL-7B 阅读器下,Q-CueGraph 在 V*Bench 上达到 0.833 准确率,而全图推理仅需 19% 图像面积预算时准确率为 0.696;在 InfographicVQA 上实现全图 ANLS 的 92%,仅使用约一半图像面积。在六个基准测试中,当证据可定位、问题能区分位置且全图阅读受分辨率限制时,显式观测最有效。
原文摘要 · Abstract (English)
High-resolution pixels and crop or zoom tools give multimodal large language models the ability to inspect an image, but they do not provide a reliable task-conditioned policy for deciding where to inspect. Q-CueGraph makes this decision explicit. It maps a question and an image representation to budgeted, coordinate-level observations for a frozen reader. Text-rich images use a reusable OCR/layout graph; natural-image search instantiates query-conditioned visual nodes behind the same selection, composition, and budgeting interface. Optional utility refinement learns which candidate crops the frozen reader can use from training-answer correctness, without region-box supervision. With a frozen Qwen2.5-VL-7B reader, Q-CueGraph reaches 0.833 accuracy on V*Bench versus 0.696 for full-image inference from a 19% image-area budget, and reaches 92% of full-image ANLS on InfographicVQA from about half the image area. Across six benchmarks, explicit observation is most valuable when evidence is localizable, the question discriminates its location, and resolution limits full-image reading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。