arXiv:2602.23615cs.CV2026-02被引 3

无需标注数据,让大模型自动聚焦高分辨率图像关键区域并自我验证。

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework

  • 构建闭环框架,让模型自主定位并验证高分辨率图像中的关键区域。
  • 在多个高分辨率评测集上超越强基线,4K/8K图像任务性能显著提升。
  • 适合需要高效处理高清视觉输入的场景,如医疗影像、遥感分析。

当前大型多模态模型在处理高分辨率视觉输入时存在瓶颈,因图像标记数量随分辨率呈平方增长,带来大量冗余与无关信息。通常做法是识别关键图像区域,并在推理中引用其高分辨率版本,但此类方法依赖外部视觉监督信号,需人工标注,成本高昂。本文提出一种无需标注的高分辨率推理技术HART,通过闭环框架使模型能自主聚焦并自验证关键视觉区域。采用后训练范式,设计优势偏好分组相对策略优化(AP-GRPO),在无外部视觉标注条件下提升模型定位准确性。实验在MME-RealWorld-Lite、TreeBench、V* Bench、HR-Bench-4K/8K和MMStar等多个数据集上验证,HART在多种高分辨率视觉任务中持续优于主流基线,展现出更强的可解释推理路径与高效优化能力。

原文摘要 · Abstract (English)

Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resolution, introducing substantial redundancy and irrelevant information. A common practice is to identify key image regions and refer to their high-resolution counterparts during reasoning, typically trained with external visual supervision. However, such visual supervision cues require costly grounding labels from human annotators. Meanwhile, it remains an open question how to enhance a model's grounding abilities to support reasoning without relying on additional annotations. In this paper, we propose High-resolution Annotation-free Reasoning Technique (HART), a closed-loop framework that enables LMMs to focus on and self-verify key regions of high-resolution visual inputs. HART incorporates a post-training paradigm in which we design Advantage Preference Group Relative Policy Optimization (AP-GRPO) to encourage accurate localization of key regions without external visual annotations. Notably, HART provides explainable reasoning pathways and enables efficient optimization of localization. Extensive experiments on MME-RealWorld-Lite, TreeBench, V* Bench, HR-Bench-4K/8K, and MMStar demonstrate that HART improves performance across a wide range of high-resolution visual tasks, consistently outperforming strong baselines.

多模态模型高分辨率自监督推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。