用程序生成任务提升多模态模型的细粒度视觉理解能力
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs

- 通过在图像上叠加几何图形生成密集监督信号
- 在多个基准上实现最高+20%的性能提升
- 适合需要提升空间推理与视觉定位的开发者
尽管多模态大模型取得了显著进展,但在细粒度理解任务上仍表现不佳。本文提出程序生成任务(PGT),一种数据驱动框架,兼具促进细粒度视觉理解与低成本诊断感知失败来源双重功能。通过在图像上叠加清晰的几何图元,PGT生成额外的密集监督信号,将视觉定位能力与语义先验解耦。在关系、量化及3D/深度理解基准上的实验表明,PGT在多种架构中均带来显著提升。在增强后的LLaVA-v1.5-Instruct上进行指令微调,使What'sUp基准提升最高达+20%,CV-Bench-2D提升+13.3%,同时保持通用感知能力。对顶尖MLLMs使用PGT数据微调,可实现What'sUp最高+5.5%、CV-Bench-2D最高+8.3%的提升。结果表明,细粒度感知瓶颈主要源于监督信号不足,而非模型架构或分辨率限制。
原文摘要 · Abstract (English)
Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this work, we propose Procedurally Generated Tasks (PGT), a simple data-driven framework that serves a dual purpose: inducing fine-grained visual understanding and acting as a low-cost diagnostic tool to identify the source of perception failures. By overlaying unambiguous geometric primitives on images, PGT generate additional dense supervision that disentangles visual grounding capability from semantic priors. Extensive experiments on relational, quantitative, and 3D/depth understanding benchmarks show that PGT yields remarkable gains across diverse architectures. Instruction tuning MLLMs on LLaVA-v1.5-Instruct augmented with PGT data results in improvements of up to +20% on the What'sUp benchmark and +13.3% on CV-Bench-2D, while maintaining general perception capabilities. Moreover, finetuning state-of-the-art MLLMs on PGT data leads to boosts of up to +5.5% on What'sUp and +8.3% on CV-Bench-2D. These findings demonstrate that PGT effectively address the bottleneck of fine-grained perception, revealing that many spatial reasoning deficits stem from inadequate supervision signals rather than inherent architectural or resolution limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。