用模型自身不确定性指导视觉任务,无需微调即可提升细粒度感知能力
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- 利用模型响应不确定性自动筛选关键视觉信息
- 在三个复杂任务上达到微调系统水平的性能
- 适合希望快速部署、避免训练的开发者使用
多模态大语言模型在细粒度感知任务中表现不佳,例如高分辨率图像中的小物体识别或长视频的关键时刻检测。现有方法通常依赖复杂的特定任务微调,降低泛化性并增加系统复杂度。本文提出一种无需训练的框架,利用模型内在不确定性作为主动引导。核心思想是:当提供相关视觉信息时,模型不确定性会下降。我们设计统一机制,通过响应不确定性对候选视觉输入打分,使模型自主聚焦于最相关信息。该方法应用于视觉搜索、长视频理解与时间定位三个挑战性任务,使现成的MLLM在性能上媲美专用微调系统。结果表明,利用内在不确定性是提升细粒度多模态表现的有效策略。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long videos. Existing methods typically rely on complex, task-specific fine-tuning, which reduces generalizability and increases system complexity. In this work, we propose an effective, training-free framework that uses an MLLM's intrinsic uncertainty as proactive guidance. Our core insight is that a model's uncertainty decreases when provided with relevant visual information. We introduce a unified mechanism that scores candidate visual inputs by response uncertainty, enabling the model to autonomously focus on the most informative data. We apply this simple principle to three challenging visual tasks: Visual Search, Long Video Understanding, and Temporal Grounding, allowing off-the-shelf MLLMs to achieve performance competitive with specialized, fine-tuned systems. Our results demonstrate that leveraging intrinsic uncertainty is a powerful strategy for improving fine-grained multimodal performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。