根据任务需求动态调整图像分辨率,提升视觉大模型表现
Task-Aware Resolution Optimization for Visual Large Language Models
- 基于图像复杂度和模型不确定性,建立分辨率优化公式
- 在多个任务上验证,分辨率自适应使性能显著提升
- 无需重训练,轻量微调即可扩展模型输入分辨率
真实场景的视觉语言应用对感知粒度要求各异。现有视觉大模型(如LLaVA)普遍采用固定输入分辨率,导致性能不佳。本文首次系统研究不同任务的分辨率偏好,发现其与图像复杂度及模型在不同分辨率下的不确定性方差相关。据此提出经验公式,综合两项因素确定最优分辨率。进一步设计一种参数高效微调方法,将预训练模型的视觉输入分辨率扩展至最优值。大量实验验证了该方法在多类视觉语言任务上的有效性。
原文摘要 · Abstract (English)
Real-world vision-language applications demand varying levels of perceptual granularity. However, most existing visual large language models (VLLMs), such as LLaVA, pre-assume a fixed resolution for downstream tasks, which leads to subpar performance. To address this problem, we first conduct a comprehensive and pioneering investigation into the resolution preferences of different vision-language tasks, revealing a correlation between resolution preferences with image complexity, and uncertainty variance of the VLLM at different image input resolutions. Building on this insight, we propose an empirical formula to determine the optimal resolution for a given vision-language task, combining these two factors. Second, based on rigorous experiments, we propose a novel parameter-efficient fine-tuning technique to extend the visual input resolution of pre-trained VLLMs to the identified optimal resolution. Extensive experiments on various vision-language tasks validate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。