arXiv:2512.06726cs.CVcs.AI2025-12

通过调控熵提升视觉定位任务表现,平衡探索与利用。

The Role of Entropy in Visual Grounding: Analysis and Optimization

  • 提出可解释的熵控制算法ECVGPO,优化视觉定位策略
  • 在多个基准上实现模型性能提升,验证熵控有效性
  • 适合关注多模态模型训练效率与稳定性的研究者

最近基于强化学习微调多模态大语言模型(MLLMs)取得了显著进展,尤其得益于各类熵控制技术的引入。然而,熵在感知类任务如视觉定位中的作用及特性,以及有效的控制策略仍不明确。为此,本文聚焦于视觉定位任务,对比分析其与推理任务中熵的作用与特征。基于上述发现,我们提出ECVGPO(Entropy Control Visual Grounding Policy Optimization),一种可解释的熵调控算法,能更好平衡探索与利用的权衡。实验表明,该方法在多种基准和模型上均取得广泛性能提升。

原文摘要 · Abstract (English)

Recent advances in fine-tuning multimodal large language models (MLLMs) using reinforcement learning have achieved remarkable progress, particularly with the introduction of various entropy control techniques. However, the role and characteristics of entropy in perception-oriented tasks like visual grounding, as well as effective strategies for controlling it, remain largely unexplored. To address this issue, we focus on the visual grounding task and analyze the role and characteristics of entropy in comparison to reasoning tasks. Building on these findings, we introduce ECVGPO (Entropy Control Visual Grounding Policy Optimization), an interpretable algorithm designed for effective entropy regulation. Through entropy control, the trade-off between exploration and exploitation is better balanced. Experiments show that ECVGPO achieves broad improvements across various benchmarks and models.

视觉定位熵控制多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。