arXiv:2604.17504cs.CVcs.AI2026-04

提出混合奖励机制,破解遥感图像理解中的视觉惯性问题

RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding

论文配图:RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding
图 1 · 摘自论文原文
  • 设计三重奖励:空间推理激活、感知正确性、视觉语义路径演化
  • 30亿参数模型在多个任务上超越70亿参数模型,零样本性能领先超3%
  • 适合需要深度推理和泛化能力的遥感图像分析研究者

强化学习后训练能显著提升遥感视觉语言模型性能。但在处理需全面视觉扫描的复杂遥感图像时,模型易依赖局部显著特征快速推断,形成称为‘感知惯性’的偏差。该偏差使模型为最大化奖励而偏好快速拟合,导致认知上过度依赖特定特征,操作上难以灵活转移视觉焦点。为此,我们提出RS-HyRe-R1,一种用于遥感图像理解的混合奖励框架。它包含:(1) 空间推理激活奖励,强制结构化视觉推理;(2) 感知正确性奖励,提供跨任务自适应质量锚点,确保几何与语义对齐;(3) 视觉-语义路径演化奖励,惩罚重复推理,促进互补线索探索以构建更丰富证据链。实验表明,该方法有效缓解‘感知惯性’,推动更深入、多样的推理。仅30亿参数即在REC、OVD、VQA任务上达到顶尖水平,优于最大达70亿参数的模型。零样本泛化能力突出,在VQA、OVD、REC上分别领先第二名3.16%、3.97%、2.72%。代码与数据集见https://github.com/geox-lab/RS-HyRe-R1。

原文摘要 · Abstract (English)

Reinforcement learning (RL) post-training substantially improves remote sensing vision-language models (RS-VLMs). However, when handling complex remote sensing imagery (RSI) requiring exhaustive visual scanning, models tend to rely on localized salient cues for rapid inference. We term this RL-induced bias "perceptual inertia". Driven by reward maximization, models favor quick outcome fitting, leading to two limitations: cognitively, overreliance on specific features impedes complete evidence construction; operationally, models struggle to flexibly shift visual focus across tasks. To address this bias and encourage comprehensive visual evidence mining, we propose RS-HyRe-R1, a hybrid reward framework for RSI understanding. It introduces: (1) a spatial reasoning activation reward that enforces structured visual reasoning; (2) a perception correctness reward that provides adaptive quality anchors across RS tasks, ensuring accurate geometric and semantic alignment; and (3) a visual-semantic path evolution reward that penalizes repetitive reasoning and promotes exploration of complementary cues to build richer evidence chains. Experiments show RS-HyRe-R1 effectively mitigates "perceptual inertia", encouraging deeper, more diverse reasoning. With only 3B parameters, it achieves state-of-the-art performance on REC, OVD, and VQA tasks, outperforming models up to 7B parameters. It also demonstrates strong zero-shot generalization, surpassing the second-best model by 3.16%, 3.97%, and 2.72% on VQA, OVD, and REC, respectively. Code and datasets are available at https://github.com/geox-lab/RS-HyRe-R1.

遥感图像强化学习视觉推理奖励机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。