通过强化学习优化模糊决策,提升少样本变化检测问答性能。
Improving Few-Shot Change Detection Visual Question Answering via Decision-Ambiguity-guided Reinforcement Fine-Tuning
- 基于模糊样本挖掘与相对优势优化,增强模型判别力。
- 在少样本设置下相比监督微调提升显著,准确率更高。
- 适合需要高鲁棒性的遥感图像问答任务研究者。
变化检测视觉问答(CDVQA)需基于双时相遥感图像推理语义变化并回答文本问题。当前主流方法通过监督微调(SFT)提升通用视觉语言模型性能。然而我们发现,大量失败并非源于错误预测,而是决策模糊——模型对正确答案与强干扰项赋予相近置信度。为此,我们定义决策模糊样本(DAS)为真实答案与最强干扰项间概率差距较小的实例。提出DARFT框架,先用SFT模型挖掘DAS,再在该子集上采用组内相对策略优化。通过多样本解码与组内相对优势,DARFT抑制强干扰项、锐化决策边界,无需额外标注。大量实验表明,在少样本场景下持续优于SFT基线。
原文摘要 · Abstract (English)
Change detection visual question answering (CDVQA) requires answering text queries by reasoning about semantic changes in bi-temporal remote sensing images. A straightforward approach is to boost CDVQA performance with generic vision-language models via supervised fine-tuning (SFT). Despite recent progress, we observe that a significant portion of failures do not stem from clearly incorrect predictions, but from decision ambiguity, where the model assigns similar confidence to the correct answer and strong distractors. To formalize this challenge, we define Decision-Ambiguous Samples (DAS) as instances with a small probability margin between the ground-truth answer and the most competitive alternative. We argue that explicitly optimizing DAS is crucial for improving the discriminability and robustness of CDVQA models. To this end, we propose DARFT, a Decision-Ambiguity-guided Reinforcement Fine-Tuning framework that first mines DAS using an SFT-trained reference policy and then applies group-relative policy optimization on the mined subset. By leveraging multi-sample decoding and intra-group relative advantages, DARFT suppresses strong distractors and sharpens decision boundaries without additional supervision. Extensive experiments demonstrate consistent gains over SFT baselines, particularly under few-shot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。