TACO提升视觉语言模型长链推理一致性与数据效率。
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
- 通过思维-答案一致性机制,确保回答基于连贯推理。
- 在REC和VQA任务上,性能显著优于现有方法。
- 适合需要稳定长链推理的多模态应用开发者。
DeepSeek R1显著提升了大语言模型的复杂推理能力。尽管近期方法尝试在多模态场景中复现R1的推理能力,但仍存在推理与最终答案不一致、长链探索中模型不稳定甚至崩溃、数据学习效率低等问题。为此,我们提出TACO,一种用于视觉推理的新型强化学习算法。基于广义强化策略优化(GRPO),TACO引入思维-答案一致性机制,紧密耦合推理过程与答案一致性,确保答案根植于深入思考。同时提出回滚重采样策略,动态移除异常样本并重新注入采样器,实现稳定的长链探索与未来学习机会。此外,TACO采用自适应学习调度,聚焦中等难度样本以优化数据效率。还提出测试时分辨率缩放方案,缓解因推理中分辨率变化导致的性能下降,同时控制计算开销。在REC和VQA任务的分布内与分布外基准上,微调LVLMs均带来显著性能提升。
原文摘要 · Abstract (English)
DeepSeek R1 has significantly advanced complex reasoning for large language models (LLMs). While recent methods have attempted to replicate R1's reasoning capabilities in multimodal settings, they face limitations, including inconsistencies between reasoning and final answers, model instability and crashes during long-chain exploration, and low data learning efficiency. To address these challenges, we propose TACO, a novel reinforcement learning algorithm for visual reasoning. Building on Generalized Reinforcement Policy Optimization (GRPO), TACO introduces Think-Answer Consistency, which tightly couples reasoning with answer consistency to ensure answers are grounded in thoughtful reasoning. We also introduce the Rollback Resample Strategy, which adaptively removes problematic samples and reintroduces them to the sampler, enabling stable long-chain exploration and future learning opportunities. Additionally, TACO employs an adaptive learning schedule that focuses on moderate difficulty samples to optimize data efficiency. Furthermore, we propose the Test-Time-Resolution-Scaling scheme to address performance degradation due to varying resolutions during reasoning while balancing computational overhead. Extensive experiments on in-distribution and out-of-distribution benchmarks for REC and VQA tasks show that fine-tuning LVLMs leads to significant performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。