arXiv:2509.08500cs.AI2025-09EMNLP被引 5

用思维对齐优化提升视觉语言模型在动态环境中的决策能力

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making

  • 以分步思维对比方式优化模型推理过程
  • 在ALFWorld中实现26.67%成功率,比RL4VLM提升6%
  • 适合需要稳定推理的具身智能系统研究者

将视觉语言模型(VLMs)的有效泛化能力应用于具身人工智能的特定动态任务仍具挑战。尽管监督微调模型能更好对齐真实物理世界,但在动态环境中仍存在响应迟缓和幻觉问题,需进一步对齐。现有后微调方法依赖强化学习与思维链(CoT),受限于稀疏奖励和仅动作优化,导致样本效率低、一致性差及模型退化。本文提出思维中心偏好优化(TCPO),采用分步偏好优化方法,将稀疏奖励转化为更丰富的步骤样本对,强调模型中间推理过程对齐,缓解模型退化。通过引入动作策略一致性约束(APC),进一步约束输出一致性。在ALFWorld环境的实验表明,平均成功率达26.67%,较RL4VLM提升6%,验证了该方法在微调后缓解模型退化的有效性。结果表明,将基于偏好的学习与思维链结合,可有效提升视觉语言模型在具身代理中的决策能力。

原文摘要 · Abstract (English)

Using effective generalization capabilities of vision language models (VLMs) in context-specific dynamic tasks for embodied artificial intelligence remains a significant challenge. Although supervised fine-tuned models can better align with the real physical world, they still exhibit sluggish responses and hallucination issues in dynamically changing environments, necessitating further alignment. Existing post-SFT methods, reliant on reinforcement learning and chain-of-thought (CoT) approaches, are constrained by sparse rewards and action-only optimization, resulting in low sample efficiency, poor consistency, and model degradation. To address these issues, this paper proposes Thought-Centric Preference Optimization (TCPO) for effective embodied decision-making. Specifically, TCPO introduces a stepwise preference-based optimization approach, transforming sparse reward signals into richer step sample pairs. It emphasizes the alignment of the model's intermediate reasoning process, mitigating the problem of model degradation. Moreover, by incorporating Action Policy Consistency Constraint (APC), it further imposes consistency constraints on the model output. Experiments in the ALFWorld environment demonstrate an average success rate of 26.67%, achieving a 6% improvement over RL4VLM and validating the effectiveness of our approach in mitigating model degradation after fine-tuning. These results highlight the potential of integrating preference-based learning techniques with CoT processes to enhance the decision-making capabilities of vision-language models in embodied agents.

具身智能思维链偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。