arXiv:2604.15705cs.LG2026-04被引 4

提出CPO++框架,应对多模态模型推理中的内在漂移问题。

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning

论文配图:Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning
图 1 · 摘自论文原文
  • 引入反事实推理与领域知识,动态调控思维与感知过程。
  • 在医疗诊断和自动驾驶中实现更高推理一致性和抗干扰能力。
  • 适合安全关键场景下需要稳定多模态推理的研究者使用。

强化微调(RFT)已成为多模态大语言模型(MLLMs)对齐复杂人类价值观和领域特定需求的关键范式。然而,现有研究主要关注由数据因素引发的外部分布偏移,而模型内部推理过程的非平稳性——即内生推理漂移——仍缺乏深入探讨。本文揭示了MLLMs的一个关键弱点:在自回归生成过程中,其思维与感知均易受内生推理漂移影响,表现为独立于外部扰动的不可预测分布变化。为此,我们首次将该现象理论定义为多模态概念漂移,并提出反事实偏好优化++(CPO++),一种针对多模态概念漂移的综合性自主适应框架。该框架融合反事实推理与领域知识,对思维与感知进行可控扰动,通过偏好优化解耦虚假相关。在医疗诊断与自动驾驶两个高度动态且安全性要求极高的领域进行广泛实验,结果表明该框架在推理连贯性、决策精度及极端干扰下的鲁棒性方面均显著优于现有方法。该方法还展现出卓越的零样本跨域泛化能力,为安全关键应用中的可靠多模态推理提供了原则性基础。

原文摘要 · Abstract (English)

Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with complex human values and domain-specific requirements. Nevertheless, current research primarily focuses on mitigating exogenous distribution shifts arising from data-centric factors, the non-stationarity inherent in the endogenous reasoning remains largely unexplored. In this work, a critical vulnerability is revealed within MLLMs: they are highly susceptible to endogenous reasoning drift, across both thinking and perception perspectives. It manifests as unpredictable distribution changes that emerge spontaneously during the autoregressive generation process, independent of external environmental perturbations. To adapt it, we first theoretically define endogenous reasoning drift within the RFT of MLLMs as the multi-modal concept drift. In this context, this paper proposes Counterfactual Preference Optimization ++ (CPO++), a comprehensive and autonomous framework adapted to the multi-modal concept drift. It integrates counterfactual reasoning with domain knowledge to execute controlled perturbations across thinking and perception, employing preference optimization to disentangle spurious correlations. Extensive empirical evaluations across two highly dynamic and safety-critical domains: medical diagnosis and autonomous driving. They demonstrate that the proposed framework achieves superior performance in reasoning coherence, decision-making precision, and inherent robustness against extreme interference. The methodology also exhibits exceptional zero-shot cross-domain generalization, providing a principled foundation for reliable multi-modal reasoning in safety-critical applications.

多模态推理鲁棒性强化学习医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。