提出新方法缓解大模型微调中的推理偏移问题
Walking the Tightrope: Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-Tuning
- 将推理过程建模为非平稳分布,识别有益与有害的漂移
- 通过反事实轨迹解耦优化,提升医疗领域微调稳定性
- 开源32万条反事实推理数据集,适合医疗AI研究者
本文揭示了多模态大模型在非平稳强化微调(RFT)过程中链式思维(CoT)推理出现有害概念漂移的现象:推理标记分布不可预测地演化,导致最终预测严重偏差。我们首次建立概念漂移理论与RFT的理论桥梁,将CoT的自回归标记流形式化为经历任意时间偏移的非平稳分布。基于此框架,提出新型反事实感知RFT,利用概念图增强的LLM专家生成反事实推理路径,系统解耦有益分布适应与有害概念漂移。所提方法Counterfactual Preference Optimization(CPO)在非平稳环境下实现稳定微调,尤其在医疗领域通过反事实感知偏好对齐实现定制化优化。大量实验表明其在鲁棒性、泛化性和协作性上表现卓越。此外,我们还构建了大规模数据集CXR-CounterFact(CCF),包含320,416条从MIMIC-CXR精心构建的反事实推理轨迹。代码与数据已公开。
原文摘要 · Abstract (English)
This paper uncovers a critical yet overlooked phenomenon in multi-modal large language models (MLLMs): detrimental concept drift within chain-of-thought (CoT) reasoning during non-stationary reinforcement fine-tuning (RFT), where reasoning token distributions evolve unpredictably, thereby introducing significant biases in final predictions. To address this, we are pioneers in establishing the theoretical bridge between concept drift theory and RFT processes by formalizing CoT's autoregressive token streams as non-stationary distributions undergoing arbitrary temporal shifts. Leveraging this framework, we propose a novel counterfact-aware RFT that systematically decouples beneficial distribution adaptation from harmful concept drift through concept graph-empowered LLM experts generating counterfactual reasoning trajectories. Our solution, Counterfactual Preference Optimization (CPO), enables stable RFT in non-stationary environments, particularly within the medical domain, through custom-tuning of counterfactual-aware preference alignment. Extensive experiments demonstrate our superior performance of robustness, generalization and coordination within RFT. Besides, we also contributed a large-scale dataset CXR-CounterFact (CCF), comprising 320,416 meticulously curated counterfactual reasoning trajectories derived from MIMIC-CXR. Our code and data are public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。