用强化学习让视觉语言模型在手术领域适配时保持通用能力
Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning
- 通过结构化推理链实现手术领域知识融入
- 在分布内与分布外测试中均提升准确率与泛化性
- 适合需要保留模型通用能力的医疗视觉任务
在特定领域数据上进行传统微调会无意中改变模型预训练的多模态先验,导致泛化能力下降。为此,我们提出链式适配(Chain-of-Adaptation, CoA)框架,可在融入领域知识的同时保持模型固有的推理与感知能力。CoA引入结构化推理格式,通过强化学习增强领域对齐,同时不损害模型的通用多模态性能。在标准手术基准上的实验表明,CoA在分布内与分布外设置下均优于监督微调,在准确性、泛化性和行为稳定性方面表现更优。消融研究进一步验证了CoA有效保留了模型的核心视觉-语言能力,为视觉语言模型的领域专业化提供可靠路径。
原文摘要 · Abstract (English)
Conventional fine-tuning on domain-specific datasets can inadvertently alter a model's pretrained multimodal priors, leading to reduced generalization. To address this, we propose Chain-of-Adaptation (CoA), an adaptation framework designed to integrate domain knowledge while maintaining the model's inherent reasoning and perceptual capabilities. CoA introduces a structured reasoning format that enhances domain alignment without sacrificing general multimodal competence by reinforcement learning. Experiments on standard surgical benchmarks, under both in-distribution and out-of-distribution settings, demonstrate that CoA achieves higher accuracy, stronger generalization, and more stable behavior than supervised fine-tuning. Furthermore, ablation studies confirm that CoA effectively preserves the model's core visual-language abilities, providing a reliable pathway for domain specialization in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。