arXiv:2601.16419cs.CLcs.CV2026-01

让多模态大模型通过强化学习注入领域知识,提升专业任务表现

Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning

  • 将领域知识转为优化目标中的约束与奖励信号
  • 在遥感与医疗领域多个数据集上达到顶尖性能
  • 适合需要高精度领域理解的科研应用

多模态大语言模型在多模态感知与理解任务中表现出色,但在遥感、医学影像等专业领域效果仍有限。传统方法通过文本指令或辅助描述注入领域知识,但实验发现即使明确提供知识,性能提升微乎其微。这表明当前模型无法仅通过语言内化领域先验,必须在优化层面融合领域知识。为此,我们提出一种强化微调框架,将领域知识编码为输出空间的约束与奖励信号,直接引导模型行为。在遥感与医疗多个数据集上的实验显示,该方法持续带来显著性能提升,达到多模态领域任务的最先进水平。结果揭示了现有模型在文本领域条件下的根本局限,并强调优化层面知识整合的必要性。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown remarkable capabilities in multimodal perception and understanding tasks. However, their effectiveness in specialized domains, such as remote sensing and medical imaging, remains limited. A natural approach to domain adaptation is to inject domain knowledge through textual instructions, prompts, or auxiliary captions. Surprisingly, we find that such input-level domain knowledge injection yields little to no improvement on scientific multimodal tasks, even when the domain knowledge is explicitly provided. This observation suggests that current MLLMs fail to internalize domain-specific priors through language alone, and that domain knowledge must be integrated at the optimization level. Motivated by this insight, we propose a reinforcement fine-tuning framework that incorporates domain knowledge directly into the learning objective. Instead of treating domain knowledge as descriptive information, we encode it as domain-informed constraints and reward signals, shaping the model's behavior in the output space. Extensive experiments across multiple datasets in remote sensing and medical domains consistently demonstrate good performance gains, achieving state-of-the-art results on multimodal domain tasks. Our results highlight the necessity of optimization-level domain knowledge integration and reveal a fundamental limitation of textual domain conditioning in current MLLMs.

多模态强化学习领域知识大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。