用对话式微调注入专业知识,不丢通用能力
Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting
- 三阶段对话结构:保基础、辨语义、加专知
- 多领域实验验证,专业能力提升同时保留通用性能
- 适合需要更新知识又怕模型“变傻”的研究者
大视觉语言模型在多模态预训练后展现出强大通用能力,但在融入训练分布外的专有知识时面临困境:直接微调常导致基础视觉-语言能力严重退化。本文提出结构化对话微调(SDFT),通过三阶段对话机制实现知识注入与通用能力保留的平衡。第一阶段“基础保全”利用图像描述任务维持预训练对齐;第二阶段“对比消歧”引入反事实样例强化语义边界;第三阶段“知识专化”通过链式推理嵌入领域知识。实验在多个领域验证了该方法的有效性,其核心贡献包括数据驱动的对话模板、加权多轮监督框架,以及涵盖多样化知识类型的全面评估。
原文摘要 · Abstract (English)
Large Vision Language Models have demonstrated impressive versatile capabilities through extensive multimodal pre-training, but face significant limitations when incorporating specialized knowledge domains beyond their training distribution. These models struggle with a fundamental dilemma: direct adaptation approaches that inject domain-specific knowledge often trigger catastrophic forgetting of foundational visual-linguistic abilities. We introduce Structured Dialogue Fine-Tuning (SDFT), an effective approach that effectively injects domain-specific knowledge while minimizing catastrophic forgetting. Drawing inspiration from supervised fine-tuning in LLMs and subject-driven personalization in text-to-image diffusion models, our method employs a three-phase dialogue structure: Foundation Preservation reinforces pre-trained visual-linguistic alignment through caption tasks; Contrastive Disambiguation introduces carefully designed counterfactual examples to maintain semantic boundaries; and Knowledge Specialization embeds specialized information through chain-of-thought reasoning. Experimental results across multiple domains confirm SDFT's effectiveness in balancing specialized knowledge acquisition with general capability retention. Our key contributions include a data-centric dialogue template that balances foundational alignment with targeted knowledge integration, a weighted multi-turn supervision framework, and comprehensive evaluation across diverse knowledge types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。