arXiv:2510.01236cs.CLcs.LG2025-10

低资源下提升皮肤病视觉语言模型推理能力

GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings

  • 改进GRPO算法,实现高效稳定训练
  • 在皮肤科数据集上性能显著优于传统微调
  • 适合医疗领域小样本场景的模型开发

视觉语言模型在医学图像分析中展现潜力,但在皮肤病等复杂领域受限于数据稀缺和高计算成本。为此,我们提出DermIQ-VLM,采用多阶段、资源高效的训练方法,模拟皮肤科医生诊断流程。核心贡献是改进版分组相对策略优化(GRPO++),稳定了原数据密集型框架。训练流程先用GRPO++进行疾病推理识别,再通过监督微调提升对话能力;为减少此阶段引入的事实错误,进一步采用直接偏好优化(DPO)对齐模型,利用基于知识图谱的系统作为可扩展的专家偏好代理。在精心构建的皮肤科数据集上的初步评估显示,该方法显著优于标准微调方案,验证了其在资源受限环境下构建专业可靠视觉语言模型的可行性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) show promise in medical image analysis, yet their capacity for structured reasoning in complex domains like dermatology is often limited by data scarcity and the high computational cost of advanced training techniques. To address these challenges, we introduce DermIQ-VLM, a VLM developed through a multi-stage, resource-efficient methodology designed to emulate a dermatologist's diagnostic process. Our primary contribution is a modified version of Grouped Relative Policy Optimization (GRPO), called GRPO++, which stabilizes the powerful but data-intensive GRPO framework. Our proposed training pipeline first employs GRPO++ for reasoning-oriented disease recognition, followed by supervised fine-tuning for conversational ability. To mitigate factual errors introduced during this step, we then align the model using Direct Preference Optimization (DPO), leveraging a Knowledge Graph-based system as a scalable proxy for expert preference. A preliminary evaluation on a curated dermatological dataset demonstrates that our proposed methodology yields notable performance gains over standard fine-tuning approaches. These findings validate the potential of our pipeline as a feasible pathway for developing specialized, reliable VLMs in resource-constrained environments.

视觉语言模型皮肤病低资源推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。