让视觉语言模型自动调用专家模型自我提升,无需更大模型或人工干预。
AIDE: Agentically Improve Visual Language Model with Domain Experts
- 通过四阶段流程让模型自主识别问题并调用领域专家改进自身
- 在多个基准上实现显著性能提升,不依赖更大模型或人工标注
- 适合缺乏更强模型时持续优化视觉语言模型的场景
视觉语言模型(VLMs)的增强传统上依赖于从更大、更强大模型中进行知识蒸馏。这种依赖带来了根本性瓶颈,尤其当不存在更优模型时。我们提出AIDE(通过领域专家实现智能改进),一种新框架,使VLM能自主利用专业化领域专家模型来提升自身能力。AIDE采用四阶段流程:(1)识别需优化的实例,(2)调用领域专家进行针对性分析,(3)融合专家输出与已有数据,(4)将优化后的实例融入训练流程。在多组基准测试(包括MMMU、MME、MMBench等)上的实验表明,AIDE可在不依赖更大VLM或人工监督的情况下实现显著性能提升。该框架提供了一种可扩展、资源高效的VLM持续改进方法,解决了现有方法的关键局限,尤其在无法获取更大模型时极具价值。
原文摘要 · Abstract (English)
The enhancement of Visual Language Models (VLMs) has traditionally relied on knowledge distillation from larger, more capable models. This dependence creates a fundamental bottleneck for improving state-of-the-art systems, particularly when no superior models exist. We introduce AIDE (Agentic Improvement through Domain Experts), a novel framework that enables VLMs to autonomously enhance their capabilities by leveraging specialized domain expert models. AIDE operates through a four-stage process: (1) identifying instances for refinement, (2) engaging domain experts for targeted analysis, (3) synthesizing expert outputs with existing data, and (4) integrating enhanced instances into the training pipeline. Experiments on multiple benchmarks, including MMMU, MME, MMBench, etc., demonstrate AIDE's ability to achieve notable performance gains without relying on larger VLMs nor human supervision. Our framework provides a scalable, resource-efficient approach to continuous VLM improvement, addressing critical limitations in current methodologies, particularly valuable when larger models are unavailable to access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。