让纯文本大模型指导视觉语言模型生成更符合需求的描述
Can a Unimodal Language Agent Provide Preferences to Tune a Multimodal Vision-Language Model?
- 用纯文本大模型分析自身信息需求,给出反馈优化多模态模型
- 多模态描述准确率最高提升13%,显著优于基线方法
- 适合对多模态生成质量有要求的研究者和开发者
为探索向现有大语言模型(LLM)添加多模态能力的更可扩展路径,本文探讨一个核心问题:仅依赖文本的单模态大模型能否识别自身的信息需求,并有效提供反馈以优化多模态模型?为此,我们提出一种方法,使语言代理能够向视觉-语言模型(VLM)提供反馈,从而调整文本生成以匹配代理偏好。实验结果验证了该假设,表明大模型偏好反馈能显著提升VLM的描述质量。使用本方法,VLM生成的多模态场景描述帮助大模型更好地理解上下文,使准确率最高提升13%(绝对值)。此外,人工评估验证了该反馈的有效性,大模型选择与人类判断的偏好一致性达到64.6%。大量实验揭示了方法的工作机制及其局限性。
原文摘要 · Abstract (English)
To explore a more scalable path for adding multimodal capabilities to existing LLMs, this paper addresses a fundamental question: Can a unimodal LLM, relying solely on text, reason about its own informational needs and provide effective feedback to optimize a multimodal model? To answer this, we propose a method that enables a language agent to give feedback to a vision-language model (VLM) to adapt text generation to the agent's preferences. Our results from different experiments affirm this hypothesis, showing that LLM preference feedback significantly enhances VLM descriptions. Using our proposed method, we find that the VLM can generate multimodal scene descriptions to help the LLM better understand multimodal context, leading to improvements of maximum 13% in absolute accuracy compared to the baseline multimodal approach. Furthermore, a human study validated our AI-driven feedback, showing a 64.6% preference alignment rate between the LLM's choices and human judgments. Extensive experiments provide insights on how and why the method works and its limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。