让AI主动提问,精准调用视觉音频等多模态信息。
Chain of Questions: Guiding Multimodal Curiosity in Language Models
- 通过自动生成问题引导模型选择性激活视觉、音频等模态。
- 在融合WebGPT等数据集上,推理准确率显著提升。
- 适合需要动态感知环境的机器人、智能助手场景。
大型语言模型(LLMs)的推理能力已通过思维链等方法大幅提升,但这些进展尚未充分延伸至多模态场景。在复杂真实环境中,模型需主动决定何时调用视觉、音频或空间感知等感官模态。本文提出链式问题(Chain of Questions, CoQ)框架,一种基于好奇心的推理机制,促使多模态语言模型动态生成针对周围环境的特定问题。这些问题引导模型选择性激活相关模态,以获取关键信息,从而实现更准确的推理与回答生成。我们在一个新构建的多模态基准数据集上评估该方法,该数据集整合了WebGPT、ScienceQA、AVSD和ScanQA。实验结果表明,CoQ方法显著提升了基础模型识别并融合相关感官信息的能力,进而提高了推理准确性、可解释性,并增强了推理过程与多样化多模态任务的一致性。
原文摘要 · Abstract (English)
Reasoning capabilities in large language models (LLMs) have substantially advanced through methods such as chain-of-thought and explicit step-by-step explanations. However, these improvements have not yet fully transitioned to multimodal contexts, where models must proactively decide which sensory modalities such as vision, audio, or spatial perception to engage when interacting with complex real-world environments. In this paper, we introduce the Chain of Questions (CoQ) framework, a curiosity-driven reasoning approach that encourages multimodal language models to dynamically generate targeted questions regarding their surroundings. These generated questions guide the model to selectively activate relevant modalities, thereby gathering critical information necessary for accurate reasoning and response generation. We evaluate our framework on a novel multimodal benchmark dataset, assembled by integrating WebGPT, ScienceQA, AVSD, and ScanQA datasets. Experimental results demonstrate that our CoQ method improves a foundation model's ability to effectively identify and integrate pertinent sensory information. This leads to improved accuracy, interpretability, and alignment of the reasoning process with diverse multimodal tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。