arXiv:2410.03321cs.CV2024-10ICLR被引 25

让模型像人一样多轮看图推理,解决指令模糊问题

Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

  • 设计多轮多模态思维链,模拟人类看图解惑过程
  • 在模糊指令上提升各智能水平模型性能,通用性强
  • 适合需要理解复杂现实指令的AI应用开发者

随着大规模模型发展,语言指令在多模态任务中应用日益广泛。由于人类语言习惯,实际场景中的指令常含歧义,需结合视觉上下文或常识才能准确理解。然而,即使高智能大模型在处理模糊指令时仍表现不佳,其消歧推理能力弱易导致灾难性错误。为此,本文提出Visual-O1:一种多模态多轮思维链推理框架。该框架模拟人类多模态多轮推理过程,为高智能模型提供实例经验,或为一般智能模型提供实证经验,以理解模糊指令。与传统方法不同,本框架不显著增加计算开销,且对一般智能模型同样有效。实验表明,该方法不仅显著提升不同智能水平模型在模糊指令上的表现,也改善其在通用数据集上的性能。研究展示了人工智能在现实不确定性与歧义场景下类人协作的潜力。代码与数据将公开。

原文摘要 · Abstract (English)

As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. However, even highly intelligent large models exhibit significant performance limitations on ambiguous instructions, where weak reasoning abilities of disambiguation can lead to catastrophic errors. To address this issue, this paper proposes Visual-O1, a multi-modal multi-turn chain-of-thought reasoning framework. It simulates human multi-modal multi-turn reasoning, providing instantial experience for highly intelligent models or empirical experience for generally intelligent models to understand ambiguous instructions. Unlike traditional methods that require models to possess high intelligence to understand long texts or perform lengthy complex reasoning, our framework does not significantly increase computational overhead and is more general and effective, even for generally intelligent models. Experiments show that our method not only significantly enhances the performance of models of different intelligence levels on ambiguous instructions but also improves their performance on general datasets. Our work highlights the potential of artificial intelligence to work like humans in real-world scenarios with uncertainty and ambiguity. We will release our data and code.

多模态推理指令理解思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。