用语音指令实现视觉问答与目标定位,模型表现超越现有水平
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
- 基于语音+图像+文本三模态融合,端到端处理语音指令
- 在MMMU和ScienceQA上达到当前最优性能,支持复杂语音推理
- 适用于智能助手、无障碍交互等需要语音驱动推理的场景
视觉语言模型在视觉问答和图像描述等任务中表现出色,但多数依赖文本指令,限制了人机交互效果。尤其在语音指令下,推理与提示技术(如COT)尚未充分探索。为此,我们提出SilVar,一种端到端多模态模型,可使用语音指令进行视觉问答中的推理。该模型结合CLIP、Whisper与LLaMA 3.1-8B,支持用户通过语音或文本提供指令。我们构建了一个新数据集,用于挑战模型在语音指令下的目标定位与推理能力,推动从物体识别向基于推理的交互演进。实验表明,尽管面临语音指令的挑战,SilVar在MMMU和ScienceQA基准上仍取得当前最优表现。我们的代码与数据集已公开。
原文摘要 · Abstract (English)
Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based instructions, limiting their effectiveness in human-machine interactions. Moreover, the quality of language models depends on reasoning and prompting techniques, such as COT, which remain underexplored when using speech instructions. To address these challenges, we propose SilVar, a novel end-to-end multimodal model that uses speech instructions for reasoning in visual question answering. In addition, we investigate reasoning techniques with levels including conversational, simple, and complex speech instruction. SilVar is built upon CLIP, Whisper, and LLaMA 3.1-8B, enabling intuitive interactions by allowing users to provide verbal or text instructions. To this end, we introduce a dataset designed to challenge models with speech-based reasoning tasks for object localization. This dataset enhances the model ability to process and explain visual scenes from spoken input, moving beyond object recognition to reasoning-based interactions. The experiments show that SilVar achieves SOTA performance on the MMMU and ScienceQA benchmarks despite the challenge of speech-based instructions. We believe SilVar will inspire next-generation multimodal reasoning models, toward expert artificial general intelligence. Our code and dataset are available here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。