arXiv:2504.01049cs.CVcs.LG2025-04被引 1

直接理解语音提问的图文问答模型,无需转写文本。

SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering

  • 端到端处理语音输入,跳过文字转录环节
  • 在SBVQA数据集上达75.62%准确率,语音+文字混合输入提升至78.85%
  • 适合语音交互、无障碍访问等场景的多模态应用

融合语音与视觉的多模态模型在人机交互中具有重要意义,尤其在基于语音的图文问答(SBVQA)任务中,需直接理解关于图像的口头提问。现有方法多聚焦于文本-视觉融合,忽视了语音-视觉模态因本质差异带来的挑战。为此,我们提出统一的语音-视觉模型SViQA,直接处理口语问题而无需文字转录。基于LLaVA架构,本框架通过两项关键创新实现音视模态对齐:(1) 端到端语音特征提取,避免中间文本转换;(2) 跨模态对齐优化,实现语音信号与视觉内容的有效融合。在SBVQA基准上的实验表明,所提模型达到75.62%准确率,具备领先性能和良好多模态泛化能力。采用语音-文本混合输入可进一步提升至78.85%,较纯语音输入提高3.23%,验证了模型更强鲁棒性与有效的跨模态注意力对齐机制。

原文摘要 · Abstract (English)

Multimodal models integrating speech and vision hold significant potential for advancing human-computer interaction, particularly in Speech-Based Visual Question Answering (SBVQA) where spoken questions about images require direct audio-visual understanding. Existing approaches predominantly focus on text-visual integration, leaving speech-visual modality gaps underexplored due to their inherent heterogeneity. To this end, we introduce SViQA, a unified speech-vision model that directly processes spoken questions without text transcription. Building upon the LLaVA architecture, our framework bridges auditory and visual modalities through two key innovations: (1) end-to-end speech feature extraction eliminating intermediate text conversion, and (2) cross-modal alignment optimization enabling effective fusion of speech signals with visual content. Extensive experimental results on the SBVQA benchmark demonstrate the proposed SViQA's state-of-the-art performance, achieving 75.62% accuracy, and competitive multimodal generalization. Leveraging speech-text mixed input boosts performance to 78.85%, a 3.23% improvement over pure speech input, highlighting SViQA's enhanced robustness and effective cross-modal attention alignment.

多模态语音理解图文问答端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。