arXiv:2505.23308eess.AScs.AI2025-05中稿 · Interspeech 2025被引 1

让模型听懂图片里的问题,用合成语音训练效果接近真实数据。

Spoken question answering for visual queries

  • 融合文本、语音和图像三模态,实现听图问答
  • 仅用合成语音训练的模型接近真实语音表现
  • 不同语音合成模型对准确率影响很小

问答系统通常处理自然语言问题。视觉问答(VQA)和语音问答(SQA)系统分别扩展了文本问答以支持图像和语音输入。本文旨在构建一个支持语音与图像交互的系统,通过融合文本、语音和图像模态来解决语音视觉问答(SVQA)任务。该多模态模型接收文本、语音和图像输入,可回答关于图像的语音问题。训练和评估SVQA模型需要包含三模态的数据集,但目前尚无此类数据集。为此,我们利用两个零样本语音合成(TTS)模型合成VQA数据集。初步结果表明,仅使用合成语音训练的模型性能几乎达到基于文本问答训练的上界模型水平。此外,实验显示不同TTS模型的选择对准确率影响较小。

原文摘要 · Abstract (English)

Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spoken input respectively. This work aims to create a system that enables user interaction through both speech and images. That is achieved through the fusion of text, speech, and image modalities to tackle the task of spoken VQA (SVQA). The resulting multi-modal model has textual, visual, and spoken inputs and can answer spoken questions on images. Training and evaluating SVQA models requires a dataset for all three modalities, but no such dataset currently exists. We address this problem by synthesizing VQA datasets using two zero-shot TTS models. Our initial findings indicate that a model trained only with synthesized speech nearly reaches the performance of the upper-bounding model trained on textual QAs. In addition, we show that the choice of the TTS model has a minor impact on accuracy.

语音问答多模态合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。