首个评估语音助手多模态理解能力的基准,涵盖声音细节与视觉信息融合。
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
- 构建包含1000段真人对话的多模态数据集,含语音情感、音色等细粒度特征。
- 10个主流模型在该基准上表现均不及人类,难以生成情境化响应。
- 适合研究多模态交互、语音理解与人机对话的学者和开发者参考。
大型语言模型(LLMs)的发展使通用模型能够作为语音助手处理口语对话。这些模型可处理文本以外的多模态输入,如语音和视觉数据,实现更富情境感的交互。然而,现有基准在全面评估模型生成情境化响应的能力方面存在不足,尤其在隐含理解细微语音特征(如音高、情绪、音色、音量)或环境声学背景(如背景噪音)方面。此外,它们未能充分评估模型将副语言线索与互补视觉信号对齐以指导响应的能力。为弥补这些缺陷,我们提出 MultiVox,首个面向通用语音助手的基准,用于评估其整合语音与视觉线索(包括副语言语音特征)实现真正多模态理解的能力。MultiVox 包含1000段人工标注并录制的语音对话,涵盖多样化的副语言特征及图像、视频等多种视觉线索。我们在10个前沿模型上进行评估发现,尽管人类表现优异,当前模型仍持续难以生成情境相关的回应。
原文摘要 · Abstract (English)
The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more context-aware interactions. However, current benchmarks fall short in comprehensively evaluating how well these models generate context-aware responses, particularly when it comes to implicitly understanding fine-grained speech characteristics, such as pitch, emotion, timbre, and volume or the environmental acoustic context such as background sounds. Additionally, they inadequately assess the ability of models to align paralinguistic cues with complementary visual signals to inform their responses. To address these gaps, we introduce MultiVox, the first omni voice assistant benchmark designed to evaluate the ability of voice assistants to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. Specifically, MultiVox includes 1000 human-annotated and recorded speech dialogues that encompass diverse paralinguistic features and a range of visual cues such as images and videos. Our evaluation on 10 state-of-the-art models reveals that, although humans excel at these tasks, current models consistently struggle to produce contextually grounded responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。