构建语音控制智能家居的多轮交互数据集,揭示大模型在真实场景中的能力差距。
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

- 设计语音驱动的多轮代码生成任务,模拟智能家居设备操作
- 发现开源与闭源多模态大模型在任务上存在显著性能差距
- 适合研究语音助手、物理世界推理和混合主动交互的学者
物联网设备在现实世界中的普及,要求语音接口具备处理复杂用户体验的能力。尽管现代大语言模型已展现出强大的工具使用能力,但建模真实世界的物联网设备仍是一个困难且研究不足的挑战,需结合时空约束建模、语音输入、动态状态追踪及混合主动交互模式。我们提出MIST(多模态交互式语音工具调用数据集),一个基于物联网设备的合成多轮语音驱动代码生成任务。实验发现,开放权重与闭源多模态大模型在MIST上的表现存在显著差距,即使最先进的闭源模型仍有巨大提升空间。我们公开发布MIST及可扩展的数据生成框架,以促进对考虑物理世界约束的混合主动语音助手的研究。
原文摘要 · Abstract (English)
The rise of Internet of Things (IoT) devices in the physical world necessitates voice-based interfaces capable of handling complex user experiences. While modern Large Language Models (LLMs) already demonstrate strong tool-usage capabilities, modeling real-world IoT devices presents a difficult, understudied challenge which combines modeling spatiotemporal constraints with speech inputs, dynamic state tracking, and mixed-initiative interaction patterns. We introduce MIST (the Multimodal Interactive Speech-based Tool-calling Dataset), a synthetic multi-turn, voice-driven code generation task that operates over IoT devices. We find that there is a significant gap between open- and closed-weight multimodal LLMs on MIST, and that even frontier closed-weight LLMs have substantial headroom. We release MIST and an extensible data generation framework to build related datasets in order to facilitate research on mixed-initiative voice assistants which reason about physical world constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。