arXiv:2509.16670cs.SDcs.MM2025-09

让机器听懂话直接找物体,无需预设类别

Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection

  • 用可学习查询压缩语音特征,生成紧凑语义表示
  • 引入高效参数调整架构,实现跨模态深度适配
  • 支持开放集识别,适合人机交互等实际场景

语音定位(即语音驱动的开放集目标检测)旨在直接从语音中定位并识别物体,实现对预定义类别的泛化能力。该任务在人机交互等场景中至关重要,因文本输入不切实际。然而,该领域进展受限于大规模成对音视频数据稀缺,且以往方法依赖间接的文本中介流程。本文提出Speech-to-See(Speech2See),一种基于预训练与微调范式的端到端方法。预训练阶段,设计查询引导语义聚合模块,利用可学习查询将冗余语音嵌入压缩为紧凑语义表征;微调阶段,引入参数高效混合式LoRA专家(MoLE)架构,实现更深层、更细致的跨模态适应。大量实验表明,Speech2See在多个基准上均表现稳健且具备强泛化能力,展现出广泛适用性。

原文摘要 · Abstract (English)

Audio grounding, or speech-driven open-set object detection, aims to localize and identify objects directly from speech, enabling generalization beyond predefined categories. This task is crucial for applications like human-robot interaction where textual input is impractical. However, progress in this domain faces a fundamental bottleneck from the scarcity of large-scale, paired audio-image data, and is further constrained by previous methods that rely on indirect, text-mediated pipelines. In this paper, we introduce Speech-to-See (Speech2See), an end-to-end approach built on a pre-training and fine-tuning paradigm. Specifically, in the pre-training stage, we design a Query-Guided Semantic Aggregation module that employs learnable queries to condense redundant speech embeddings into compact semantic representations. During fine-tuning, we incorporate a parameter-efficient Mixture-of-LoRA-Experts (MoLE) architecture to achieve deeper and more nuanced cross-modal adaptation. Extensive experiments show that Speech2See achieves robust and adaptable performance across multiple benchmarks, demonstrating its strong generalization ability and broad applicability.

语音定位跨模态开放集检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。