arXiv:2511.09958cs.ROcs.SD2025-11被引 3

让机器人通过声音感知抓取过程,提升操作精度与评估全面性。

Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation

  • 融合视觉与接触音效,实现对操作过程的多模态感知。
  • 在仿真和真实任务中显著优于纯视觉模型,动态感知能力更强。
  • 提出新评估指标TCR,量化机器人对操作过程的理解水平。

视觉-语言-动作模型(VLA)在机器人操作中取得显著进展,但仅依赖视觉会限制对交互过程的感知。本文提出Audio-VLA,一种利用接触音频感知接触事件与动态反馈的多模态操作策略,突破了纯视觉VLA的局限。该模型采用预训练DINOv2和SigLIP作为视觉编码器,AudioCLIP作为音频编码器,Llama2作为语言模型骨干,并通过LoRA微调实现跨模态理解。引入多模态投影层将不同模态特征对齐至统一空间。同时,在RLBench和LIBERO仿真环境中添加基于碰撞的音频生成,提供真实声学反馈。针对现有评估侧重最终结果而忽略动态过程的问题,本文提出任务完成率(TCR)指标,系统衡量机器人在操作过程中对动态状态的感知能力。大量实验在LIBERO、RLBench及两个真实任务上验证了Audio-VLA性能优势,且TCR能有效量化动态过程感知能力。

原文摘要 · Abstract (English)

The Vision-Language-Action models (VLA) have achieved significant advances in robotic manipulation recently. However, vision-only VLA models create fundamental limitations, particularly in perceiving interactive and manipulation dynamic processes. This paper proposes Audio-VLA, a multimodal manipulation policy that leverages contact audio to perceive contact events and dynamic process feedback. Audio-VLA overcomes the vision-only constraints of VLA models. Additionally, this paper introduces the Task Completion Rate (TCR) metric to systematically evaluate dynamic operational processes. Audio-VLA employs pre-trained DINOv2 and SigLIP as visual encoders, AudioCLIP as the audio encoder, and Llama2 as the large language model backbone. We apply LoRA fine-tuning to these pre-trained modules to achieve robust cross-modal understanding of both visual and acoustic inputs. A multimodal projection layer aligns features from different modalities into the same feature space. Moreover RLBench and LIBERO simulation environments are enhanced by adding collision-based audio generation to provide realistic sound feedback during object interactions. Since current robotic manipulation evaluations focus on final outcomes rather than providing systematic assessment of dynamic operational processes, the proposed TCR metric measures how well robots perceive dynamic processes during manipulation, creating a more comprehensive evaluation metric. Extensive experiments on LIBERO, RLBench, and two real-world tasks demonstrate Audio-VLA's superior performance over vision-only comparative methods, while the TCR metric effectively quantifies dynamic process perception capabilities.

多模态机器人音频感知评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。