arXiv:2505.12665cs.RO2025-05

用声音和视觉融合识别果树接触物类型,提升农业机器人操作安全

Audio-Visual Contact Classification for Tree Structures in Agriculture

  • 融合振动音频与视觉信息进行接触分类
  • 零样本迁移至机械臂传感器,F1达0.82
  • 适合需要精准触觉反馈的农业机器人场景

农业中如修剪、采摘等高接触任务要求机器人在杂乱枝叶间物理互动。仅靠视觉难以判断是否接触刚性或柔性物体,因遮挡和视角受限。为此,我们提出一种多模态分类框架,融合振动(音频)与视觉输入,识别接触类别:叶、枝、干或环境。关键洞察是接触引发的振动携带材料特异性信号,使音频有效检测接触事件并区分材质,而视觉提供互补语义线索,支持更细粒度分类。通过手持传感器采集训练数据,实现零样本迁移至机器人搭载探头,获得0.82的F1分数。结果表明,音视频学习在复杂接触环境中具有巨大潜力。

原文摘要 · Abstract (English)

Contact-rich manipulation tasks in agriculture, such as pruning and harvesting, require robots to physically interact with tree structures to maneuver through cluttered foliage. Identifying whether the robot is contacting rigid or soft materials is critical for the downstream manipulation policy to be safe, yet vision alone is often insufficient due to occlusion and limited viewpoints in this unstructured environment. To address this, we propose a multi-modal classification framework that fuses vibrotactile (audio) and visual inputs to identify the contact class: leaf, twig, trunk, or ambient. Our key insight is that contact-induced vibrations carry material-specific signals, making audio effective for detecting contact events and distinguishing material types, while visual features add complementary semantic cues that support more fine-grained classification. We collect training data using a hand-held sensor probe and demonstrate zero-shot generalization to a robot-mounted probe embodiment, achieving an F1 score of 0.82. These results underscore the potential of audio-visual learning for manipulation in unstructured, contact-rich environments.

农业机器人多模态感知触觉识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。