arXiv:2510.15448cs.CV2025-10被引 2

融合多视角信息提升微型无人机动作识别准确率

MAVR-Net: Robust Multi-View Learning for MAV Action Recognition with Cross-View Attention

  • 用RGB、光流和分割图三路输入,通过跨视图注意力增强特征交互
  • 在短/中/长任务数据集上分别达到97.8%、96.5%、92.8%准确率
  • 适合需要高鲁棒性动作识别的无人机协同系统研究者

微型飞行器(MAVs)的动作识别对自主空中集群的协作感知与控制至关重要。现有基于RGB图像的方法难以捕捉复杂的时空运动特征,导致动作区分能力有限。为此,本文提出MAVR-Net,一种基于多视角学习的无人机动作识别框架。该方法融合原始RGB帧、光流和分割掩码三类互补数据,采用基于ResNet的编码器提取各视角判别性特征,并使用多尺度特征金字塔保留运动模式的时空细节。引入跨视图注意力模块建模不同模态与特征尺度间的依赖关系,设计多视图对齐损失以确保语义一致性并强化跨视图特征表示。在标准MAV动作识别数据集上的实验表明,本方法显著优于现有方法,在Short MAV、Medium MAV、Long MAV数据集上分别达到97.8%、96.5%、92.8%的准确率。

原文摘要 · Abstract (English)

Recognizing the motion of Micro Aerial Vehicles (MAVs) is crucial for enabling cooperative perception and control in autonomous aerial swarms. Yet, vision-based recognition models relying only on RGB data often fail to capture the complex spatial temporal characteristics of MAV motion, which limits their ability to distinguish different actions. To overcome this problem, this paper presents MAVR-Net, a multi-view learning-based MAV action recognition framework. Unlike traditional single-view methods, the proposed approach combines three complementary types of data, including raw RGB frames, optical flow, and segmentation masks, to improve the robustness and accuracy of MAV motion recognition. Specifically, ResNet-based encoders are used to extract discriminative features from each view, and a multi-scale feature pyramid is adopted to preserve the spatiotemporal details of MAV motion patterns. To enhance the interaction between different views, a cross-view attention module is introduced to model the dependencies among various modalities and feature scales. In addition, a multi-view alignment loss is designed to ensure semantic consistency and strengthen cross-view feature representations. Experimental results on benchmark MAV action datasets show that our method clearly outperforms existing approaches, achieving 97.8\%, 96.5\%, and 92.8\% accuracy on the Short MAV, Medium MAV, and Long MAV datasets, respectively.

动作识别多视角学习无人机注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。