arXiv:2506.21080cs.CVcs.AI2025-06ICCV被引 4

EgoAdapt通过自适应多模态蒸馏与策略学习,大幅降低场景感知模型计算开销。

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

  • 自适应跨模态蒸馏+任务感知策略模块,动态优化推理过程
  • 在3个数据集上减少89%的GMACs、82%参数、9.6倍能耗
  • 适合边缘设备部署,兼顾效率与性能,尤其适合动作识别等任务

当前针对多模态具身感知任务的感知模型虽表现优异,但计算成本高,难以在资源受限环境下部署。本文提出EgoAdapt框架,通过自适应跨模态蒸馏与策略学习,在具身动作识别、主动说话者定位和行为预测等任务中实现高效推理。其策略模块可适配不同任务的动作空间,具备广泛适用性。在EPIC-Kitchens、EasyCom和Aria Everyday Activities三个具身感知数据集上的实验表明,该方法显著提升效率:最大降低89.09%的GMACs、82.02%的参数量、9.6倍能耗,同时性能与现有最先进模型相当甚至更优。

原文摘要 · Abstract (English)

Modern perception models, particularly those designed for multisensory egocentric tasks, have achieved remarkable performance but often come with substantial computational costs. These high demands pose challenges for real-world deployment, especially in resource-constrained environments. In this paper, we introduce EgoAdapt, a framework that adaptively performs cross-modal distillation and policy learning to enable efficient inference across different egocentric perception tasks, including egocentric action recognition, active speaker localization, and behavior anticipation. Our proposed policy module is adaptable to task-specific action spaces, making it broadly applicable. Experimental results on three challenging egocentric datasets EPIC-Kitchens, EasyCom, and Aria Everyday Activities demonstrate that our method significantly enhances efficiency, reducing GMACs by up to 89.09%, parameters up to 82.02%, and energy up to 9.6x, while still on-par and in many cases outperforming, the performance of corresponding state-of-the-art models.

具身感知多模态蒸馏高效推理边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。