arXiv:2606.02246cs.CV2026-06

首个面向实时动作分割的多模态节能基准,助力智能设备在低功耗下持续感知

Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark

论文配图:Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark
图 1 · 摘自论文原文
  • 动态选择传感器激活策略,在严格能耗预算下实现在线动作分割
  • 超过100小时未剪辑的视角视频,覆盖5种模态,验证不同场景最优路由策略
  • 适合研究节能感知、边缘计算与多模态融合的开发者和学者

为在物理世界中运行,具身智能体需以“始终在线”方式感知环境,通过选择性访问最有效传感器,在能耗约束与任务精度间取得平衡。尽管这对资源受限设备至关重要,但现有研究大多假设计算资源无限,能源感知感知仍被忽视。为此,我们提出Ego-METAS:首个面向视角视频的在线多模态节能时序动作分割基准。该基准整合了来自EgoExo4D、CMU-MMAC和CaptainCook4D的超过100小时未剪辑视角视频,涵盖RGB、音频、注视、IMU和单色相机共5种模态。我们定义了一个在线时序动作分割任务,要求模型在每个时间步动态决定激活哪些传感器,同时严格遵守硬件级能量预算。我们还发布了统一数据划分、清洗后的标注、预提取特征及多样化的基线路由策略。评估表明,最优路由高度依赖具体场景,而现有基于修剪片段设计的策略学习方法难以适应连续未剪辑环境。然而,即使简单的互补模态动态融合(如随机路由)也对在严苛能耗下保持预测准确率至关重要。Ego-METAS为开发鲁棒、成本敏感的自主持续感知智能系统提供了标准化基础。

原文摘要 · Abstract (English)

To operate in the physical world, embodied agents must perceive their environment in an "always-on" fashion, selectively accessing the most informative sensors to balance energy constraints and task accuracy. Despite its importance for resource-constrained devices, energy-aware perception remains under-explored, with most prior work assuming unlimited compute. To address this, we introduce Ego-METAS: the first Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark. Ego-METAS provides a unified testbed of more than 100 hours of untrimmed egocentric video from EgoExo4D, CMU-MMAC, and CaptainCook4D, spanning 5 modalities (RGB, audio, gaze, IMU, and monochrome camera). We formulate an online temporal action segmentation task where models must dynamically select which sensors to activate at each timestep while strictly adhering to hardware-representative energy budgets. Alongside the benchmark, we release unified splits, cleaned annotations, pre-extracted features, and a diverse suite of baseline routing policies. Our evaluations show that optimal routing is highly scenario-dependent, and that existing policy-learning methods, designed primarily for trimmed clips, struggle to adapt to continuous, untrimmed environments. However, even simple dynamic fusion of complementary modalities (e.g., via random routing) proves critical for balancing predictive accuracy against strict energy budgets. Ultimately, Ego-METAS provides a standardized foundation to develop robust, cost-aware policies for autonomous, always-on embodied AI.

多模态节能动作分割边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。