根据场景重要性动态调度感知模块,提升人机协作实时性
Perceive What Matters: Relevance-Driven Scheduling for Multimodal Streaming Perception
- 基于前帧输出预估当前关键感知任务,实现轻量级实时调度
- 计算延迟降低27.52%,姿态激活召回率提升72.73%,关键帧准确率达98%
- 适合需要低延迟、高效率的多模态实时感知系统开发者
在现代人机协作(HRC)应用中,多个感知模块协同提取视觉、听觉和上下文线索以实现全面场景理解,使机器人能智能地为人类提供协助。尽管逐帧执行多个感知模块可提升离线场景下的感知质量,但在流式感知场景中会累积显著延迟,导致系统性能下降。现有研究提出的‘相关性’概念为高效方法奠定了基础,但当前感知流水线仍面临信息冗余与计算资源分配不合理的问题。受相关性概念及HRC事件中信息稀疏性的启发,本文提出一种新型轻量级感知调度框架,通过利用前帧输出实时估计并调度必要的感知模块。实验表明,该框架相比传统并行感知管道,计算延迟最高降低27.52%,MMPose激活召回率提升72.73%,关键帧准确率最高达98%。结果验证了其在不显著牺牲精度的前提下,显著提升实时感知效率的能力。该框架具备作为可扩展、系统化解决方案应用于多模态流式感知系统的潜力。
原文摘要 · Abstract (English)
In modern human-robot collaboration (HRC) applications, multiple perception modules jointly extract visual, auditory, and contextual cues to achieve comprehensive scene understanding, enabling the robot to provide appropriate assistance to human agents intelligently. While executing multiple perception modules on a frame-by-frame basis enhances perception quality in offline settings, it inevitably accumulates latency, leading to a substantial decline in system performance in streaming perception scenarios. Recent work in scene understanding, termed Relevance, has established a solid foundation for developing efficient methodologies in HRC. However, modern perception pipelines still face challenges related to information redundancy and suboptimal allocation of computational resources. Drawing inspiration from the Relevance concept and the information sparsity in HRC events, we propose a novel lightweight perception scheduling framework that efficiently leverages output from previous frames to estimate and schedule necessary perception modules in real-time based on scene context. The experimental results demonstrate that the proposed perception scheduling framework effectively reduces computational latency by up to 27.52% compared to conventional parallel perception pipelines, while also achieving a 72.73% improvement in MMPose activation recall. Additionally, the framework demonstrates high keyframe accuracy, achieving rates of up to 98%. The results validate the framework's capability to enhance real-time perception efficiency without significantly compromising accuracy. The framework shows potential as a scalable and systematic solution for multimodal streaming perception systems in HRC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。