arXiv:2607.18897eess.IV2026-07中稿 · publication in the…

动态融合多传感器数据,实现超低功耗下的高精度深度估计。

Thinking Fast, Thinking Slow: Adaptive Multimodal Transformer-based Sensor Fusion for Depth Estimation on Ultra-low-power MCUs

论文配图:Thinking Fast, Thinking Slow: Adaptive Multimodal Transformer-based Sensor Fusion for Depth Estimation on Ultra-low-power MCUs
图 1 · 摘自论文原文
  • 根据置信度动态选择传感器输入,逐步增加信息量
  • 相比全量传感器方案节能90%,仅损失4.8%准确率
  • 参数量仅为移动模型1/9,适合嵌入式设备部署

基于人工智能的多模态传感器融合在超低功耗嵌入式与信息物理系统中日益重要,可提升真实场景下的可靠性、精度与鲁棒性。然而,在功耗低于100 mW的资源受限平台上增加传感器需权衡能耗与预测精度。为此,本文提出一种新型自适应AI方法,结合相机、超声波与飞行时间(ToF)传感器的多模态融合,采用轻量级递归Transformer架构(688 k参数)。通过迭代中的置信度门控机制,动态决定是否引入更丰富但功耗更高的传感器;同时利用特征传播保持时间一致性。我们设计了一块集成三类传感器的新型电路板,搭载超低功耗GWT GAP9多核SoC进行原型验证。在NYUv2数据集上,相较于使用所有传感器与全部迭代的基线,本方法仅损失4.8%的δ1精度,却实现90%能耗降低(2.44 mJ/帧)。相比仅使用9倍参数的MobileDepth模型,本方法δ1精度仅低5.6%;与同在GAP9运行的先进模型相比,精度提升31.8%,且平均功耗控制在约400 mW内。

原文摘要 · Abstract (English)

Artificial intelligence (AI)-based multimodal sensor fusion is a relevant topic gaining ever more traction across ultra-low-power (ULP) embedded and cyber-physical systems, as it improves reliability, accuracy, and robustness under real-world constraints. However, adding more and more sensors to ultra-constrained sub-100 mW platforms requires balancing energy consumption against prediction accuracy. To achieve this ambitious goal, we present a novel adaptive AI methodology that combines multimodal sensor fusion (camera, ultrasound, and Time-of-Flight sensors) with a lightweight recurrent Transformer-based architecture (688 k parameters). We address the depth map estimation task with a mechanism that combines token propagation across iterations with incremental sensor utilization. At each iteration, a confidence-based gating mechanism dynamically decides whether to continue the computation by adding progressively richer but more power-demanding sensors as input. Token propagation ensures temporal consistency by forwarding context features across time. To deploy our algorithm and test a first real-world prototype, we design a novel printed circuit board featuring all three sensors, coupled with an ULP GWT GAP9 multicore System-on-Chip. When comparing our adaptive system against the same pipeline using all sensors and iterations on the NYUv2 dataset, we lose only 4.8% of the δ1 accuracy in exchange for 90% energy saving (2.44 mJ/frame). Finally, our adaptive method marks only 5.6% lower δ1 accuracy than MobileDepth despite using 9x fewer parameters. Compared with a state-of-the-art model also running on GAP9, our method improves δ1 accuracy by 31.8% thanks to our adaptive sensor fusion while operating within the same average power budget (~400 mW).

深度估计多模态融合超低功耗轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。