arXiv:2604.21926cs.CV2026-04被引 1

仅用可穿戴传感器实现人体与场景的4D动态重建,突破视觉依赖。

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs

论文配图:Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs
图 1 · 摘自论文原文
  • 利用耳塞、手表等传感器数据,通过大模型实现非视觉时空理解
  • 在多个数据集上生成更连贯稳定的4D人体运动与粗粒度场景结构
  • 适合隐私敏感、低功耗场景下的智能感知应用

理解人类活动及其周围环境通常依赖视觉感知,但摄像头在隐私、安全、能效和可扩展性方面存在持续挑战。本文探索一种无视觉的4D感知方法:仅凭日常可穿戴传感器重建人体运动与3D场景布局。为此提出IMU-to-4D框架,将大规模语言模型用于非视觉的人-场景动态时空理解。该框架使用来自耳塞、手表或智能手机的少量惯性传感器数据,预测详细4D人体运动及粗粒度场景结构。在多个多样化的真人-场景数据集上实验表明,IMU-to-4D生成结果比当前最优级联流水线更具一致性与时间稳定性,表明仅靠可穿戴运动传感器即可支持丰富的4D理解。

原文摘要 · Abstract (English)

Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in privacy, safety, energy efficiency, and scalability. We explore an alternative: 4D perception without vision. Its goal is to reconstruct human motion and 3D scene layouts purely from everyday wearable sensors. For this we introduce IMU-to-4D, a framework that repurposes large language models for non-visual spatiotemporal understanding of human-scene dynamics. IMU-to-4D uses data from a few inertial sensors from earbuds, watches, or smartphones and predicts detailed 4D human motion together with coarse scene structure. Experiments across diverse human-scene datasets show that IMU-to-4D yields more coherent and temporally stable results than SoTA cascaded pipelines, suggesting wearable motion sensors alone can support rich 4D understanding.

4D感知可穿戴传感非视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。