用大模型提升可穿戴惯性传感器的实时动作捕捉精度
Mojito: LLM-Aided Motion Instructor with Jitter-Reduced Inertial Tokens
- 结合大模型与惯性传感器,实现交互式动作捕捉
- 有效降低信号抖动和漂移,提升长时间追踪稳定性
- 适合需要隐私保护的实时行为分析场景
人体运动蕴含动作意图与认知过程的关键信息,现有多模态系统主要依赖语言、视觉和音频理解运动,难以捕捉3D运动中的动态力与力矩。惯性测量单元(IMUs)提供轻量、可穿戴且保护隐私的运动感知方案,但流式IMU数据处理面临无线传输不稳、传感器噪声与漂移等问题,限制其在长期实时动作捕捉(MoCap)及在线行为分析中的应用。为此,我们提出Mojito,一种融合惯性传感与大语言模型(LLMs)的智能动作代理,实现交互式动作捕捉与行为分析。
原文摘要 · Abstract (English)
Human bodily movements convey critical insights into action intentions and cognitive processes, yet existing multimodal systems primarily focused on understanding human motion via language, vision, and audio, which struggle to capture the dynamic forces and torques inherent in 3D motion. Inertial measurement units (IMUs) present a promising alternative, offering lightweight, wearable, and privacy-conscious motion sensing. However, processing of streaming IMU data faces challenges such as wireless transmission instability, sensor noise, and drift, limiting their utility for long-term real-time motion capture (MoCap), and more importantly, online motion analysis. To address these challenges, we introduce Mojito, an intelligent motion agent that integrates inertial sensing with large language models (LLMs) for interactive motion capture and behavioral analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。