从视频中学习可视觉解释的软体机器人动力学模型。
Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video
- 用注意力广播解码器生成像素级注意力图,定位潜变量贡献并过滤背景。
- 构建2D潜变振荡网络,实现质量、刚度和力的图像化可视化,误差降低3.5倍以上。
- 无需先验知识,自动发现振子链结构,适合机器人控制与可解释建模研究。
从视频中学习软体连续机器人(SCR)动力学具有灵活性,但现有方法缺乏可解释性或依赖先验假设。基于模型的方法需要先验知识和手动设计。本文提出:(1) 注意力广播解码器(ABCD),一种可即插即用的模块,用于基于自编码器的潜空间动力学学习,能生成像素级注意力图,定位每个潜变量的贡献并过滤静态背景,实现空间对齐的潜变量与图像叠加的可解释性;(2) 视觉振荡网络(VONs),将2D潜空间振荡网络与ABCD注意力图结合,实现对学习到的质量、耦合刚度和力的图像化可视化,提升机械可解释性。在单段与双段SCR上验证,基于ABCD的模型显著提升多步预测精度,在双段机器人上使Koopman算子误差降低5.8倍,振荡网络误差降低3.5倍。VONs自主发现振子链结构。该全数据驱动方法生成紧凑且具机械可解释性的模型,适用于未来控制应用。
原文摘要 · Abstract (English)
Learning soft continuum robot (SCR) dynamics from video offers flexibility but existing methods lack interpretability or rely on prior assumptions. Model-based approaches require prior knowledge and manual design. We bridge this gap by introducing: (1) The Attention Broadcast Decoder (ABCD), a plug-and-play module for autoencoder-based latent dynamics learning that generates pixel-accurate attention maps localizing each latent dimension's contribution while filtering static backgrounds, enabling visual interpretability via spatially grounded latents and on-image overlays. (2) Visual Oscillator Networks (VONs), a 2D latent oscillator network coupled to ABCD attention maps for on-image visualization of learned masses, coupling stiffness, and forces, thereby enabling mechanical interpretability. We validate our approach on single- and double-segment SCRs, demonstrating that ABCD-based models significantly improve multi-step prediction accuracy with 5.8x error reduction for Koopman operators and 3.5x for oscillator networks on a two-segment robot. VONs autonomously discover a chain structure of oscillators. This fully data-driven approach yields compact, mechanically interpretable models with potential relevance for future control applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。