arXiv:2508.02512cs.ROcs.CV2025-08中稿 · CoRL被引 8

为四足机器人生成可控的全景视频,解决真实数据稀缺问题。

QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots

  • 模拟四足机器人运动特性,通过垂直抖动编码生成真实感全景视频。
  • 在360度场景中提升多目标追踪性能,显著改善感知效果。
  • 适合研究机器人视觉感知、仿真数据生成与运动控制的学者。

全景相机可捕捉360度环境信息,适用于四足机器人在复杂环境中的感知与交互。然而,由于固有的运动学限制和复杂的传感器校准挑战,高质量全景训练数据匮乏,严重制约了面向此类实体平台的鲁棒感知系统发展。为此,我们提出QuaDreamer——首个专为四足机器人设计的全景数据生成引擎。该模型通过模仿四足机器人运动模式,生成高度可控且逼真的全景视频,为下游任务提供数据支持。为有效捕捉四足行走时特有的垂直振动特征,引入垂直抖动编码(VJE),通过频域特征过滤提取可控的垂直信号并生成高质量提示。为实现抖动信号控制下的高质量全景视频生成,提出场景-物体控制器(SOC),利用注意力机制有效管理物体运动,并增强背景抖动控制。针对广视角视频生成中的全景畸变问题,提出全景增强器(PE)——一种双流架构,协同进行局部细节的频域-纹理优化与全局几何一致性空间-结构修正。进一步验证表明,生成的视频序列可作为四足机器人全景视觉感知模型的训练数据,显著提升360度场景中的多目标追踪性能。源代码与模型权重将公开于 https://github.com/losehu/QuaDreamer。

原文摘要 · Abstract (English)

Panoramic cameras, capturing comprehensive 360-degree environmental data, are suitable for quadruped robots in surrounding perception and interaction with complex environments. However, the scarcity of high-quality panoramic training data-caused by inherent kinematic constraints and complex sensor calibration challenges-fundamentally limits the development of robust perception systems tailored to these embodied platforms. To address this issue, we propose QuaDreamer-the first panoramic data generation engine specifically designed for quadruped robots. QuaDreamer focuses on mimicking the motion paradigm of quadruped robots to generate highly controllable, realistic panoramic videos, providing a data source for downstream tasks. Specifically, to effectively capture the unique vertical vibration characteristics exhibited during quadruped locomotion, we introduce Vertical Jitter Encoding (VJE). VJE extracts controllable vertical signals through frequency-domain feature filtering and provides high-quality prompts. To facilitate high-quality panoramic video generation under jitter signal control, we propose a Scene-Object Controller (SOC) that effectively manages object motion and boosts background jitter control through the attention mechanism. To address panoramic distortions in wide-FoV video generation, we propose the Panoramic Enhancer (PE)-a dual-stream architecture that synergizes frequency-texture refinement for local detail enhancement with spatial-structure correction for global geometric consistency. We further demonstrate that the generated video sequences can serve as training data for the quadruped robot's panoramic visual perception model, enhancing the performance of multi-object tracking in 360-degree scenes. The source code and model weights will be publicly available at https://github.com/losehu/QuaDreamer.

全景视频四足机器人数据生成视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。