arXiv:2606.27325cs.CV2026-06

提出分层动作建模方法,提升高自由度操作的视频预测精度。

Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model

论文配图:Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model
图 1 · 摘自论文原文
  • 将动作序列分维处理,避免统一压缩导致的信息失真
  • 在EgoDex和EgoVerse上FID、FVD、PCK均显著提升
  • 适合需要精细动作建模的机器人控制与虚拟交互场景

近期动作条件世界模型在复杂交互建模和多动作序列未来状态预测方面取得进展,但动作条件机制仍缺乏深入探索。现有方法常将整个动作序列压缩为单一表示,在低自由度控制中表现良好,但在高自由度场景下可靠性下降。我们观察到高自由度灵巧动作具有内在异质性,大尺度运动与细微信号共存,统一聚合导致优化不平衡,影响细粒度效应建模和动作保真度。为此,提出DexAC-WM,将动作条件视为结构化过程:通过动作分词保留维度级语义,并结合局部精炼与全局调制对齐动作信号与视觉动态。为进一步解决现有模型高层语义支撑不足问题,引入语义分支提供丰富的物体-场景先验,使模型既能捕捉动态视觉细节,又能支持高自由度动作条件视频预测。在EgoDex和EgoVerse上的实验表明,结合语义分支与DexAC显著提升FID、FVD和PCK指标,验证了视觉-时序真实性和动作跟随一致性。进一步验证表明DexAC可适配多种骨干网络,证明其结构化动作建模设计具备可扩展性。结果表明,实现高自由度控制需兼顾结构化动作建模与语义接地。

原文摘要 · Abstract (English)

Recent advances in action-conditioned world models show promising progress in modeling complex interactions and forecasting future states under diverse action sequences. While these models are often driven by stronger visual representations and model capacity, action conditioning itself remains underexplored. Most existing approaches compress the entire action sequence into a single representation, which works well for low-DoF control but becomes less reliable in high-DoF scenarios. We observe that high-DoF dexterous actions are inherently heterogeneous, spanning multiple orders of magnitude, where large-scale motions coexist with subtle but important signals. When uniformly aggregated, optimization exhibits an imbalance across action components, which hinders the modeling of fine-grained effects and affects action fidelity. We therefore propose DexAC-WM, which treats action conditioning as a structured process rather than global compression. DexAC preserves dimension-level semantics via action tokenization and aligns action signals with visual dynamics through local refinement and global modulation. To address the limited high-level semantic grounding in existing world models, we further introduce a semantic branch that provides rich object-scene priors, which enables world model to capture dynamic visual details while supporting high-DoF action-conditioned video prediction. Experiments on EgoDex and EgoVerse show that combining the semantic branch with DexAC significantly improves FID, FVD, and PCK, demonstrating gains in visual-temporal realism and action-following consistency. We further verify that DexAC extends to other backbones, showing the scalability of our structured action-conditioning design. These results suggest that scaling world models to high-DoF control requires both structured action modeling and semantic grounding.

世界模型动作建模灵巧操作视频预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。