arXiv:2605.15497cs.CVcs.GR2026-05

从非人类角色视频中生成可编辑的人类动作,保留原动作核心动态。

AnyAct: Towards Human Reenactment of Character Motion From Video

论文配图:AnyAct: Towards Human Reenactment of Character Motion From Video
图 1 · 摘自论文原文
  • 用稀疏局部2D关节运动作为桥梁,实现跨角色动作迁移。
  • 在自建基准上生成高保真人类重演动作,保留原视频动作特征。
  • 适合动画创作、虚拟角色驱动等需要快速生成人类动作的场景。

我们研究如何直接从单目非人类角色视频中生成初始人类重演动作。目标不是重建源角色本身,而是将其动作重新诠释为真实可信且可编辑的人类表演,以支持后续动画创作。该任务挑战在于:现有基于视频的动作捕获方法多局限于人体结构空间,而动作重定向方法通常需已知三维结构和拓扑。我们的关键洞察是:稀疏局部刚性运动线索可在显著结构差异下保留本质动力学,构成从角色视频到人类重演的稳定桥梁。基于此,提出AnyAct,将角色视频驱动的人类重演建模为从可迁移的稀疏局部2D关节运动出发的条件人类动作生成。为实现这一目标,引入三项关键设计:仅使用人体动作的增强3D到2D投影监督、渐进式3D到2D训练以缓解条件模糊性、全局-局部运动解耦以实现可靠局部控制。进一步构建了涵盖多样非人类角色视频的基准数据集。实验表明,AnyAct能生成高保真初始人类重演,有效保留参考视频中角色的核心动力学;消融实验证明其核心设计的有效性。

原文摘要 · Abstract (English)

We study the problem of directly deriving an initial human reenactment from a monocular video of a non-human character. Our goal is not to reconstruct the source character itself but to reinterpret its motion as a plausible and editable human performance for downstream animation authoring. This task is challenging because existing video-based motion capture methods are largely restricted to human-centric structural spaces, while motion retargeting methods typically require structured 3D source motions and known source topologies. Our key insight is that sparse local articulated motion cues can preserve essential dynamics across large structural differences, providing a stable bridge from character video to human reenactment. Based on this observation, we propose AnyAct, which formulates character-video-driven human reenactment as conditional human motion generation from transferable sparse local 2D articulated motion. To make this practical, we introduce three key designs: human-motion-only supervision via augmented 3D-to-2D projection, progressive 3D-to-2D training to alleviate conditioning ambiguity, and global-local motion decoupling for reliable local motion control. We further construct a benchmark primarily covering diverse non-human character videos. Experiments on the benchmark show that AnyAct produces high-fidelity initial human reenactments that preserve the essential dynamics of the characters in reference videos, and further ablation studies validate the effectiveness of its core designs.

动作重演视频驱动动作生成非人角色

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。