arXiv:2607.15717cs.CVcs.GR2026-07

让文字生成动作的每一帧都精准到位,解决动作顺序与时间错乱问题。

Per-Stroke Temporal Control for Text-to-Motion via Action Units and Action-Detection Guidance

论文配图:Per-Stroke Temporal Control for Text-to-Motion via Action Units and Action-Detection Guidance
图 1 · 摘自论文原文
  • 用动作单元(AU)显式控制每段动作的时间、轨迹和起止点。
  • 在StrokeBench上单段动作定位准确率显著超越现有方法,且运动质量最优。
  • 适合需要精确控制动作时序的动画设计与人机交互场景。

文本到动作生成模型能准确表达动作名称,但在动作出现时间上不可靠:如左右交替出拳常无法生成四个分离的独立动作。本文提出一种称为动作单元(Action Units, AUs)的时序事件类型,将每个动作片段(包括身体轨迹、动作类别、时间窗口及冲击时刻)作为显式条件信号。通过轻量级门控适配器,将冻结的文本到动作主干模型与AU集合对齐,注入两路信息流(逐动作标记与逐帧相位通道)。推理时,利用冻结的帧级检测器无训练梯度修正残余时间误差。在StrokeBench数据集上评估,其提示中包含动作数量、顺序、轨迹与核心帧位置,搭配经审计的动作语料库。相比最强基线接口,AU对齐显著提升单个动作定位准确率,在文本、间隔与帧级基线中达到最佳运动质量。提示中的核心帧进一步成为可调节的控制轴。

原文摘要 · Abstract (English)

Text-to-motion models are competent at the action a prompt names but unreliable at when each stroke lands: four punches alternating left and right rarely return four separable strokes. We introduce typed temporal events called Action Units (AUs) that make the individual stroke -- its body track, action class, time window, and impact timing -- an explicit conditioning signal. We ground a frozen text-to-motion backbone on the AU set through a lightweight gated adapter injecting two streams (per-stroke tokens and a per-frame phase channel), and at inference close residual timing errors with a training-free classifier gradient from a frozen frame-level detector. We measure per-stroke control on StrokeBench, whose prompts specify count, ordering, track, and core-frame placement, paired with an audited stroke corpus. AU grounding markedly raises the rate of correctly placed single strokes over the strongest prior interface, at the best motion quality among text-, interval-, and frame-level baselines. The prompted core frame emerges as a further steerable axis.

动作生成时序控制文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。