从片段级标注中学习动作与文本的细粒度对齐,提升动作生成精度。
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

- 将长描述切分为动作语义短语,通过最优传输建模帧与短语的多对多关系
- 在SnapMoGen数据集上,对齐效果优于基线方法,显著提升文本-动作关联能力
- 无需人工标注帧级对齐,适合大规模动作生成任务
文本条件的人体动作生成因大规模动作-语言数据集的发展取得快速进展。然而,即使具有丰富长文本描述的数据集,通常也仅提供片段级别的监督,缺乏动作帧与语言之间的显式时间对应关系,限制了细粒度的文本-动作定位与精确时序生成。我们提出FineMoLA,一种弱监督框架,可直接从片段级标注中学习细粒度的帧-短语对应关系。该方法首先将长文本分割为承载动作语义的短语,然后将动作-语言对齐建模为带全局约束的最优传输问题,自然捕捉帧与文本间的多对多关系。结合熵正则化与Sinkhorn迭代,FineMoLA高效推断出伪帧级对齐,无需人工标注。在SnapMoGen上的实验表明,所学对齐在文本-动作定位任务上优于基线方法。
原文摘要 · Abstract (English)
Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form descriptions typically provide supervision only at the clip level, without explicit temporal correspondence between motion frames and language. This limits fine-grained motion--text grounding and temporally precise generation. We propose FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations. Our method first segments long-form descriptions into action-bearing phrases, and then formulates motion--language alignment as an optimal transport problem, which naturally models many-to-many relations between motion frames and text under global constraints. With entropic regularization and Sinkhorn iterations, FineMoLA efficiently infers pseudo frame-level alignments without human labeling. Experiments on SnapMoGen demonstrate that the learned alignments outperform baselines in motion--text grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。