arXiv:2602.13751cs.CV2026-02被引 2

构建首个针对文本到动作生成的分布外评测基准,揭示现有模型在复杂场景下的短板。

T2MBench: A Benchmark for Out-of-Distribution Text-to-Motion Generation

  • 设计1025条分布外文本描述,构建专用于OOS评估的基准数据集
  • 通过多维度评测框架发现多数模型在精细动作准确性上表现不佳
  • 适合关注生成模型泛化能力与真实应用落地的研究者参考

现有文本到动作生成的评估大多局限于分布内文本输入和有限评价指标,难以系统评估模型在复杂分布外(OOD)文本条件下的泛化与生成能力。为解决这一问题,我们提出一个专门针对分布外文本到动作生成的评测基准,包含对14个代表性基线模型的全面分析及基于评估结果衍生的两个数据集。我们构建了一个包含1025条文本描述的分布外提示数据集,并在此基础上引入统一评估框架,整合大语言模型评估、多因素动作评估与细粒度准确率评估。实验结果显示,尽管不同基线模型在语义对齐、动作泛化性和物理合理性方面各有优势,但多数模型在细粒度准确率评估中表现较弱。这些发现揭示了现有方法在分布外场景下的局限性,为未来生产级文本到动作模型的设计与评估提供了实用指导。

原文摘要 · Abstract (English)

Most existing evaluations of text-to-motion generation focus on in-distribution textual inputs and a limited set of evaluation criteria, which restricts their ability to systematically assess model generalization and motion generation capabilities under complex out-of-distribution (OOD) textual conditions. To address this limitation, we propose a benchmark specifically designed for OOD text-to-motion evaluation, which includes a comprehensive analysis of 14 representative baseline models and the two datasets derived from evaluation results. Specifically, we construct an OOD prompt dataset consisting of 1,025 textual descriptions. Based on this prompt dataset, we introduce a unified evaluation framework that integrates LLM-based Evaluation, Multi-factor Motion evaluation, and Fine-grained Accuracy Evaluation. Our experimental results reveal that while different baseline models demonstrate strengths in areas such as text-to-motion semantic alignment, motion generalizability, and physical quality, most models struggle to achieve strong performance with Fine-grained Accuracy Evaluation. These findings highlight the limitations of existing methods in OOD scenarios and offer practical guidance for the design and evaluation of future production-level text-to-motion models.

文本生成动作合成评测基准泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。