构建首个覆盖多场景、多粒度的人体动作-文本检索基准,推动跨模态对齐评估
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

- 设计多阶段数据清洗流程,确保动作与描述语义清晰且平衡
- 包含3390段动作、10170条细粒度描述,覆盖118个精细类别
- 提出轻量级粒度感知模型,提升多粒度检索性能且不牺牲标准表现
人体动作-文本检索为评估跨模态对齐提供了严格标准。现有基准普遍存在室内动作同质化、动作分布不均、文本简单重复等问题,阻碍了跨领域和跨粒度对齐的可靠评估。为此,我们提出MRBench,一个全面的动作-文本检索基准,涵盖异构动作、广泛且均衡的类别覆盖,以及可靠、可区分的多粒度描述。MRBench通过精心设计的多阶段数据整理流程构建,包括候选筛选与平衡、语义对齐验证,以及在多个粒度层级上生成动作相关描述。最终基准包含来自动作捕捉、真实场景视频、合成视频及动作生成模型的3,390段动作,覆盖118个细粒度类别。每段动作配有简洁、标准、细粒度的描述,共生成10,170条标题。对代表性检索基线在MRBench上的广泛评估揭示了显著的跨数据集泛化差距和对查询粒度的强烈敏感性。我们提出一种基于冻结标准标题对齐检索模型的轻量级粒度感知模型。利用大语言模型生成的简洁细粒度标题作为伪监督信号,训练额外分支的粒度特定动作提取器与文本适配器。推理时,粒度感知得分融合整合全局与适配相似性,同时严格保持所有描述层级间的得分可比性。该模型在不损害标准标题性能的前提下,提升了混合粒度检索效果。我们认为MRBench为推进动作-语言对齐评估提供了全面测试平台。
原文摘要 · Abstract (English)
Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distributions, and oversimplified, repetitive texts, which hinder the reliable measurement of cross-domain and cross-granularity alignment. We thus introduce MRBench, a comprehensive motion-text retrieval benchmark featuring heterogeneous motions, broad and balanced category coverage, and reliable, discriminative, multi-granular descriptions. MRBench is constructed through a meticulously designed multi-stage data curation pipeline, which filters and balances candidates, verifies unambiguous semantic alignment, and generates motion-grounded descriptions at multiple granularities. The resulting benchmark contains 3,390 motions drawn from motion capture, in-the-wild videos, synthetic videos, and motion generative models, covering 118 fine-grained categories. Each motion is paired with concise, standard, and fine-grained descriptions, yielding 10,170 captions. Extensive evaluations of representative retrieval baselines on MRBench reveal a substantial cross-dataset generalization gap and pronounced sensitivity to query granularity. We propose a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model. LLM-based concise and fine-grained captions provide pseudo-supervision for extra-branch granularity-specific motion extractors and text adapters. For inference, granularity-aware score fusion integrates global and adapted similarities while strictly maintaining score comparability across all description levels. The resulting model improves mixed-granularity retrieval without compromising standard-caption performance. We believe that our MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。