arXiv:2510.03200cs.CV2025-10ICCV被引 2

首个统一建模动作、场景、文本的检索模型,实现三者精准对齐。

MonSTeR: a Unified Model for Motion, Scene, Text Retrieval

论文配图:MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
图 1 · 摘自论文原文
  • 构建统一潜在空间,融合单模态与跨模态表示
  • 在多任务检索中超越仅用单模态的模型表现
  • 支持零样本场景物体摆放与动作描述生成,适合多模态研究者

人类在复杂环境中的运动由意图驱动,但这种运动只有在周围环境支持时才能发生。尽管这一机制直观,但现有研究尚缺乏评估骨骼动作(运动)、意图(文本)与周围环境(场景)之间对齐程度的工具。本文提出MonSTeR,首个用于运动-场景-文本检索的统一模型。受高阶关系建模启发,MonSTeR通过融合单模态与跨模态表示构建统一潜在空间,捕捉模态间的复杂依赖关系,实现灵活且稳健的跨任务检索。实验表明,MonSTeR优于仅依赖单模态表示的三模态模型。通过专门用户研究验证了检索分数与人类偏好的一致性。我们还展示了其潜在空间在零样本场景物体放置和动作描述生成任务上的泛化能力。代码与预训练模型已开源至github.com/colloroneluca/MonSTeR。

原文摘要 · Abstract (English)

Intention drives human movement in complex environments, but such movement can only happen if the surrounding context supports it. Despite the intuitive nature of this mechanism, existing research has not yet provided tools to evaluate the alignment between skeletal movement (motion), intention (text), and the surrounding context (scene). In this work, we introduce MonSTeR, the first MOtioN-Scene-TExt Retrieval model. Inspired by the modeling of higher-order relations, MonSTeR constructs a unified latent space by leveraging unimodal and cross-modal representations. This allows MonSTeR to capture the intricate dependencies between modalities, enabling flexible but robust retrieval across various tasks. Our results show that MonSTeR outperforms trimodal models that rely solely on unimodal representations. Furthermore, we validate the alignment of our retrieval scores with human preferences through a dedicated user study. We demonstrate the versatility of MonSTeR's latent space on zero-shot in-Scene Object Placement and Motion Captioning. Code and pre-trained models are available at github.com/colloroneluca/MonSTeR.

多模态检索动作理解统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。