arXiv:2603.25892cs.CV2026-03被引 3

一个模型搞定人体4D感知,用文本控制完成多种视觉任务。

THFM: A Unified Video Foundation Model for 4D Human Perception and Beyond

  • 基于文生视频扩散模型,统一处理稠密与稀疏感知任务。
  • 仅在合成数据训练,性能媲美甚至超越专用模型。
  • 能泛化到多人或多类物体,展现新奇涌现能力。

我们提出THFM,一种面向人体中心感知的统一视频基础模型,可在单一架构中联合解决密集任务(深度、法向、分割、稠密姿态)和稀疏任务(2D/3D关键点估计)。THFM源自预训练的文本到视频扩散模型,被改造为单次前向传播的感知模型,并通过可学习标记增强稀疏预测能力。受文本提示调制,该模型可执行多种感知任务。关键的是,尽管仅在合成数据上训练(未使用真实世界或特定基准数据),其在多个基准上的表现仍达到或超过当前最优专用模型。我们进一步揭示了模型的有趣涌现特性,归因于基于扩散的视频表示。例如,仅在单人场景视频上训练的模型,能泛化至多人及其它对象类别,如拟人角色和动物,这一能力此前未被实现。

原文摘要 · Abstract (English)

We present THFM, a unified video foundation model for human-centric perception that jointly addresses dense tasks (depth, normals, segmentation, dense pose) and sparse tasks (2d/3d keypoint estimation) within a single architecture. THFM is derived from a pretrained text-to-video diffusion model, repurposed as a single-forward-pass perception model and augmented with learnable tokens for sparse predictions. Modulated by the text prompt, our single unified model is capable of performing various perception tasks. Crucially, our model is on-par or surpassing state-of-the-art specialized models on a variety of benchmarks despite being trained exclusively on synthetic data (i.e.~without training on real-world or benchmark specific data). We further highlight intriguing emergent properties of our model, which we attribute to the underlying diffusion-based video representation. For example, our model trained on videos with a single human in the scene generalizes to multiple humans and other object classes such as anthropomorphic characters and animals -- a capability that hasn't been demonstrated in the past.

视频生成扩散模型人体感知多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。