arXiv:2512.17012cs.CV2025-12被引 2

让大模型理解视频中物体的时空变化,支持区域级提问。

4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation

  • 用感知蒸馏技术从专家模型学习4D视觉表示。
  • 在新基准上性能超越现有方法,提升显著。
  • 适合做视频理解、动态场景分析的研究者。

尽管多模态大模型(MLLMs)取得进展,其对3D结构和时序动态的推理能力仍受限于薄弱的4D感知与时间理解。现有3D及4D视频问答(VQA)基准侧重静态场景,缺乏区域级提示。本文提出:(a) 4D-RGPT,一种专用于从视频输入捕捉4D表示并增强时序感知的MLLM;(b) 感知4D蒸馏(P4D),一种将冻结专家模型的4D表示迁移至4D-RGPT的训练框架;(c) R4D-Bench,一个基于混合自动化与人工验证流程构建的、支持深度感知动态场景与区域级提示的新基准。4D-RGPT在现有4D VQA基准及所提R4D-Bench上均取得显著提升。

原文摘要 · Abstract (English)

Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains limited, constrained by weak 4D perception and temporal understanding. Existing 3D and 4D Video Question Answering (VQA) benchmarks also emphasize static scenes and lack region-level prompting. We tackle these issues by introducing: (a) 4D-RGPT, a specialized MLLM designed to capture 4D representations from video inputs with enhanced temporal perception; (b) Perceptual 4D Distillation (P4D), a training framework that transfers 4D representations from a frozen expert model into 4D-RGPT for comprehensive 4D perception; and (c) R4D-Bench, a benchmark for depth-aware dynamic scenes with region-level prompting, built via a hybrid automated and human-verified pipeline. Our 4D-RGPT achieves notable improvements on both existing 4D VQA benchmarks and the proposed R4D-Bench benchmark.

视频理解4D感知多模态蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。