arXiv:2602.20354cs.CV2026-02

用3D语义点自编码器自动评估视频真实感,无需参考视频。

3DSPA: A 3D Semantic Point Autoencoder for Evaluating Video Realism

  • 融合3D轨迹、深度与语义特征,构建统一视频表征
  • 能识别违反物理规律的视频,对运动伪影更敏感
  • 评估结果更贴近人类判断,适合生成视频模型基准测试

AI视频生成技术快速发展。为使生成视频适用于机器人、影视制作等场景,必须持续产出逼真视频。然而,当前视频真实感评估仍依赖人工标注或特定数据集,覆盖范围有限。本文提出3DSPA——一种无需参考视频的自动化视频真实感评估框架。该方法基于3D时空点自编码器,整合3D点轨迹、深度信息与DINO语义特征,构建统一表征,建模物体运动与场景语义变化,实现对真实性、时序一致性与物理合理性的真实评估。实验表明,3DSPA能有效识别违反物理规律的视频,对运动伪影更敏感,且在多个数据集上与人类判断高度一致。结果证明,在轨迹表征中融入3D语义可为生成视频模型提供更强评估基础,并隐式捕捉物理规则违反。代码与预训练权重将开源于https://github.com/TheProParadox/3dspa_code。

原文摘要 · Abstract (English)

AI video generation is evolving rapidly. For video generators to be useful for applications ranging from robotics to film-making, they must consistently produce realistic videos. However, evaluating the realism of generated videos remains a largely manual process -- requiring human annotation or bespoke evaluation datasets which have restricted scope. Here we develop an automated evaluation framework for video realism which captures both semantics and coherent 3D structure and which does not require access to a reference video. Our method, 3DSPA, is a 3D spatiotemporal point autoencoder which integrates 3D point trajectories, depth cues, and DINO semantic features into a unified representation for video evaluation. 3DSPA models how objects move and what is happening in the scene, enabling robust assessments of realism, temporal consistency, and physical plausibility. Experiments show that 3DSPA reliably identifies videos which violate physical laws, is more sensitive to motion artifacts, and aligns more closely with human judgments of video quality and realism across multiple datasets. Our results demonstrate that enriching trajectory-based representations with 3D semantics offers a stronger foundation for benchmarking generative video models, and implicitly captures physical rule violations. The code and pretrained model weights will be available at https://github.com/TheProParadox/3dspa_code.

视频评估3D表示自编码器真实感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。