arXiv:2606.26916cs.CV2026-06中稿 · ECCV被引 1

用检索增强生成提升视频生成的物理合理性

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation

论文配图:PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
图 1 · 摘自论文原文
  • 基于检索增强生成,从物理数据库中注入知识到扩散模型
  • 筛选出7000条高质量物理视频,提升模型对真实物理规律的遵循
  • 适合关注物理仿真与视频生成交叉研究的学者

构建物理感知视频生成模型面临挑战,因难以捕捉热力学、力学、光学等多样物理现象。本文提出PhysRAG,通过检索增强生成(RAG)提升视频生成的物理意识。针对高质量数据稀缺问题,基于WISA-80K数据集设计两阶段过滤流程,获得7000条高质视频用于训练。同时构建物理视频数据库,并开发可学习查询机制,将物理知识注入视频扩散模型。在PhyGenBench和VBench等基准上,该方法在视觉质量和物理规则符合度上均达领先水平。通过大量消融实验验证了数据过滤、RAG机制及物理信息提取的有效性。代码、数据与模型将开源发布于https://github.com/sediment1024/PhysRAG。

原文摘要 · Abstract (English)

Developing physically aware video generation models remains a significant challenge due to the difficulty in capturing diverse physical phenomena, such as thermal dynamics, mechanics, and optics. In this work, we introduce PhysRAG, a novel pipeline that enhances physical awareness in video generation through Retrieval-Augmented Generation (RAG). To address the issue of limited high-quality data, we design a two-stage data filtering pipeline based on the WISA-80K dataset, resulting in a curated set of 7K high-quality videos for training. Furthermore, we construct a physical video database and develop a mechanism to inject physical knowledge into a video diffusion model using learnable queries. Our method achieves state-of-the-art performance in both visual quality and physical rule compliance, surpassing existing models in benchmarks such as PhyGenBench and VBench. We conduct extensive ablation studies to validate the effectiveness of our key components, including the data filtering pipeline, RAG mechanism, and method for physical information extraction. To facilitate future research, our code, data, and models are prepared for release at https://github.com/sediment1024/PhysRAG.

视频生成物理感知扩散模型RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。