用3D相机轨迹生成真实房屋导览视频,无需专业设备
HouseTour: A Virtual Real Estate A(I)gent
- 通过扩散模型生成符合几何结构的平滑摄像机路径
- 结合3D重建与视觉语言模型,生成更准确的空间描述
- 适合房地产、旅游等行业快速制作高质量导览视频
我们提出HouseTour,一种从一组描绘三维空间的图像中生成空间感知的3D摄像机轨迹和自然语言摘要的方法。不同于现有视觉语言模型(VLM)在几何推理上的不足,我们的方法利用已知相机位姿约束扩散过程,生成平滑视频路径,并将此信息融入VLM以生成3D锚定的描述。最终视频通过3D高斯泼溅技术渲染沿路径的新视角。为支持该任务,我们构建了HouseTour数据集,包含超过1,200段带相机位姿、3D重建和房产描述的房屋导览视频。实验表明,将3D摄像机轨迹融入文本生成可显著提升性能。我们评估了单任务与端到端表现,引入新的联合评价指标。本工作实现了无需专业技能或设备即可自动生成专业级房产与旅游导览视频。
原文摘要 · Abstract (English)
We introduce HouseTour, a method for spatially-aware 3D camera trajectory and natural language summary generation from a collection of images depicting an existing 3D space. Unlike existing vision-language models (VLMs), which struggle with geometric reasoning, our approach generates smooth video trajectories via a diffusion process constrained by known camera poses and integrates this information into the VLM for 3D-grounded descriptions. We synthesize the final video using 3D Gaussian splatting to render novel views along the trajectory. To support this task, we present the HouseTour dataset, which includes over 1,200 house-tour videos with camera poses, 3D reconstructions, and real estate descriptions. Experiments demonstrate that incorporating 3D camera trajectories into the text generation process improves performance over methods handling each task independently. We evaluate both individual and end-to-end performance, introducing a new joint metric. Our work enables automated, professional-quality video creation for real estate and touristic applications without requiring specialized expertise or equipment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。