arXiv:2604.09639cs.CV2026-04

无需相机位姿即可实现多视角3D风格迁移,保持几何结构稳定。

3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation

论文配图:3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation
图 1 · 摘自论文原文
  • 基于特征匹配的跨视角一致性损失,避免纹理漂移
  • 结合深度与结构约束,在Tanks and Temples等数据集上提升重建精度
  • 适合需要保留几何信息的3D内容创作与视觉任务

艺术风格迁移在图像和视频中已较成熟,但扩展至多视角3D场景仍具挑战,因风格化可能破坏几何感知流程所需的对应关系。独立的单视图风格化常导致纹理漂移、边缘扭曲和阴影不一致,影响SLAM、深度预测与多视图重建。本文提出一种无需假设相机位姿或显式3D表示的前馈式风格化网络,采用场景级测试时优化,联合外观迁移与几何保持。通过冻结的VGG-19编码器设计类似AdaIN的损失,匹配风格图像的通道统计量;为稳定多视角结构,引入基于SuperPoint与SuperGlue的对应一致性损失,约束风格化锚视图描述子与原始多视图匹配描述子的一致性。同时使用MiDaS/DPT进行深度保持,并通过全局颜色对齐减少深度模型域偏移。采用分阶段权重调度逐步引入几何与深度约束。在Tanks and Temples与Mip-NeRF 360数据集上评估,以图像与重建指标衡量。风格保真度与结构保留分别通过色彩直方图距离(CHD)与结构距离(DSD)评估;3D一致性通过单目DROID-SLAM轨迹与反投影点云的对称切尔姆夫距离衡量。消融实验表明,对应与深度正则化显著降低结构失真,提升SLAM稳定性与重建几何质量;相较于MuVieCAST基线,本方法在保持竞争性风格化的同时,实现更优的轨迹与点云一致性。

原文摘要 · Abstract (English)

Artistic style transfer is well studied for images and videos, but extending it to multi-view 3D scenes remains difficult because stylization can disrupt correspondences needed by geometry-aware pipelines. Independent per-view stylization often causes texture drift, warped edges, and inconsistent shading, degrading SLAM, depth prediction, and multi-view reconstruction. This thesis addresses multi-view stylization that remains usable for downstream 3D tasks without assuming camera poses or an explicit 3D representation during training. We introduce a feed-forward stylization network trained with per-scene test-time optimization under a composite objective coupling appearance transfer with geometry preservation. Stylization is driven by an AdaIN-inspired loss from a frozen VGG-19 encoder, matching channel-wise moments to a style image. To stabilize structure across viewpoints, we propose a correspondence-based consistency loss using SuperPoint and SuperGlue, constraining descriptors from a stylized anchor view to remain consistent with matched descriptors from the original multi-view set. We also impose a depth-preservation loss using MiDaS/DPT and use global color alignment to reduce depth-model domain shift. A staged weight schedule introduces geometry and depth constraints. We evaluate on Tanks and Temples and Mip-NeRF 360 using image and reconstruction metrics. Style adherence and structure retention are measured by Color Histogram Distance (CHD) and Structure Distance (DSD). For 3D consistency, we use monocular DROID-SLAM trajectories and symmetric Chamfer distance on back-projected point clouds. Across ablations, correspondence and depth regularization reduce structural distortion and improve SLAM stability and reconstructed geometry; on scenes with MuVieCAST baselines, our method yields stronger trajectory and point-cloud consistency while maintaining competitive stylization.

3D风格迁移几何保持多视角一致性SLAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。