用量子电路压缩3D场景建模,参数量少一半还能画得更好。
QNeRF: Neural Radiance Fields on a Simulated Gate-Based Quantum Computer
- 用量子叠加和纠缠编码视角与空间信息,压缩模型规模。
- 在中等分辨率图像上训练,参数少于一半,性能不输甚至超越经典模型。
- 适合对轻量化3D建模有需求的研究者或量子计算初探者。
最近,量子视觉场(QVFs)在学习2D或3D信号时展现出模型更紧凑、收敛更快的优势。与此同时,神经辐射场(NeRFs)在新视角合成方面取得显著进展,通过从2D图像学习紧凑表示来渲染3D场景,但代价是模型庞大且训练耗时。本文提出QNeRF,首个用于从2D图像进行新视角合成的混合量子-经典模型。QNeRF利用参数化量子电路,通过量子叠加和纠缠编码空间与视角依赖信息,相比经典模型实现更紧凑的结构。我们设计两种架构:全量子版最大化利用所有量子振幅以增强表征能力;双分支版通过分离空间与视角的量子态制备引入任务感知先验,大幅降低操作复杂度,提升可扩展性与硬件兼容性。实验表明,在中等分辨率图像上训练时,QNeRF在参数量不足经典模型一半的情况下,仍能匹配或超越其性能。结果表明,量子机器学习可作为计算机视觉中层任务(如从2D观测学习3D表示)的有力替代方案。
原文摘要 · Abstract (English)
Recently, Quantum Visual Fields (QVFs) have shown promising improvements in model compactness and convergence speed for learning the provided 2D or 3D signals. Meanwhile, novel-view synthesis has seen major advances with Neural Radiance Fields (NeRFs), where models learn a compact representation from 2D images to render 3D scenes, albeit at the cost of larger models and intensive training. In this work, we extend the approach of QVFs by introducing QNeRF, the first hybrid quantum-classical model designed for novel-view synthesis from 2D images. QNeRF leverages parameterised quantum circuits to encode spatial and view-dependent information via quantum superposition and entanglement, resulting in more compact models compared to the classical counterpart. We present two architectural variants. Full QNeRF maximally exploits all quantum amplitudes to enhance representational capabilities. In contrast, Dual-Branch QNeRF introduces a task-informed inductive bias by branching spatial and view-dependent quantum state preparations, drastically reducing the complexity of this operation and ensuring scalability and potential hardware compatibility. Our experiments demonstrate that -- when trained on images of moderate resolution -- QNeRF matches or outperforms classical NeRF baselines while using less than half the number of parameters. These results suggest that quantum machine learning can serve as a competitive alternative for continuous signal representation in mid-level tasks in computer vision, such as 3D representation learning from 2D observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。