实现129帧/秒的全高清3D高斯溅射实时渲染,功耗低且面积小。
A 129FPS Full HD Real-Time Accelerator for 3D Gaussian Splatting

- 硬件友好压缩:融合迭代剪枝、渐进球谐降阶与向量量化
- 129帧/秒渲染1080p画面,能效比提升7.5倍
- 适用于AR/VR设备的轻量级部署,适合资源受限场景
在AR/VR级设备上渲染大规模无界场景受限于3D高斯溅射(3DGS)的计算、带宽和存储开销。本文提出一种低功耗、低成本的3DGS硬件加速器,可实现实时全高清图像渲染。配套的硬件友好压缩流水线结合了迭代高斯剪枝与微调、逐级球谐(SH)阶数降低以及所有SH系数和颜色的向量量化,实现51.6倍模型尺寸压缩,仅损失0.743 dB PSNR。加速器采用帧级流水线,集成基于点的剔除与投影、基于瓦片的排序与光栅化,跳过零雅可比矩阵乘法(使处理单元减少63%,计算量降低53%),并采用无比较的瓦片排序以保证确定性延迟。该设计在台积电28纳米工艺下运行于800 MHz,面积为0.66 mm²,含114.38万逻辑门和120 kB SRAM,功耗0.219 W,能效达1219 Mpixels/J,峰值吞吐率267.5 Mpixels/s,支持1080p下129帧/秒运行。相比先前3DGS加速器,面积缩小5.98倍,吞吐提升5.94倍,能效提高7.5倍。
原文摘要 · Abstract (English)
Rendering large-scale, unbounded scenes on AR/VR-class devices is constrained by the computation, bandwidth, and storage cost of 3D Gaussian Splatting (3DGS). We propose a low-power, low-cost 3DGS hardware accelerator that renders full-HD images in real time, together with a hardware-friendly compression pipeline that combines iterative Gaussian pruning and fine-tuning, progressive spherical harmonics (SH) degree reduction, and vector quantization of all SH coefficients and colors. The scheme achieves a $51.6\times$ model-size reduction with a 0.743 dB PSNR loss. The accelerator uses a frame-level pipeline that integrates point-based culling and projection with tile-based sorting and rasterization, skips zero-Jacobian matrix multiplications (reducing processing elements by 63\% and computation by 53\%), and adopts comparison-free tile-based sorting with deterministic latency. Implemented in a TSMC 28-nm process at 800 MHz, the design occupies $0.66~\text{mm}^2$ with 1.1438 M gates and 120 kB SRAM, consumes 0.219 W, and delivers 1219 Mpixels/J at 267.5 Mpixels/s, enabling 1080p at 129 FPS. Overall, it is $5.98\times$ smaller in area, $5.94\times$ higher throughput, and delivers $7.5\times$ higher energy efficiency than prior 3DGS accelerators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。