arXiv:2512.07230cs.CV2025-12中稿 · WACV 2026被引 1

让3D场景中的文字更清晰可读,提升重建精度。

STRinGS: Selective Text Refinement in Gaussian Splatting

  • 分区域优化:先精修文字区,再融合非文字区
  • 7000次迭代下文字识别错误率降低63.6%
  • 专为含文字场景设计,适合智能导航与虚拟现实

文本作为标识、标签或指令,是真实场景中传递关键上下文信息的重要元素。现有的3D表示方法如3D高斯点阵(3DGS)难以在保持高视觉保真度的同时保留细粒度文本细节,微小的文本重建误差会导致显著语义丢失。本文提出STRinGS——一种面向文本的有选择性精修框架,对文本与非文本区域分别处理:优先精修文本区域,随后与非文本区域合并进行全场景优化。该方法能在复杂配置下生成清晰可读的文本。我们引入基于OCR字符错误率(CER)的可读性评估指标,并在新构建的多场景数据集STRinGS-360上验证效果。STRinGS在仅7000次迭代下相较原始3DGS实现63.6%的相对性能提升,推动了文本丰富环境下的3D场景理解边界,为更鲁棒的文本感知重建方法奠定基础。

原文摘要 · Abstract (English)

Text as signs, labels, or instructions is a critical element of real-world scenes as they can convey important contextual information. 3D representations such as 3D Gaussian Splatting (3DGS) struggle to preserve fine-grained text details, while achieving high visual fidelity. Small errors in textual element reconstruction can lead to significant semantic loss. We propose STRinGS, a text-aware, selective refinement framework to address this issue for 3DGS reconstruction. Our method treats text and non-text regions separately, refining text regions first and merging them with non-text regions later for full-scene optimization. STRinGS produces sharp, readable text even in challenging configurations. We introduce a text readability measure OCR Character Error Rate (CER) to evaluate the efficacy on text regions. STRinGS results in a 63.6% relative improvement over 3DGS at just 7K iterations. We also introduce a curated dataset STRinGS-360 with diverse text scenarios to evaluate text readability in 3D reconstruction. Our method and dataset together push the boundaries of 3D scene understanding in text-rich environments, paving the way for more robust text-aware reconstruction methods.

3D重建文本感知高斯点阵可读性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。