arXiv:2609.02664cs.CV2026-09

用关键词重写提升4D高斯表示中复杂物体分割的准确性

Query Rewriting for Complex Object Segmentation in 4D Gaussian Representations

论文配图:Query Rewriting for Complex Object Segmentation in 4D Gaussian Representations
图 1 · 摘自论文原文
  • 将长句描述转为简洁关键词,降低语言噪声
  • 时间定位准确率从60.92%提至92.21%,空间分割vIoU从20.08%升至76.94%
  • 无需微调,适合动态场景理解任务的研究与应用

近期的4D高斯表示框架在语言引导的动态场景理解中表现出色,但对冗长、叙述性查询(含噪声上下文信息)仍高度敏感。本文研究了查询重写在4D高斯表示中复杂物体分割的影响。受检索增强语言模型和关键词引导查询重构的启发,我们提出一种无需训练的重解释策略,将长篇描述性查询转化为简洁的关键词锚定形式。该方法逐步减少语言噪声,同时保留与物体中心表示相关的语义锚点。在HyperNeRF和Neu3D上的实验表明,简洁重写后的查询显著提升了时序定位与空间分割性能。具体而言,平均时序准确率从60.92%提升至92.21%,平均vIoU从20.08%提高到76.94%,且无需额外微调。大量消融实验进一步显示,更短、以关键词为导向的查询能保持稳定的视频特征相似性分布,并更好对齐物体中心高斯表示。

原文摘要 · Abstract (English)

Recent 4D Gaussian representation frameworks have demonstrated strong performance in language-guided dynamic scene understanding. However, these methods remain highly sensitive to verbose and narrative-style queries that contain noisy contextual information. In this paper, we investigate the impact of query rewriting for complex object segmentation in 4D Gaussian representations. Inspired by recent findings in retrieval-augmented language models and keyword-guided query reformulation, we propose a training-free reinterpretation strategy that transforms long descriptive queries into concise keyword-grounded forms. Our approach progressively reduces linguistic noise while preserving semantic anchors relevant to object-centric representations. Experiments on HyperNeRF and Neu3D demonstrate that concise rewritten queries significantly improve both temporal localization and spatial segmentation performance. In particular, our method improves average temporal accuracy from 60.92% to 92.21% and average vIoU from 20.08% to 76.94% without any additional fine-tuning. Extensive ablation studies further reveal that shorter, keyword-focused queries consistently yield stable video-feature similarity distributions and better alignment with object-centric Gaussian representations

4D高斯物体分割查询重写

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。