首个跨谱视觉定位基准,支持红外+可见光融合,提升复杂环境下的定位鲁棒性。
RGBT-GroundBench: Visual Grounding Beyond RGB in Complex Real-World Scenarios
- 构建多模态(RGB+热成像)标注数据集,覆盖3类场景、6种环境与4类物体属性
- 发现低光照下性能显著下降,且模型对场景复杂度敏感,现有方法普遍不鲁棒
- 提出简洁可复现的基准模型,通过三重机制实现可靠跨模态融合,适合工业部署
视觉定位(VG)旨在从自然语言描述中定位图像中的目标对象。在真实感知场景中,低光照和恶劣天气常导致可见光信息退化,使定位任务更具挑战性。然而,现有基准大多仅基于可见光(RGB),对复杂条件的覆盖有限,难以系统评估模型鲁棒性或进行跨谱比较。本文提出 RGBT-GroundBench,首个面向复杂环境的可见光-热成像(RGB-TIR)视觉定位大规模基准。数据集包含超过4万张图像(21,535组RGB-TIR配对),38,760个物体实例,附带指代表达、边界框及三级细粒度标注:场景类型、环境条件(光照与天气)、物体属性(尺寸与遮挡)。作为基准套件,RGBT-GroundBench不仅提供精心标注的数据,还统一评估协议,支持纯可见光、纯热成像及双模态输入。在此协议下,我们对11个代表性模型在多种场景与环境下进行了评测。结果表明,定位精度与场景复杂度强相关;基于LoRA的模型在复杂场景中更鲁棒;而低光照条件导致的性能下降此前极少被研究。基于此,我们提出 RGBT-VGNet,一种简单且可复现的参考基线,在统一协议下引入非对称模态适配、语言感知视觉协同与三先验融合机制,实现可靠性导向的跨模态集成。相关资源、标注、代码、权重及评估脚本均已公开。
原文摘要 · Abstract (English)
Visual grounding (VG) localizes target objects in an image from natural-language expressions. In real-world perception, RGB cues often degrade under low illumination and adverse weather, making visual grounding substantially more challenging. However, existing VG benchmarks are largely RGB-only and provide limited, structured coverage of such conditions, hindering systematic robustness evaluation and cross-spectral comparison. We present RGBT-GroundBench, the first large-scale benchmark for RGB-Thermal (TIR) visual grounding in complex environments. It contains over 40K images (21,535 RGB-TIR pairs) and 38,760 object instances with referring expressions, bounding boxes, and fine-grained annotations at three levels: scene types, environmental conditions (illumination and weather), and object properties (size and occlusion). As a benchmark suite, RGBT-GroundBench provides not only curated RGB-TIR grounding annotations but also a unified evaluation protocol supporting RGB-only, TIR-only, and RGB+TIR inputs. Under this protocol, we benchmark 11 representative VG models across diverse scenes and environmental conditions. Our results show that grounding accuracy is strongly correlated with scene complexity, LoRA-based models are more robust in complex scenes, and low-illumination conditions cause significant performance degradation that has been rarely explored. Guided by these observations, we introduce RGBT-VGNet, a simple and reproducible reference baseline under the unified protocol, featuring Asymmetric Modality Adaptation, Language-Aware Visual Synergy, and Tri-Prior Fusion for reliability-aware RGB-TIR integration. Resources, annotations, code, checkpoints, and evaluation scripts have been publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。