arXiv:2605.14795cs.CV2026-05

通过增强观察与反事实学习,提升复杂场景下多目标指代追踪的精准度。

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking

论文配图:COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking
图 1 · 摘自论文原文
  • 用视觉语言模型注入语义信息,扩大观察空间以增强区分性。
  • 利用大模型推理生成反事实监督,强制验证属性特征,提升组合识别能力。
  • 框架融合外部知识,适合高同质复杂场景下的多目标追踪研究。

指代多目标追踪(RMOT)面临高区分度需求与稀疏语义监督之间的根本矛盾,尤其在高度同质、需精细区分复杂组合语义的场景中更为突出。稀疏监督导致模型过拟合于显著但不充分的线索,引发捷径学习和语义坍塌。为此,本文提出COAL(反事实与观测增强对齐学习)框架,通过知识正则化突破孤立结构优化。首先,引入显式语义注入(ESI),借助视觉语言模型丰富观测空间,提升实例区分性;其次,基于大语言模型推理设计反事实学习(CFL),扩充监督信号,强制属性验证以增强鲁棒的组合识别能力。二者统一于分层多流融合(HMSI)架构,将外部知识提炼为领域特定的判别表示。在Refer-KITTI和Refer-KITTI-V2基准上的实验表明,该方法在极具挑战性的Refer-KITTI-V2上超越现有最优模型7.28% HOTA,验证了知识正则化对解决稀疏性-区分性悖论的有效性。

原文摘要 · Abstract (English)

Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic supervision. This mismatch is particularly acute in highly homogeneous scenarios that require fine-grained discrimination over complex compositional semantics. However, under sparse supervision, models overfit to salient yet insufficient cues, thereby encouraging shortcut learning and semantic collapse. To resolve this, we propose COAL (Counterfactual and Observation-enhanced Alignment Learning), a framework that advances RMOT beyond isolated structural optimization through knowledge regularization. First, we introduce Explicit Semantic Injection (ESI) via a VLM to densify the observation space and enhance instance discriminability. Second, leveraging LLM reasoning, we propose Counterfactual Learning (CFL) to augment supervision, enforcing strict attribute verification for robust compositional recognition. These strategies are unified within a Hierarchical Multi-Stream Integration (HMSI) architecture, which distills external knowledge into domain-specific discriminative representations. Experiments on Refer-KITTI and Refer-KITTI-V2 benchmarks validate COAL's efficacy. Notably, it surpasses the state-of-the-art by 7.28% HOTA on the highly challenging Refer-KITTI-V2. These results demonstrate the effectiveness of knowledge regularization for resolving the sparsity-discriminability paradox in RMOT.

多目标追踪指代理解视觉语言模型知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。