arXiv:2507.17664cs.CVcs.RO2025-07NeurIPS被引 6

首个基于事件相机的语言引导物体定位基准,支持动态场景理解。

Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras

  • 构建多属性事件流语言锚定框架,融合外观、状态等四类信息。
  • 在3万+真实驾驶数据上实现超越现有方法的定位精度。
  • 适合机器人感知、自动驾驶等需要时序敏感语言理解的场景。

事件相机具备微秒级延迟和抗运动模糊特性,适合动态环境感知。然而,如何将异步事件流与人类语言关联仍是开放挑战。我们提出Talk2Event,首个面向事件感知的语言驱动物体定位大规模基准,基于真实驾驶数据构建,包含超3万条经验证的指代表达,每条均标注外观、状态、与观察者关系及与其他物体关系四类属性,打通空间、时间与关系推理。为充分挖掘这些线索,我们设计EventRefer框架,通过事件-属性专家混合(MoEE)动态融合多属性表示。该方法适应不同模态与场景动态,在纯事件、纯帧及事件-帧融合设置中均优于当前最优基线。期望本数据集与方法能推动现实世界机器人与自主系统中多模态、时序敏感、语言驱动感知的发展。

原文摘要 · Abstract (English)

Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challenge. We introduce Talk2Event, the first large-scale benchmark for language-driven object grounding in event-based perception. Built from real-world driving data, we provide over 30,000 validated referring expressions, each enriched with four grounding attributes -- appearance, status, relation to viewer, and relation to other objects -- bridging spatial, temporal, and relational reasoning. To fully exploit these cues, we propose EventRefer, an attribute-aware grounding framework that dynamically fuses multi-attribute representations through a Mixture of Event-Attribute Experts (MoEE). Our method adapts to different modalities and scene dynamics, achieving consistent gains over state-of-the-art baselines in event-only, frame-only, and event-frame fusion settings. We hope our dataset and approach will establish a foundation for advancing multimodal, temporally-aware, and language-driven perception in real-world robotics and autonomy.

事件相机语言理解物体定位动态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。