首个基于事件相机的语言引导目标定位基准,支持动态场景下精准语义理解。
Visual Grounding from Event Cameras
- 构建事件数据驱动的语言定位框架,以结构化属性增强语义表达
- 包含5567个场景、超过3万条指代表达,覆盖复杂时空关系
- 适合研究机器人感知、人机交互等需要实时动态理解的领域
事件相机以微秒级精度捕捉亮度变化,在运动模糊和极端光照条件下仍保持稳定,为建模高动态场景提供显著优势。然而,其与自然语言理解的融合尚未受到足够关注,导致多模态感知存在空白。为此,我们提出Talk2Event,首个基于事件数据的语言驱动目标定位大规模基准。该数据集基于真实驾驶场景构建,包含5,567个场景、13,458个标注对象和超过30,000条经过严格验证的指代表达。每条表达均附加四个结构化属性:外观、状态、与观察者的关系以及与周围物体的关系,显式捕捉空间、时间与关系线索。这一属性中心设计支持可解释且可组合的定位,使分析超越简单识别,实现动态环境中的上下文推理。我们期望Talk2Event成为推进多模态与时间感知感知的基础,应用于机器人、人机交互等领域。
原文摘要 · Abstract (English)
Event cameras capture changes in brightness with microsecond precision and remain reliable under motion blur and challenging illumination, offering clear advantages for modeling highly dynamic scenes. Yet, their integration with natural language understanding has received little attention, leaving a gap in multimodal perception. To address this, we introduce Talk2Event, the first large-scale benchmark for language-driven object grounding using event data. Built on real-world driving scenarios, Talk2Event comprises 5,567 scenes, 13,458 annotated objects, and more than 30,000 carefully validated referring expressions. Each expression is enriched with four structured attributes -- appearance, status, relation to the viewer, and relation to surrounding objects -- that explicitly capture spatial, temporal, and relational cues. This attribute-centric design supports interpretable and compositional grounding, enabling analysis that moves beyond simple object recognition to contextual reasoning in dynamic environments. We envision Talk2Event as a foundation for advancing multimodal and temporally-aware perception, with applications spanning robotics, human-AI interaction, and so on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。