首个支持任意事件语义分割的多粒度框架,让语言指令轻松定位事件目标。
Segment Any Events with Language
- 基于视觉提示统一实现事件分割与开放词汇掩码分类
- 在四个新构建基准上显著超越基线,参数高效且推理更快
- 适合需要灵活响应语言指令的智能感知系统开发者
以自由语言形式进行场景理解已在图像、点云和激光雷达等多模态中广泛研究,但针对事件传感器的相关工作仍较少,且多局限于语义层面的理解。本文提出SEAL——首个面向开放词汇事件实例分割(OV-EIS)的语义感知事件分割框架。给定视觉提示,该模型可在实例级和部件级等多个粒度层次上统一支持事件分割与开放词汇掩码分类。为全面评估OV-EIS任务,我们构建了四个涵盖粗到细标签粒度及从实例级到部件级语义粒度的基准数据集。大量实验证明,SEAL在性能与推理速度上均显著优于现有基线,且具有参数高效特性。附录中还提出了一个无需用户输入视觉提示的简化版本,可实现通用时空开放词汇事件分割。
原文摘要 · Abstract (English)
Scene understanding with free-form language has been widely explored within diverse modalities such as images, point clouds, and LiDAR. However, related studies on event sensors are scarce or narrowly centered on semantic-level understanding. We introduce SEAL, the first Semantic-aware Segment Any Events framework that addresses Open-Vocabulary Event Instance Segmentation (OV-EIS). Given the visual prompt, our model presents a unified framework to support both event segmentation and open-vocabulary mask classification at multiple levels of granularity, including instance-level and part-level. To enable thorough evaluation on OV-EIS, we curate four benchmarks that cover label granularity from coarse to fine class configurations and semantic granularity from instance-level to part-level understanding. Extensive experiments show that our SEAL largely outperforms proposed baselines in terms of performance and inference speed with a parameter-efficient architecture. In the Appendix, we further present a simple variant of our SEAL achieving generic spatiotemporal OV-EIS that does not require any visual prompts from users in the inference. Check out our project page in https://0nandon.github.io/SEAL
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。