用事件相机的边缘信号统一多模态数据,提升运动估计精度。
$x^2$-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Space
- 以事件相机的时空边缘构建统一表征空间,实现跨模态对齐。
- 在真实和合成数据上均达到领先性能,极端条件下优势更明显。
- 适合做多传感器动态场景理解的研究者和工程师参考。
稠密2D光流和3D场景流估计对动态场景理解至关重要。现有方法融合图像、激光雷达与事件数据预测运动,但大多在异构特征空间中操作。缺乏所有模态可对齐的共享潜在空间,导致依赖多个模态专用模块,交叉传感器不匹配问题未解,融合过程复杂化。事件相机天然提供时空边缘信号,可视为内在边缘场,用于锚定统一潜在表示,称为事件边缘空间(Event Edge Space)。基于此,我们提出$x^2$-Fusion,将多模态融合重构为表示统一:事件生成的时空边缘定义以边缘为中心的同质空间,图像与激光雷达特征在此共享空间中显式对齐。在此空间内,采用可靠性感知自适应融合,评估模态可靠性并强化退化条件下的稳定线索。进一步引入跨维度对比学习,紧密耦合2D光流与3D场景流。在合成与真实基准上的大量实验表明,$x^2$-Fusion在标准条件下达到最先进精度,并在挑战性场景中实现显著提升。
原文摘要 · Abstract (English)
Estimating dense 2D optical flow and 3D scene flow is essential for dynamic scene understanding. Recent work combines images, LiDAR, and event data to jointly predict 2D and 3D motion, yet most approaches operate in separate heterogeneous feature spaces. Without a shared latent space that all modalities can align to, these systems rely on multiple modality-specific blocks, leaving cross-sensor mismatches unresolved and making fusion unnecessarily complex.Event cameras naturally provide a spatiotemporal edge signal, which we can treat as an intrinsic edge field to anchor a unified latent representation, termed the Event Edge Space. Building on this idea, we introduce $x^2$-Fusion, which reframes multimodal fusion as representation unification: event-derived spatiotemporal edges define an edge-centric homogeneous space, and image and LiDAR features are explicitly aligned in this shared representation.Within this space, we perform reliability-aware adaptive fusion to estimate modality reliability and emphasize stable cues under degradation. We further employ cross-dimension contrast learning to tightly couple 2D optical flow with 3D scene flow. Extensive experiments on both synthetic and real benchmarks show that $x^2$-Fusion achieves state-of-the-art accuracy under standard conditions and delivers substantial improvements in challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。