arXiv:2608.14952cs.ROcs.CV2026-08被引 1

用声音感知缺失来推断隐藏危险,让自动驾驶在看不见时仍能预警。

Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails

论文配图:Evidence of Absence: Cross-Modal Abductive Risk Perception to Sustain World Models When Vision Fails
图 1 · 摘自论文原文
  • 通过声音缺失反推视觉盲区中的潜在风险源
  • 提前1.7秒预警,误报率降低42%,定位精度达3.4度
  • 适合高安全需求的自动驾驶场景,尤其视线受阻时

当视觉感知因遮挡或退化失效时,系统难以维持世界模型。本文提出一种基于跨模态反事实推理的框架:将预期共现证据的缺失视为隐藏原因的证据。该方法以声学为补充模态,通过麦克风阵列估计发动机与轮胎源方位,并提取接近速率信息(有稳定音调时用多普勒,否则用宽带逼近信号)。当存在事件特征但无视觉对应证据时,触发对隐藏道路使用者的反演推理,生成校准的风险提示而非控制指令。通过分析隐状态可恢复性,将线索识别建模为在明确虚警预算下的奈曼-皮尔逊检测。在真实盲区路口的遮挡接近数据上,方法平均在视线进入前1.7秒发出警告,与现有声学基线相比检测率相当但误报减少42%,一旦可见后定位中位误差为3.4度,校准误差仅为0.034,在视觉通道退化至0.03时仍保持0.87以上的危险意识水平。此外,校准性能几乎无损迁移至未见路口,而签名分类器则不具泛化能力,移动自车噪声是实际部署的主要限制。

原文摘要 · Abstract (English)

A structured world-state (entities, relations, context, and predictive cues) is designed to preserve prediction-critical content when perception degrades, but it presumes observations to populate it; when the primary visual modality is occluded or degraded, those observations may be missing. We address how to sustain the world model from a complementary modality by treating the absence of expected co-evidence as evidence of a hidden cause. The abductive framework is modality-agnostic; this article instantiates it acoustically. A microphone-array front-end estimates the bearing of engine and tire sources and extracts approach-rate evidence (Doppler when a stable tone exists, a broadband looming readout otherwise); the event "signature present, visual co-evidence absent" then triggers abductive inference of a hidden road user, emitting a calibrated risk advisory rather than a control command. Recoverability of the hidden state is analyzed as an identifiability question separating shared from modality-unique information, and cueing is cast as Neyman-Pearson detection under an explicit false-alarm budget. On real occluded-approach recordings at blind junctions, the method warns a mean 1.7 seconds before line-of-sight entry, matches the sustained-window variant of the published acoustic baseline's detection rate with 42% fewer false alarms, localizes to 3.4 degrees median once in view, is well calibrated (expected calibration error 0.034), and keeps hazard awareness above 0.87 under staged vision degradation that collapses a vision-only channel to 0.03. We also measure the method's limits: calibration transfers to an unseen junction almost losslessly, the signature classifier does not, and moving-ego noise is the binding deployment constraint.

自动驾驶多模态感知风险预警反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。