arXiv:2608.27584cs.CVcs.AI2026-08

用概率模型实时处理单光子信号,让机器人在极暗环境中也能稳定感知。

Quanta Perception as Probabilistic Events

论文配图:Quanta Perception as Probabilistic Events
图 1 · 摘自论文原文
  • 将光子流建模为递归贝叶斯信念状态,实现低延迟感知。
  • 在0.05勒克斯下完成跑步者姿态估计,无需重训练视觉模型。
  • 支持每秒超5万帧光子流处理,速度比现有方法快1000倍以上。

自主系统依赖光信息感知,但在极端环境(如夜间导航、高速机器人)中仍易失效。传统传感器固定曝光时间,在灵敏度、动态范围和时间分辨率间存在权衡,导致光子稀少或动态快速时性能下降。量子传感器可探测单个光子,但其数据流远超实时计算能力。本文提出“概率事件”这一计算原语,通过推断自上次亮度变化以来的时间后验分布,将光子流表示为递归信念状态。相比固定阈值事件相机触发机制,该递归贝叶斯框架生成三种低延迟信号:自适应运动的场景通量、高保真活动图与基于熵的感知不确定性。该表征使系统可在极端条件下运行,包括在约0.05勒克斯下对奔跑人体进行姿态估计,且无需重新训练视觉模型。本方法在通用GPU上处理超过50,000量子帧/秒的输入流,输出达千赫兹级,比当前最先进的量子重建基线快达四数量级,即使对兆像素阵列亦然。通过跳过帧重建,直接对光子流进行概率推理,本工作实现了量子计数传感与机器人视觉的融合。

原文摘要 · Abstract (English)

Autonomous systems rely on extracting information from light, yet remain brittle in extreme environments, from nighttime navigation to high-speed robotics. Conventional sensors aggregate photons over fixed exposures, imposing trade-offs between sensitivity, dynamic range, and temporal resolution that degrade perception when photons are scarce or dynamics are rapid. Quanta sensors detect individual photons, but their streams exceed real-time compute and latency budgets by orders of magnitude. Here we introduce $\textit{probabilistic events}$, a computational primitive for real-time quanta perception from individual photon detections. By computing the posterior over the time since the last intensity change, we represent photon streams as recursive belief states. Rather than fixed-threshold event-camera triggers, this recursive Bayesian formulation yields three low-latency signals: motion-adaptive scene flux, high-fidelity activity maps, and entropy-based perceptual uncertainty. This representation enables perception in extreme conditions, including pose estimation of a running person at $\sim$0.05 lux---without retraining vision models. Our approach processes input streams exceeding 50{,}000 quanta frames per second on commodity GPU hardware---yielding kilohertz-scale outputs up to four orders of magnitude faster than state-of-the-art quanta reconstruction baselines, even for megapixel arrays. By replacing frame reconstruction with direct probabilistic inference over photon streams, this work bridges photon-counting quanta sensing with robotic vision.

量子感知概率推理极暗视觉事件相机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。