arXiv:2511.16225cs.LG2025-11

提出自适应时间窗机制,实时应对多模态数据延迟不确定性

Real-Time Inference for Distributed Multimodal Systems under Communication Delay Uncertainty

  • 用动态时间窗替代固定参考模态,灵活应对不同流的延迟变化
  • 在AVEL任务中实现比现有方法更优的实时推理鲁棒性
  • 无需离线调参,适合网络波动大的分布式多模态系统

联网的网络物理系统基于多个数据流的实时输入执行推理。跨数据流的通信延迟不确定性破坏了推理过程的时间连续性。现有最先进(SotA)的非阻塞推理方法依赖参考模态范式,要求某一模态完全接收后才开始处理,并依赖昂贵的离线性能分析。本文提出一种新型神经启发式非阻塞推理范式,主要采用自适应时间窗集成(TWIs),可动态适应异构数据流中的随机延迟模式,同时放宽对参考模态的依赖。所提出的通信延迟感知框架实现了鲁棒的实时推理,对精度-延迟权衡具有更细粒度的控制能力。在音频-视觉事件定位(AVEL)任务上的实验表明,该方法相比SotA方法对网络动态变化展现出更强的适应性。

原文摘要 · Abstract (English)

Connected cyber-physical systems perform inference based on real-time inputs from multiple data streams. Uncertain communication delays across data streams challenge the temporal flow of the inference process. State-of-the-art (SotA) non-blocking inference methods rely on a reference-modality paradigm, requiring one modality input to be fully received before processing, while depending on costly offline profiling. We propose a novel, neuro-inspired non-blocking inference paradigm that primarily employs adaptive temporal windows of integration (TWIs) to dynamically adjust to stochastic delay patterns across heterogeneous streams while relaxing the reference-modality requirement. Our communication-delay-aware framework achieves robust real-time inference with finer-grained control over the accuracy-latency tradeoff. Experiments on the audio-visual event localization (AVEL) task demonstrate superior adaptability to network dynamics compared to SotA approaches.

多模态推理实时系统延迟鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。