arXiv:2410.24028cs.LGcs.HC2024-10被引 5

让手机多模态数据一到就推理,不等慢的,还能提升准确率。

AdaFlow: Opportunistic Inference on Asynchronous Mobile Data with Generalized Affinity Control

  • 基于层次化矩阵动态控制多模态数据关联性,支持异步输入
  • 推理延迟降低79.9%,准确率最高提升61.9%
  • 无需重训练即可适配不同传感器和任务,适合移动端应用

配备多种传感器(如LiDAR、摄像头)的移动设备推动了分布式多模态智能在智慧座舱、驾驶辅助等场景的应用。然而,由于模态数据大小与网络状态差异,数据到达时间不一致,导致等待慢数据造成延迟,或提前推理引发精度下降。现有方法虽关注模态间一致性与互补性(即模态亲和性),但缺乏在开放世界移动环境中对亲和性的计算控制。为此,本文提出AdaFlow,首次将结构化跨模态亲和性建模为基于分层分析的归一化矩阵,可适应多样的模态类型与数量变化。结合亲和性注意力条件生成对抗网络(ACGAN),实现无需重训练的灵活数据补全,适配多种模态与下游任务。实验表明,该方法可使推理延迟降低高达79.9%,准确率提升达61.9%,显著优于现有方法。

原文摘要 · Abstract (English)

The rise of mobile devices equipped with numerous sensors, such as LiDAR and cameras, has spurred the adoption of multi-modal deep intelligence for distributed sensing tasks, such as smart cabins and driving assistance. However, the arrival times of mobile sensory data vary due to modality size and network dynamics, which can lead to delays (if waiting for slower data) or accuracy decline (if inference proceeds without waiting). Moreover, the diversity and dynamic nature of mobile systems exacerbate this challenge. In response, we present a shift to \textit{opportunistic} inference for asynchronous distributed multi-modal data, enabling inference as soon as partial data arrives. While existing methods focus on optimizing modality consistency and complementarity, known as modal affinity, they lack a \textit{computational} approach to control this affinity in open-world mobile environments. AdaFlow pioneers the formulation of structured cross-modality affinity in mobile contexts using a hierarchical analysis-based normalized matrix. This approach accommodates the diversity and dynamics of modalities, generalizing across different types and numbers of inputs. Employing an affinity attention-based conditional GAN (ACGAN), AdaFlow facilitates flexible data imputation, adapting to various modalities and downstream tasks without retraining. Experiments show that AdaFlow significantly reduces inference latency by up to 79.9\% and enhances accuracy by up to 61.9\%, outperforming status quo approaches.

多模态移动端异步推理生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。