arXiv:2601.15676cs.SDcs.LG2026-01

让边缘音频系统学会分阶段推理,关键时刻才调用云端工具。

Bridging the Perception Gap: A Lightweight Coarse-to-Fine Architecture for Edge Audio Systems

  • 先本地快速处理,发现不确定时才触发云端辅助
  • 在MMAR上准确率从27.2%提升至53.6%
  • 适合对延迟和隐私敏感的边缘音频应用

将音频语言模型(Audio-LLMs)部署在边缘设备时,感知深度与计算效率之间存在持续矛盾。轻量本地模型常产生泛化摘要,无法捕捉多步音频推理所需的细微证据;而盲目云迁移则带来不可接受的延迟、带宽开销和隐私风险。我们提出CoFi-Agent(工具增强的粗粒度到细粒度代理),一种面向边缘服务器与网关的混合架构。它在本地运行单次通过的7B Audio-LLM进行快速感知,当检测到不确定性时,由云端控制器介入并下发轻量级任务指令,如时间重听或本地ASR。在MMAR基准测试中,CoFi-Agent将准确率从27.20%提升至53.60%,且在精度-效率权衡上优于始终在线的调查流水线。整体而言,该架构在实际系统约束下,通过工具赋能的条件性边缘-云协作,弥合了感知差距。

原文摘要 · Abstract (English)

Deploying Audio-Language Models (Audio-LLMs) on edge infrastructure exposes a persistent tension between perception depth and computational efficiency. Lightweight local models tend to produce passive perception - generic summaries that miss the subtle evidence required for multi-step audio reasoning - while indiscriminate cloud offloading incurs unacceptable latency, bandwidth cost, and privacy risk. We propose CoFi-Agent (Tool-Augmented Coarse-to-Fine Agent), a hybrid architecture targeting edge servers and gateways. It performs fast local perception and triggers conditional forensic refinement only when uncertainty is detected. CoFi-Agent runs an initial single-pass on a local 7B Audio-LLM, then a cloud controller gates difficult cases and issues lightweight plans for on-device tools such as temporal re-listening and local ASR. On the MMAR benchmark, CoFi-Agent improves accuracy from 27.20% to 53.60%, while achieving a better accuracy-efficiency trade-off than an always-on investigation pipeline. Overall, CoFi-Agent bridges the perception gap via tool-enabled, conditional edge-cloud collaboration under practical system constraints.

边缘计算音频理解混合架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。