让大模型在边缘设备上高效推理,兼顾性能与隐私。
Cognitive Edge Computing: A Comprehensive Survey on Optimizing Large Models and AI Agents for Pervasive Deployment
- 构建端侧推理框架,融合量化、剪枝、蒸馏等压缩技术。
- 实现低延迟、低功耗下多步推理能力,支持跨设备协同。
- 适合边缘AI部署、智能终端开发及隐私敏感场景应用。
本文系统综述认知边缘计算,为在资源受限的网络边缘部署具备推理能力的大语言模型(LLMs)和自主AI代理提供实用路径。提出统一的认知保持框架,涵盖:(1) 模型优化(量化、稀疏性、低秩适配、蒸馏),在严苛内存与算力约束下维持多步推理能力;(2) 系统架构(本地推理、弹性卸载、云边协作),权衡延迟、能耗、隐私与容量;(3) 自适应智能(上下文压缩、动态路由、联邦个性化),根据任务难度与设备条件动态调节计算。整合高效Transformer设计、多模态融合、硬件感知编译、隐私学习与代理工具使用等进展,并映射至边缘运行边界。提出标准化评估协议,涵盖延迟、吞吐、每令牌能耗、准确率、鲁棒性、隐私与可持续性,明确测量假设以增强可比性。现存挑战包括模态感知推理基准、透明可复现的能耗报告、面向边缘的安全对齐评估及多智能体测试平台。最后给出算法、运行时与硬件协同设计的实践指南,以实现可靠、高效、私密的边缘认知能力。
原文摘要 · Abstract (English)
This article surveys Cognitive Edge Computing as a practical and methodical pathway for deploying reasoning-capable Large Language Models (LLMs) and autonomous AI agents on resource-constrained devices at the network edge. We present a unified, cognition-preserving framework spanning: (1) model optimization (quantization, sparsity, low-rank adaptation, distillation) aimed at retaining multi-step reasoning under tight memory/compute budgets; (2) system architecture (on-device inference, elastic offloading, cloud-edge collaboration) that trades off latency, energy, privacy, and capacity; and (3) adaptive intelligence (context compression, dynamic routing, federated personalization) that tailors computation to task difficulty and device constraints. We synthesize advances in efficient Transformer design, multimodal integration, hardware-aware compilation, privacy-preserving learning, and agentic tool use, and map them to edge-specific operating envelopes. We further outline a standardized evaluation protocol covering latency, throughput, energy per token, accuracy, robustness, privacy, and sustainability, with explicit measurement assumptions to enhance comparability. Remaining challenges include modality-aware reasoning benchmarks, transparent and reproducible energy reporting, edge-oriented safety/alignment evaluation, and multi-agent testbeds. We conclude with practitioner guidelines for cross-layer co-design of algorithms, runtime, and hardware to deliver reliable, efficient, and privacy-preserving cognitive capabilities on edge devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。