arXiv:2605.10384cs.AIcs.DC2026-05中稿 · AutoEdge workshop,…

小模型也能当智能代理,关键在选对模型和工具组合。

Agentic Performance at the Edge: Insights from Benchmarking

  • 用领域适配的评测方法,分析小模型在边缘设备上的表现。
  • 80亿参数以下模型仍可实现高精度任务执行,但效果依赖工具协同。
  • 适合边缘AI部署、需要优化模型与工具链配合的研发人员。

智能代理(Agentic AI)天然适用于物联网和边缘计算系统,但边缘部署通常受限于约80亿参数以内的模型规模。本文首次通过实证研究,考察了边缘场景下的模型缩放规律、通用模型与编程专用模型的差异,以及在固定协议下工具赋能的执行效果。提出一种领域条件化的评估方法,分析模型与工具的协同机制,给出受限条件下的模型选型建议,并揭示不同模型家族在语义理解与执行层面的失败模式差异。核心发现是:边缘智能体性能并非仅由参数量决定,其鲁棒性取决于模型选择与工具流程的联合设计。领域条件分析识别出准确率-延迟空间中的帕累托前沿,可依据实际运行优先级指导策略选择。

原文摘要 · Abstract (English)

Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to models around 8 billion parameters or smaller. An important question is: How much agentic-task quality is lost when model size is constrained by memory, power, and latency budgets? To address this question, in this paper, we provide an initial empirical study considering edge-focused model scaling, general-purpose versus coder-oriented model effects, and tool-enabled execution under a fixed protocol. We introduce a domain-conditioned evaluation methodology, an implementation-grounded analysis of model-tool interactions, practical guidance for model selection under constraints, and an analysis of failure modes that reveals distinct semantic versus execution failure patterns across model families. Our core finding is that edge-agent quality is not a simple function of parameter count. Robust deployment depends on the joint design of model choice and tool workflow. Domain-conditioned analysis reveals Pareto fronts in the accuracy-latency space that can guide strategy selection based on operational priorities.

边缘AI智能代理模型选型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。