arXiv:2602.16653cs.AI2026-02被引 2

小模型在工业场景下用技能框架表现有限,30B以上模型效果更优。

Agent Skill Framework: Perspectives on the Potential of Small to Medium Language Models in Industrial Environments

  • 测试270M至80B开源模型在工业任务中的技能表现
  • 30B-80B模型显著提升技能选择可靠性,小模型表现差
  • 提醒部署时权衡算力成本与智能性能

代理技能被主流智能体框架广泛支持,在专有模型上表现良好,但对中小型开源语言模型(270M-80B)的有效性仍缺乏研究。本文系统考察资源受限工业环境下的技能范式,因数据安全和预算限制,依赖专有API不现实。在两个开源任务和一个真实保险理赔分类任务中发现,极小模型难以可靠选择技能,而30B-80B模型则显著受益。思考变体虽提升性能,但导致GPU使用增加,引发过思考问题。结果揭示算力成本与代理性能间的权衡,为实际部署中小模型提供可操作建议。

原文摘要 · Abstract (English)

Agent skills are widely supported by major agentic frameworks and perform well with proprietary models, yet their effectiveness for small and medium-sized open source language models (270 M-80B) remains underexplored. We systematically study the Skill paradigm in resource-constrained industrial settings, where reliance on proprietary APIs is impractical due to data security and budget constraints. Across two open-source tasks and a real-world insurance claims classification task, we find that very small models struggle with reliable skill selection, while models around 30B-80B benefit substantially. Thinking variants do not show major levels of improvement from skills, also considering GPU usage increases due to overthinking. These findings reveal a trade-off between GPU cost and agent performance, and provide actionable insights for effective Skill configuration and SLM deployment in real world settings.

智能体小模型工业应用技能框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。