arXiv:2509.17567cs.AI2025-09被引 18

用78个高质量示范样本,让AI实现高效自主执行任务

LIMI: Less is More for Agency

  • 仅用78个精心设计的示范样本训练,突破传统数据规模依赖
  • 在综合评测中达73.5%,显著超越多个大模型,提升超50%
  • 证明高质量示范比海量数据更关键,适合追求效率的智能体开发

我们定义智能体能力为人工智能系统在环境与工具中自主发现问题、提出假设并执行解决方案的涌现能力。当前产业迫切需要的不是仅能思考的AI,而是能主动工作、操作工具并产生真实成果的自主智能体。尽管现有方法普遍认为更多数据带来更强智能,我们提出反直觉的范式:通过聚焦软件开发与科研工作流,仅用78个精心设计的示范样本,即可让智能体在综合性评估中达到73.5%得分,远超Kimi-K2-Instruct(24.1%)、DeepSeek-V3.1(11.9%)、Qwen3-235B-A22B-Instruct(27.5%)和GLM-4.5(45.1%)。最显著的是,其性能比使用10,000样本训练的模型高出53.7%,仅以1/128的数据量实现更优表现。该研究确立了‘智能体效率原则’:机器自主性并非来自数据量,而源于对高质智能行为示范的策略性筛选。

原文摘要 · Abstract (English)

We define Agency as the emergent capacity of AI systems to function as autonomous agents actively discovering problems, formulating hypotheses, and executing solutions through self-directed engagement with environments and tools. This fundamental capability marks the dawn of the Age of AI Agency, driven by a critical industry shift: the urgent need for AI systems that don't just think, but work. While current AI excels at reasoning and generating responses, industries demand autonomous agents that can execute tasks, operate tools, and drive real-world outcomes. As agentic intelligence becomes the defining characteristic separating cognitive systems from productive workers, efficiently cultivating machine autonomy becomes paramount. Current approaches assume that more data yields better agency, following traditional scaling laws from language modeling. We fundamentally challenge this paradigm. LIMI (Less Is More for Intelligent Agency) demonstrates that agency follows radically different development principles. Through strategic focus on collaborative software development and scientific research workflows, we show that sophisticated agentic intelligence can emerge from minimal but strategically curated demonstrations of autonomous behavior. Using only 78 carefully designed training samples, LIMI achieves 73.5% on comprehensive agency benchmarks, dramatically outperforming state-of-the-art models: Kimi-K2-Instruct (24.1%), DeepSeek-V3.1 (11.9%), Qwen3-235B-A22B-Instruct (27.5%), and GLM-4.5 (45.1%). Most strikingly, LIMI demonstrates 53.7% improvement over models trained on 10,000 samples-achieving superior agentic intelligence with 128 times fewer samples. Our findings establish the Agency Efficiency Principle: machine autonomy emerges not from data abundance but from strategic curation of high-quality agentic demonstrations.

智能体少样本学习自主执行高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。