让智能体主动学数据,才是AI落地的关键突破。
What's the next frontier for Data-centric AI? Data Savvy Agents
- 设计能自动找数据、补知识的智能体,不再依赖人工预设。
- 提出四能力框架:主动采数、灵活处理、动态测试、持续进化。
- 适合关注真实场景部署的开发者与研究者,推动数据驱动的下一代智能体。
近期自主通信、协作并使用多种工具的AI智能体展现出广阔应用前景,但其对数据的处理能力仍被忽视。实现可扩展自主性需智能体持续获取、处理并演化数据。本文主张将数据敏锐能力作为智能体系统设计的核心,提出四大关键能力:(1) 主动数据获取:自主收集任务关键知识或向人类求助填补数据缺口;(2) 复杂数据处理:具备上下文感知能力,灵活应对多样数据挑战;(3) 交互式测试数据生成:从静态基准转向动态生成的交互式测试数据以评估智能体;(4) 持续适应:通过迭代优化数据与背景知识,适应环境变化。当前智能体研究多聚焦推理,本文呼吁重视数据敏锐智能体作为数据中心型AI的新前沿。
原文摘要 · Abstract (English)
The recent surge in AI agents that autonomously communicate, collaborate with humans and use diverse tools has unlocked promising opportunities in various real-world settings. However, a vital aspect remains underexplored: how agents handle data. Scalable autonomy demands agents that continuously acquire, process, and evolve their data. In this paper, we argue that data-savvy capabilities should be a top priority in the design of agentic systems to ensure reliable real-world deployment. Specifically, we propose four key capabilities to realize this vision: (1) Proactive data acquisition: enabling agents to autonomously gather task-critical knowledge or solicit human input to address data gaps; (2) Sophisticated data processing: requiring context-aware and flexible handling of diverse data challenges and inputs; (3) Interactive test data synthesis: shifting from static benchmarks to dynamically generated interactive test data for agent evaluation; and (4) Continual adaptation: empowering agents to iteratively refine their data and background knowledge to adapt to shifting environments. While current agent research predominantly emphasizes reasoning, we hope to inspire a reflection on the role of data-savvy agents as the next frontier in data-centric AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。