提出轻量级领域专用AI模型,实现低功耗高效推理。
A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
- 设计10-20B参数的紧凑多模态模型,专注特定领域。
- 目标实现系统级能效提升1000倍以上,支持实时决策。
- 适合对功耗敏感、需持续学习的嵌入式智能场景。
人工智能已深刻影响社会、产业与治理,市场预计从2023年的1890亿美元增长至2033年的4.8万亿美元。当前主流为大型语言模型(LLMs),其训练需大量网络数据及50–60 GWh能源(如GPT-4)。但此类模型常出现幻觉,难以部署于关键领域。相比之下,人脑仅消耗20W功率。未来需要发展轻量级、领域专用的多模态模型,具备10–20B参数规模,在有限领域内实现推理、规划与实时决策能力,并支持持续学习。该愿景要求硬件系统实现≥1000倍于现有水平的能效提升,同时满足精度、延迟与覆盖度约束。本文提出面向未来高能效领域智能体的体系结构蓝图。
原文摘要 · Abstract (English)
The field of artificial intelligence (AI) has taken a tight hold on broad aspects of society, industry, business, and governance in ways that dictate the prosperity and might of the world's economies. The AI market size is projected to grow from {\$}189 billion in 2023 to {\$}4.8 trillion by 2033. Currently, AI is dominated by large language models (LLMs) that exhibit linguistic and visual intelligence. However, training these models requires a massive amount of data scraped from the web as well as large amounts of energy (50-60 GWh to train GPT-4). Despite these costs, these models often hallucinate, a characteristic that prevents them from being deployed in critical application domains. In contrast, the human brain consumes only 20W of power. What is needed is the next level of AI evolution in which lightweight domain-specific multimodal models, especially compact models with 10--20B parameters for bounded domains, with higher levels of intelligence can reason, plan, and make decisions in dynamic environments with real-time data and prior knowledge, while learning continuously and evolving in ways that enhance future decision-making capability. This will define the next wave of AI, progressing from today's large models, trained with vast amounts of data, to nimble energy-efficient domain-specific agents that can reason and think in a world full of uncertainty. To support such agents, hardware will need to be reimagined to allow system-level energy efficiencies $\geq {1000X}$ over the state of the art for targeted domain tasks, subject to accuracy, latency, and coverage constraints. Such a vision of future AI systems is developed in this work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。