测试17个大模型在家庭场景复合任务中的表现,发现通用智能差距明显。
Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
- 设计基于儿童日常活动的三类复合任务:物体理解、空间推理、社交行为。
- 17个主流多模态大模型在三类任务中均表现不佳,平均准确率不足40%。
- 适合关注具身智能评估与真实世界部署的研究者参考。
区分通用人工智能(AGI)与传统AI的关键特征在于能否完成需多种能力协同的复合任务。尽管由多模态大语言模型(MLLMs)驱动的具身智能体具备丰富的感知与交互能力,但其是否能解决复合任务仍缺乏系统探索。本文基于早期儿童发展观察,设计了一套模拟真实家庭环境的复合任务,涵盖物体理解、空间智能和社交活动三大核心领域。对17个主流专有及开源的MLLMs进行了评估,结果显示所有模型在三个领域均表现较差,表明当前能力与通用智能要求之间存在显著差距。该任务集为评估具身智能体的综合能力提供了初步框架,是迈向具身MLLM发展与实际应用的重要一步。
原文摘要 · Abstract (English)
A key feature differentiating artificial general intelligence (AGI) from traditional AI is that AGI can perform composite tasks that require a wide range of capabilities. Although embodied agents powered by multimodal large language models (MLLMs) offer rich perceptual and interactive capabilities, it remains largely unexplored whether they can solve composite tasks. In the current work, we designed a set of composite tasks inspired by common daily activities observed in early childhood development. Within a dynamic and simulated home environment, these tasks span three core domains: object understanding, spatial intelligence, and social activity. We evaluated 17 leading proprietary and open-source MLLMs on these tasks. The results consistently showed poor performance across all three domains, indicating a substantial gap between current capabilities and general intelligence requirements. Together, our tasks offer a preliminary framework for evaluating the general capabilities of embodied agents, marking an early but significant step toward the development of embodied MLLMs and their real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。