arXiv:2608.18184cs.CV2026-08综述

构建六层次人类中心智能框架,打通多模态研究断点

Human-Centric Intelligence in the Era of Foundation Models: A Survey

论文配图:Human-Centric Intelligence in the Era of Foundation Models: A Survey
图 1 · 摘自论文原文
  • 从视觉、动态、情境三维度构建人类上下文分类体系
  • 系统梳理数据、架构与训练优化方法,整合跨任务资源
  • 面向可扩展、可信、物理合理的人类智能提供实践指南

在基础模型时代,以人为中心的智能正朝着规模化、可迁移性和通用建模方向演进,但尚未充分融入基础模型以实现与之相当的进展。更重要的是,该领域的最新成果分散于不同任务、模态和研究社群中,其内在概念与方法论关联尚不清晰。为弥合这些分歧并重思基础模型时代下的人类中心智能,本文提出一个涵盖六个相互关联层级的全谱人类上下文分类体系:将人作为通过视觉外观与空间几何可观测的主体,作为通过运动学动力学与交互建模的动态行为者,以及作为通过世界模拟与具身代理所处的情境代理。随后,系统阐述领域的方法论基础,涵盖人类中心数据族、计算架构范式及代表性训练与推理优化策略。进一步对各层级代表性方法进行系统综述,并整理相关数据集、基准测试与评估指标。最后讨论开放挑战与有前景的研究方向,推动可扩展、可信、物理合理且可部署的人类中心智能发展,旨在为该领域提供连贯框架与实用参考。项目页面持续更新人类中心人工智能文献与资源集合。

原文摘要 · Abstract (English)

Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, and general-purpose modeling. Yet it has not fully integrated with foundation models to achieve the comparable progress seen in them. More importantly, recent advances across this broad landscape remain fragmented across tasks, modalities, and research communities, leaving their intrinsic conceptual and methodological connections unclear. To bridge these divides and rethink human-centric intelligence in the foundation-model era, we introduce a full-spectrum human context taxonomy that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency. We next present the methodological foundations of the field, covering human-centric data families, computational architecture paradigms, and representative training and inference optimization strategies. We then systematically review representative methods across these levels and organize the associated datasets, benchmarks, and evaluation metrics. We further discuss open challenges and promising research directions toward human-centric intelligence that is scalable, trustworthy, physically grounded, and deployable, aiming to provide a coherent framework and practical reference for advancing the field. Finally, we provide a systematically organized and continuously updated collection of human-centric AI literature and resources on our project page.

人类中心智能基础模型多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。