arXiv:2601.04563cs.LGcs.AI2026-01被引 1

构建融合五感的AI,让机器更懂人类体验

A Vision for Multisensory Intelligence: Sensing, Science, and Synergy

  • 拓展AI感知维度,融入触觉、生理信号等多模态输入
  • 提出统一建模框架,实现跨模态信息融合与迁移
  • 适合人机交互、智能健康、环境感知领域研究者

人类对世界的体验是多感官融合的,涵盖语言、视觉、听觉、触觉、味觉和嗅觉。然而当前人工智能主要聚焦于文本、视觉和音频等数字模态。本文展望未来十年多感官人工智能的发展愿景,旨在通过连接AI与人体感官及环境中丰富的生理、触觉、物理和社会信号,重塑人与AI的交互方式。该领域需围绕感知、科学与协同三大核心推进:首先,拓展AI获取世界信息的途径,突破数字媒介限制;其次,建立量化多模态异质性与交互关系的理论基础,发展统一的建模架构与表征体系,探索跨模态迁移机制;最后,攻克多模态融合、对齐、推理、生成、泛化及用户体验等关键技术挑战。本文附带麻省理工媒体实验室多感官智能团队的系列项目、资源与演示,详见 https://mit-mi.github.io/。

原文摘要 · Abstract (English)

Our experience of the world is multisensory, spanning a synthesis of language, sight, sound, touch, taste, and smell. Yet, artificial intelligence has primarily advanced in digital modalities like text, vision, and audio. This paper outlines a research vision for multisensory artificial intelligence over the next decade. This new set of technologies can change how humans and AI experience and interact with one another, by connecting AI to the human senses and a rich spectrum of signals from physiological and tactile cues on the body, to physical and social signals in homes, cities, and the environment. We outline how this field must advance through three interrelated themes of sensing, science, and synergy. Firstly, research in sensing should extend how AI captures the world in richer ways beyond the digital medium. Secondly, developing a principled science for quantifying multimodal heterogeneity and interactions, developing unified modeling architectures and representations, and understanding cross-modal transfer. Finally, we present new technical challenges to learn synergy between modalities and between humans and AI, covering multisensory integration, alignment, reasoning, generation, generalization, and experience. Accompanying this vision paper are a series of projects, resources, and demos of latest advances from the Multisensory Intelligence group at the MIT Media Lab, see https://mit-mi.github.io/.

多感官智能人机交互感知融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。