arXiv:2504.03603cs.AIcs.LG2025-04被引 11

推动多模态AI在医疗、交通等领域的落地应用

Towards deployment-centric multimodal AI beyond vision and language

  • 从部署角度设计多模态AI,提前考虑实际应用约束
  • 提出跨学科协作框架,解决疫情、自动驾驶等真实场景问题
  • 强调开放合作,助力多模态AI走出视觉语言局限

多模态人工智能通过融合多种数据提升理解与决策能力,在医疗、科学和工程等领域具有广泛应用。然而当前研究主要聚焦于视觉与语言模态,部署可行性仍是关键挑战。本文倡导以部署为中心的工作流程,将实际部署约束纳入早期设计,弥补数据驱动与模型驱动的不足。同时强调多层级多模态融合及跨学科协作,拓展研究边界至非视觉语言领域。为推动该范式,我们识别了跨领域的共性挑战,并分析了疫情应对、自动驾驶设计、气候适应三个实际案例,整合医疗、社会科学、工程、可持续发展、金融等多领域知识。通过促进跨学科对话与开放研究,有望加速多模态AI的落地进程,实现广泛社会影响。

原文摘要 · Abstract (English)

Multimodal artificial intelligence (AI) integrates diverse types of data via machine learning to improve understanding, prediction, and decision-making across disciplines such as healthcare, science, and engineering. However, most multimodal AI advances focus on models for vision and language data, while their deployability remains a key challenge. We advocate a deployment-centric workflow that incorporates deployment constraints early to reduce the likelihood of undeployable solutions, complementing data-centric and model-centric approaches. We also emphasise deeper integration across multiple levels of multimodality and multidisciplinary collaboration to significantly broaden the research scope beyond vision and language. To facilitate this approach, we identify common multimodal-AI-specific challenges shared across disciplines and examine three real-world use cases: pandemic response, self-driving car design, and climate change adaptation, drawing expertise from healthcare, social science, engineering, science, sustainability, and finance. By fostering multidisciplinary dialogue and open research practices, our community can accelerate deployment-centric development for broad societal impact.

多模态AI部署导向跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。