arXiv:2508.12227cs.CL2025-08被引 3

系统梳理阿拉伯语多模态机器学习的研究现状与挑战

Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges

  • 构建四维分类体系:数据集、应用、方法与挑战
  • 归纳现有研究,揭示未被充分探索的空白领域
  • 为后续研究提供结构化方向,适合该领域学者参考

多模态机器学习(MML)旨在融合文本、音频、视觉等多源信息,实现情感分析、情绪识别和多媒体检索等复杂任务。近年来,阿拉伯语多模态机器学习在基础建设上已初具规模,是时候开展全面综述。本文提出一种新型分类框架,将研究工作划分为四个核心维度:数据集、应用场景、方法技术与现存挑战。通过系统梳理现有成果,本综述揭示了尚未深入探索的研究空白,并指明关键瓶颈。研究者可据此把握发展方向,推动该领域持续进步。

原文摘要 · Abstract (English)

Multimodal Machine Learning (MML) aims to integrate and analyze information from diverse modalities, such as text, audio, and visuals, enabling machines to address complex tasks like sentiment analysis, emotion recognition, and multimedia retrieval. Recently, Arabic MML has reached a certain level of maturity in its foundational development, making it time to conduct a comprehensive survey. This paper explores Arabic MML by categorizing efforts through a novel taxonomy and analyzing existing research. Our taxonomy organizes these efforts into four key topics: datasets, applications, approaches, and challenges. By providing a structured overview, this survey offers insights into the current state of Arabic MML, highlighting areas that have not been investigated and critical research gaps. Researchers will be empowered to build upon the identified opportunities and address challenges to advance the field.

多模态阿拉伯语综述数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。