首个乌尔都语多模态命名实体识别数据集与框架
A Benchmark Dataset and a Framework for Urdu Multimodal Named Entity Recognition
- 构建乌尔都语图文联合识别框架,融合文本与视觉信息
- 在自建数据集上实现当前最优性能,超越纯文本模型
- 为低资源语言多模态研究提供基准与数据支持
多模态内容(尤其是社交媒体上的图文结合内容)的兴起使多模态命名实体识别(MNER)成为自然语言处理的重要方向。尽管英语等高资源语言已取得进展,乌尔都语等低资源语言的MNER仍严重缺乏标注数据和标准基准。为此,本文提出U-MNER框架并发布Twitter2015-Urdu数据集——首个专为乌尔都语设计的多模态命名实体识别数据集,基于广泛使用的Twitter2015数据集并采用乌尔都语语法规则进行标注。通过评估纯文本与多模态模型,建立了基准对比分析,以支持未来研究。该框架利用Urdu-BERT提取文本特征、ResNet提取视觉特征,并通过跨模态融合模块对齐与融合信息。实验表明,所提模型在Twitter2015-Urdu数据集上达到当前最优性能,为低资源语言的MNER研究奠定基础。
原文摘要 · Abstract (English)
The emergence of multimodal content, particularly text and images on social media, has positioned Multimodal Named Entity Recognition (MNER) as an increasingly important area of research within Natural Language Processing. Despite progress in high-resource languages such as English, MNER remains underexplored for low-resource languages like Urdu. The primary challenges include the scarcity of annotated multimodal datasets and the lack of standardized baselines. To address these challenges, we introduce the U-MNER framework and release the Twitter2015-Urdu dataset, a pioneering resource for Urdu MNER. Adapted from the widely used Twitter2015 dataset, it is annotated with Urdu-specific grammar rules. We establish benchmark baselines by evaluating both text-based and multimodal models on this dataset, providing comparative analyses to support future research on Urdu MNER. The U-MNER framework integrates textual and visual context using Urdu-BERT for text embeddings and ResNet for visual feature extraction, with a Cross-Modal Fusion Module to align and fuse information. Our model achieves state-of-the-art performance on the Twitter2015-Urdu dataset, laying the groundwork for further MNER research in low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。