arXiv:2511.23287cs.LGcs.CL2025-11中稿 · the 28th Internati…

用视觉和文本融合提升低资源孟加拉语意图识别准确率

Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla

  • 中间层融合多模态特征,平衡文本与图像表征
  • 达到84.11%宏F1,比之前方法提升8.4个百分点
  • 为低资源语言多模态研究提供新基准和框架

互联网和社交网络的兴起导致用户生成内容激增。作者意图理解在解析社交媒体内容中至关重要。本文针对孟加拉语社交媒体帖子中的作者意图分类问题,结合文本与视觉数据进行研究。针对以往单模态方法的局限性,系统评估了基于Transformer的语言模型(mBERT、DistilBERT、XLM-RoBERTa)与视觉架构(ViT、Swin、SwiftFormer、ResNet、DenseNet、MobileNet),使用包含3,048条帖子的Uddessho数据集,涵盖六种实际意图类别。提出一种新型中间融合策略,显著优于早期和晚期融合。实验表明,中间融合(尤其mBERT与Swin Transformer结合)达到84.11%宏F1,创下新纪录,较此前孟加拉语多模态方法提升8.4个百分点。分析显示,引入视觉上下文显著增强意图分类效果。跨模态特征在中间层级融合,实现模态特异性表征与跨模态学习的最佳平衡。本研究为孟加拉语及其他低资源语言建立新基准与方法标准。我们提出的框架命名为BangACMM(Bangla Author Content MultiModal)。

原文摘要 · Abstract (English)

The expansion of the Internet and social networks has led to an explosion of user-generated content. Author intent understanding plays a crucial role in interpreting social media content. This paper addresses author intent classification in Bangla social media posts by leveraging both textual and visual data. Recognizing limitations in previous unimodal approaches, we systematically benchmark transformer-based language models (mBERT, DistilBERT, XLM-RoBERTa) and vision architectures (ViT, Swin, SwiftFormer, ResNet, DenseNet, MobileNet), utilizing the Uddessho dataset of 3,048 posts spanning six practical intent categories. We introduce a novel intermediate fusion strategy that significantly outperforms early and late fusion on this task. Experimental results show that intermediate fusion, particularly with mBERT and Swin Transformer, achieves 84.11% macro-F1 score, establishing a new state-of-the-art with an 8.4 percentage-point improvement over prior Bangla multimodal approaches. Our analysis demonstrates that integrating visual context substantially enhances intent classification. Cross-modal feature integration at intermediate levels provides optimal balance between modality-specific representation and cross-modal learning. This research establishes new benchmarks and methodological standards for Bangla and other low-resource languages. We call our proposed framework BangACMM (Bangla Author Content MultiModal).

多模态低资源语言意图识别Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。