arXiv:2409.09504cs.CL2024-09中稿 · publication in "18…被引 2

首个面向低资源孟加拉语的多模态作者意图分类框架,提升社交媒体内容理解精度。

Uddessho: An Extensive Benchmark Dataset for Multimodal Author Intent Classification in Low-Resource Bangla Language

  • 融合文本与图像的多模态方法,结合早期与晚期融合策略分析作者意图。
  • 多模态模型准确率达76.19%,较单模态提升11.66%。
  • 适用于低资源语言社交内容分析,尤其适合研究多模态意图识别的学者。

随着互联网上日常信息分享与获取的普及,本文提出一种针对孟加拉语社交媒体帖子中意图分类的创新方法,重点关注个体表达观点与思想的语境。现有方法在低资源语言如孟加拉语中面临挑战,尤其是当作者特征与意图紧密关联时。为此,我们提出多模态作者孟加拉语意图分类(MABIC)框架,利用文本与图像数据深入理解内容背后的意图。我们构建了名为「Uddessho」的数据集,包含3,048个来自社交媒体的实例。方法包括两种意图分类路径:文本意图与多模态作者意图分类,采用早期融合与晚期融合技术。实验表明,单模态方法在解读孟加拉语文本意图时准确率为64.53%;而多模态方法显著优于传统单模态方法,准确率达到76.19%,提升11.66%。据我们所知,这是首个针对低资源孟加拉语社交媒体帖子的多模态作者意图分类研究。

原文摘要 · Abstract (English)

With the increasing popularity of daily information sharing and acquisition on the Internet, this paper introduces an innovative approach for intent classification in Bangla language, focusing on social media posts where individuals share their thoughts and opinions. The proposed method leverages multimodal data with particular emphasis on authorship identification, aiming to understand the underlying purpose behind textual content, especially in the context of varied user-generated posts on social media. Current methods often face challenges in low-resource languages like Bangla, particularly when author traits intricately link with intent, as observed in social media posts. To address this, we present the Multimodal-based Author Bangla Intent Classification (MABIC) framework, utilizing text and images to gain deeper insights into the conveyed intentions. We have created a dataset named "Uddessho," comprising 3,048 instances sourced from social media. Our methodology comprises two approaches for classifying textual intent and multimodal author intent, incorporating early fusion and late fusion techniques. In our experiments, the unimodal approach achieved an accuracy of 64.53% in interpreting Bangla textual intent. In contrast, our multimodal approach significantly outperformed traditional unimodal methods, achieving an accuracy of 76.19%. This represents an improvement of 11.66%. To our best knowledge, this is the first research work on multimodal-based author intent classification for low-resource Bangla language social media posts.

多模态语言理解低资源语言意图分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。