arXiv:2607.17182cs.CV2026-07

构建首个孟加拉语点击标题数据集,验证多模态融合有效提升识别准确率。

BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos

论文配图:BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos
图 1 · 摘自论文原文
  • 整合图文信息,用中间融合方式结合视觉与语言模型
  • 多模态模型准确率达0.84,优于单一文本或图像模型
  • 适合低资源语言内容分析与虚假信息检测研究者参考

点击标题(clickbait)通过夸大或歪曲视频内容误导用户,损害信任、浪费注意力并助长虚假信息传播。现有公开的孟加拉语多模态数据集稀缺,限制了该领域研究。为此,我们构建了BanClickThumb数据集,包含7,147个来自五个内容领域的孟加拉语视频标题-缩略图对,由十名标注员标注,一致性高(Cohen's Kappa: 0.83–0.93)。基于此数据集,我们评估了纯文本、纯图像及多模态方法。在单模态模型中,BanClickTextFormer(XLM-RoBERTa)达到0.82准确率,BanClickImageFormer(SwiftFormer)达0.68。提出的多模态模型BanClickFusionFormer,通过中间融合结合ViT与XLM-RoBERTa,取得最高准确率0.84。错误分析显示,缩略图密集文字、隐喻语言和文化特有俚语仍是挑战。结果表明,多模态融合对孟加拉语点击标题检测有效,并为低资源多模态内容分析提供了公开基准。

原文摘要 · Abstract (English)

Clickbait, where video titles and thumbnails exaggerate or misrepresent content, reduces user trust, wastes attention, and promotes misinformation on video-sharing platforms. Detecting Bengali clickbait remains challenging because publicly available multimodal datasets are limited. To address this gap, we introduce BanClickThumb, a curated dataset of 7,147 Bengali YouTube thumbnail-title pairs from five content domains, annotated by ten annotators with high agreement (Cohen's Kappa: 0.83-0.93). Using this dataset, we benchmark text-only, image-only, and multimodal approaches. Among unimodal models, BanClickTextFormer (XLM-RoBERTa) achieves 0.82 accuracy, while BanClickImageFormer (SwiftFormer) reaches 0.68. Our proposed multimodal model, BanClickFusionFormer, combines ViT and XLM-RoBERTa through intermediate fusion and achieves the best accuracy of 0.84. Error analysis shows that dense thumbnail text, figurative language, and culturally specific slang remain challenging. Our findings demonstrate the effectiveness of multimodal fusion for Bengali clickbait detection and provide a publicly available benchmark to support future research on low-resource multimodal content analysis.

点击标题检测多模态融合低资源语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。