首个尼泊尔语多模态虚假上下文检测基准,发现文本语义已足够准确
NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection

- 构建首个尼泊尔语主导的多模态虚假上下文数据集
- 仅用文本模型即达94.65%准确率,超越多模态系统
- 适合关注多语言虚假信息检测的研究者
虚假上下文(OOC) misinformation通过将真实图像与误导性标题配对制造虚假叙事,无需图像篡改,检测核心在于多模态对齐而非图像取证。尽管此类信息在尼泊尔广泛存在且影响深远,但此前缺乏公开的尼泊尔语基准。本文提出NepOOC,首个公开可用的尼泊尔语主导多模态OOC基准,包含1,090组图像-标题对(545组真实,545组OOC),涵盖五类标注(虚构、误标、时间错位、地理错位、身份错位),标注者间一致性kappa=0.84。对五种多模态架构及纯文本/纯图像基线的系统评估表明,在当前数据规模下,标题语义已足够支撑优异性能。纯文本mBERT模型取得94.65±0.20%宏平均F1,与最佳多模态系统(ResNet-50+mBERT,94.65±0.20%)无统计差异(McNemar中位数p=1.000,5个种子中0个显著,α=0.05)。纯图像模型表现接近随机(33-50%),而训练规模扩展预示数据扩充比架构复杂化或区域特化更有效。
原文摘要 · Abstract (English)
Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite the prevalence and consequences of OOC misinformation in Nepal, no public benchmark exists for Nepali. We introduce NepOOC, the first publicly available Nepali-dominant multilingual OOC benchmark, comprising 1,090 image-caption pairs (545 pristine, 545 OOC) annotated across five typologies (fabricated, miscaptioned, temporal mismatch, geographic mismatch, identity mismatch) with inter-annotator agreement kappa = 0.84. Systematic evaluation of five multimodal architectures alongside text-only and image-only baselines reveals that caption semantics appear sufficient for strong performance at the current dataset scale. A text-only mBERT model achieves 94.65+/-0.20% Macro-F1, statistically equivalent to the best multimodal system (ResNet-50+mBERT, 94.65+/-0.20%; McNemar median p = 1.000, 0/5 seeds significant at alpha = 0.05). Image-only models perform near chance (33-50%), while training-size scaling suggests that dataset expansion is a more direct path to progress than architectural sophistication or regional specialisation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。