arXiv:2511.08085cs.CLcs.AI2025-11

发现孟加拉语停用词是作者识别的关键风格特征,新数据集揭示其重要性。

BARD10: A New Benchmark Reveals Significance of Bangla Stop-Words in Authorship Attribution

  • 构建10位作者的孟加拉语文本数据集BARD10,系统测试停用词移除影响
  • 传统TF-IDF+SVM在两数据集上表现最佳,BERT模型反而落后5个百分点
  • 停用词高频成分承载作者特征,适合短文本与跨领域研究者参考

本研究深入探讨孟加拉语作者识别问题,提出新平衡基准数据集BARD10(包含10位当代作者的博客与评论文本),并系统分析停用词移除对经典与深度学习模型的影响,揭示孟加拉语停用词的风格意义。在统一预处理下,评估了SVM、Bangla BERT、XGBoost和MLP四种分类器在BARD10与基准数据集BAAD16上的表现。所有数据集中,基于TF-IDF+SVM的基线模型最优,在BAAD16上宏平均F1达0.997,在BARD10上为0.921;而Bangla BERT性能低约5个百分点。分析显示,BARD10作者对停用词修剪高度敏感,而BAAD16作者则相对稳健,表明文体依赖性。错误分析指出,高频成分传递作者特征,但被变压器模型弱化。研究总结三点:停用词是关键风格指标;微调机器学习模型在短文本中有效;BARD10连接正式文献与网络对话,为未来长文本或领域适配变压器提供可复现基准。

原文摘要 · Abstract (English)

This research presents a comprehensive investigation into Bangla authorship attribution, introducing a new balanced benchmark corpus BARD10 (Bangla Authorship Recognition Dataset of 10 authors) and systematically analyzing the impact of stop-word removal across classical and deep learning models to uncover the stylistic significance of Bangla stop-words. BARD10 is a curated corpus of Bangla blog and opinion prose from ten contemporary authors, alongside the methodical assessment of four representative classifiers: SVM (Support Vector Machine), Bangla BERT (Bidirectional Encoder Representations from Transformers), XGBoost, and a MLP (Multilayer Perception), utilizing uniform preprocessing on both BARD10 and the benchmark corpora BAAD16 (Bangla Authorship Attribution Dataset of 16 authors). In all datasets, the classical TF-IDF + SVM baseline outperformed, attaining a macro-F1 score of 0.997 on BAAD16 and 0.921 on BARD10, while Bangla BERT lagged by as much as five points. This study reveals that BARD10 authors are highly sensitive to stop-word pruning, while BAAD16 authors remain comparatively robust highlighting genre-dependent reliance on stop-word signatures. Error analysis revealed that high frequency components transmit authorial signatures that are diminished or reduced by transformer models. Three insights are identified: Bangla stop-words serve as essential stylistic indicators; finely calibrated ML models prove effective within short-text limitations; and BARD10 connects formal literature with contemporary web dialogue, offering a reproducible benchmark for future long-context or domain-adapted transformers.

作者识别孟加拉语停用词数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。