用日文评论分析作者风格,辅助威胁情报中的角色识别。
Foundational Study on Authorship Attribution of Japanese Web Reviews for Actor Analysis
- 基于风格特征,比较四种文本分析方法。
- 大规模作者时TF-IDF+LR更稳定准确,计算成本低。
- 适合做网络威胁角色分析的初学者或小规模研究者。
本研究探讨基于风格特征的作者归属技术在威胁情报中支持角色分析的适用性。作为未来应用于暗网论坛的基础步骤,我们使用来自公开网页的日本购物评论数据进行实验。构建了来自乐天市场(Rakuten Ichiba)的评论数据集,对比了四种方法:TF-IDF+逻辑回归(TF-IDF+LR)、BERT嵌入+逻辑回归(BERT-Emb+LR)、BERT微调(BERT-FT)和度量学习+k近邻(Metric+kNN)。结果表明,BERT-FT表现最佳;但当作者数量达到数百人时,训练变得不稳定,此时TF-IDF+LR在准确率、稳定性及计算成本上均更优。此外,Top-k评估显示候选筛选具实用性,错误分析揭示模板化文本、主题依赖性和短文本长度是导致误分类的主要因素。
原文摘要 · Abstract (English)
This study investigates the applicability of authorship attribution based on stylistic features to support actor analysis in threat intelligence. As a foundational step toward future application to dark web forums, we conducted experiments using Japanese review data from clear web sources. We constructed datasets from Rakuten Ichiba reviews and compared four methods: TF-IDF with logistic regression (TF-IDF+LR), BERT embeddings with logistic regression (BERT-Emb+LR), BERT fine-tuning (BERT-FT), and metric learning with $k$-nearest neighbors (Metric+kNN). Results showed that BERT-FT achieved the best performance; however, training became unstable as the number of authors scaled to several hundred, where TF-IDF+LR proved superior in terms of accuracy, stability, and computational cost. Furthermore, Top-$k$ evaluation demonstrated the utility of candidate screening, and error analysis revealed that boilerplate text, topic dependency, and short text length were primary factors causing misclassification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。