arXiv:2410.14814cs.CLcs.LG2024-10被引 1

用跨领域迁移和实体信息提升文本欺骗检测准确率

Effects of Soft-Domain Transfer and Named Entity Information on Deception Detection

  • 通过拼接BERT中间层实现多领域数据迁移学习
  • 结合命名实体增强使准确率最高提升11.2%
  • 发现杰恩-申农距离与迁移效果中度相关,适合做数据评估

在线交流日益普遍,辨别文本真实性面临挑战。为提升欺骗检测能力,研究在八个不同领域的文本数据集上,采用微调BERT模型的中间层拼接进行迁移学习,显著优于基线。实验表明,结合命名实体信息可使准确率最高提升11.2%;同时,多种数据集间距离度量中,杰恩-申农距离与迁移性能呈中度正相关,提示其可用于评估跨域适用性。

原文摘要 · Abstract (English)

In the modern age an enormous amount of communication occurs online, and it is difficult to know when something written is genuine or deceitful. There are many reasons for someone to deceive online (e.g., monetary gain, political gain) and detecting this behavior without any physical interaction is a difficult task. Additionally, deception occurs in several text-only domains and it is unclear if these various sources can be leveraged to improve detection. To address this, eight datasets were utilized from various domains to evaluate their effect on classifier performance when combined with transfer learning via intermediate layer concatenation of fine-tuned BERT models. We find improvements in accuracy over the baseline. Furthermore, we evaluate multiple distance measurements between datasets and find that Jensen-Shannon distance correlates moderately with transfer learning performance. Finally, the impact was evaluated of multiple methods, which produce additional information in a dataset's text via named entities, on BERT performance and we find notable improvement in accuracy of up to 11.2%.

欺骗检测迁移学习命名实体BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。