用模糊理论评估语言模型对线上诱拐风险的识别能力
Evaluating Language Models on Grooming Risk Estimation Using Fuzzy Theory
- 引入模糊理论构建多级诱拐风险评分体系
- 模型在高风险间接言语场景中预测偏差大
- 适合研究网络欺凌检测与安全模型改进者
语言模型在高风险领域面临隐含语义识别难题,尤其在在线儿童诱拐检测中,施害者常通过显性和隐性语言传递危害意图。尽管SBERT等基于Transformer的模型在预判诱拐行为上展现潜力,但其依赖表面特征,且常以执法或警戒对话作为代理数据,未能真实反映受害者的实际交流过程。本文首次系统检验该方法的合理性,评估SBERT在不同群体间对诱拐风险等级的区分能力。结果表明,微调虽有助于模型学习风险评分,但在高风险情境下仍存在显著预测波动,尤其在使用间接言语路径、缺乏直接性暗示的对话中。这揭示了语言模型亟需增强对间接言语行为的建模能力,尤其是在应对施害者策略时。
原文摘要 · Abstract (English)
Encoding implicit language presents a challenge for language models, especially in high-risk domains where maintaining high precision is important. Automated detection of online child grooming is one such critical domain, where predators manipulate victims using a combination of explicit and implicit language to convey harmful intentions. While recent studies have shown the potential of Transformer language models like SBERT for preemptive grooming detection, they primarily depend on surface-level features and approximate real victim grooming processes using vigilante and law enforcement conversations. The question of whether these features and approximations are reasonable has not been addressed thus far. In this paper, we address this gap and study whether SBERT can effectively discern varying degrees of grooming risk inherent in conversations, and evaluate its results across different participant groups. Our analysis reveals that while fine-tuning aids language models in learning to assign grooming scores, they show high variance in predictions, especially for contexts containing higher degrees of grooming risk. These errors appear in cases that 1) utilize indirect speech pathways to manipulate victims and 2) lack sexually explicit content. This finding underscores the necessity for robust modeling of indirect speech acts by language models, particularly those employed by predators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。