arXiv:2409.06072cs.CLcs.LG2024-09被引 13

构建多任务反欺诈检测基准,评估大模型在真实场景中的表现

DetoxBench: Benchmarking Large Language Models for Multitask Fraud & Abuse Detection

  • 设计涵盖多种滥用语言类型的综合评测基准
  • 大模型在单项任务表现良好,但跨任务差异显著
  • 尤其在识别女性歧视语言等复杂语境时能力不足

大语言模型(LLMs)在自然语言处理任务中展现出卓越能力,但在高风险领域如欺诈与滥用检测中的实际应用仍需深入探索。现有研究多聚焦于单一任务,如毒性或仇恨言论检测。本文提出一个全面的基准测试套件,用于评估 LLMs 在多种真实场景中识别和缓解欺诈性及攻击性语言的能力。该基准涵盖垃圾邮件、仇恨言论、性别歧视语言等多种任务。我们评估了 Anthropic、Mistral AI 及 AI21 家族等多个前沿 LLM 模型,结果显示:尽管模型在单个任务上表现良好,但跨任务性能差异明显,尤其在需要细致语用推理的任务(如识别多样化形式的女性歧视语言)中表现较弱。这些发现对 LLM 在高风险场景中的负责任开发与部署具有重要意义。本基准可为研究人员和实践者提供系统评估工具,推动更稳健、可信且符合伦理的反欺诈系统发展。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing tasks. However, their practical application in high-stake domains, such as fraud and abuse detection, remains an area that requires further exploration. The existing applications often narrowly focus on specific tasks like toxicity or hate speech detection. In this paper, we present a comprehensive benchmark suite designed to assess the performance of LLMs in identifying and mitigating fraudulent and abusive language across various real-world scenarios. Our benchmark encompasses a diverse set of tasks, including detecting spam emails, hate speech, misogynistic language, and more. We evaluated several state-of-the-art LLMs, including models from Anthropic, Mistral AI, and the AI21 family, to provide a comprehensive assessment of their capabilities in this critical domain. The results indicate that while LLMs exhibit proficient baseline performance in individual fraud and abuse detection tasks, their performance varies considerably across tasks, particularly struggling with tasks that demand nuanced pragmatic reasoning, such as identifying diverse forms of misogynistic language. These findings have important implications for the responsible development and deployment of LLMs in high-risk applications. Our benchmark suite can serve as a tool for researchers and practitioners to systematically evaluate LLMs for multi-task fraud detection and drive the creation of more robust, trustworthy, and ethically-aligned systems for fraud and abuse detection.

大模型评测反欺诈滥用检测伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。