arXiv:2409.00940cs.CLcs.AI2024-09被引 8

用大模型自动识别敏感内容,准确率超99%。

Large Language Models for Automatic Detection of Sensitive Topics

  • 用五种大模型检测心理健康类敏感信息
  • 最佳模型GPT-4o准确率达99.5%,F1为0.99
  • 适合内容审核系统开发者参考

敏感信息检测对于维护安全的在线社区至关重要。传统上依赖人工审核,任务繁重且耗时。大型语言模型(LLMs)具备理解自然语言的能力,可能成为辅助审核的有效工具。本研究评估了五种LLMs在两个在线数据集上检测心理健康领域敏感信息的表现,涵盖准确率、精确率、召回率、F1分数和一致性。结果显示,LLMs具备集成到审核流程中的潜力,可作为便捷高效的检测工具。表现最佳的GPT-4o模型平均准确率为99.5%,F1得分为0.99。文章讨论了在审核流程中使用LLMs的优势与潜在挑战,并建议未来研究应关注该技术应用的伦理问题。

原文摘要 · Abstract (English)

Sensitive information detection is crucial in content moderation to maintain safe online communities. Assisting in this traditionally manual process could relieve human moderators from overwhelming and tedious tasks, allowing them to focus solely on flagged content that may pose potential risks. Rapidly advancing large language models (LLMs) are known for their capability to understand and process natural language and so present a potential solution to support this process. This study explores the capabilities of five LLMs for detecting sensitive messages in the mental well-being domain within two online datasets and assesses their performance in terms of accuracy, precision, recall, F1 scores, and consistency. Our findings indicate that LLMs have the potential to be integrated into the moderation workflow as a convenient and precise detection tool. The best-performing model, GPT-4o, achieved an average accuracy of 99.5\% and an F1-score of 0.99. We discuss the advantages and potential challenges of using LLMs in the moderation workflow and suggest that future research should address the ethical considerations of utilising this technology.

大模型内容审核敏感检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。