用知识图谱增强语言模型,提升有害文本识别准确率。
Improving Harmful Text Detection with Joint Retrieval and External Knowledge
- 融合预训练模型与知识图谱进行联合检索
- 低资源和多语言场景下性能显著优于单模型
- 适合AI安全、内容审核方向的研究者参考
有害文本检测在大语言模型的开发与部署中变得至关重要,尤其随着AI生成内容在数字平台持续扩展。本文提出一种联合检索框架,将预训练语言模型与知识图谱相结合,以提升有害文本检测的准确性和鲁棒性。实验结果表明,该方法在低资源训练场景和多语言环境中显著优于单一模型基线,通过利用外部上下文信息有效捕捉细微有害内容,弥补了传统检测模型的局限。未来研究应聚焦计算效率优化、模型可解释性提升及多模态检测能力拓展,以应对不断演化的有害内容模式。本工作推动了AI安全发展,助力构建更可信可靠的内容审核系统。
原文摘要 · Abstract (English)
Harmful text detection has become a crucial task in the development and deployment of large language models, especially as AI-generated content continues to expand across digital platforms. This study proposes a joint retrieval framework that integrates pre-trained language models with knowledge graphs to improve the accuracy and robustness of harmful text detection. Experimental results demonstrate that the joint retrieval approach significantly outperforms single-model baselines, particularly in low-resource training scenarios and multilingual environments. The proposed method effectively captures nuanced harmful content by leveraging external contextual information, addressing the limitations of traditional detection models. Future research should focus on optimizing computational efficiency, enhancing model interpretability, and expanding multimodal detection capabilities to better tackle evolving harmful content patterns. This work contributes to the advancement of AI safety, ensuring more trustworthy and reliable content moderation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。