用大模型自动提取隐蔽毒内容的搜索关键词,提升检测效率。
QExplorer: Large Language Model Based Query Extraction for Toxic Content Exploration
- 基于两阶段训练:指令微调+偏好优化,生成精准查询
- 离线实验优于多个大模型和人工,线上部署毒内容检出率显著提升
- 适合安全团队、内容审核系统快速构建有效查询策略
在信息检索中,自动提取有效查询极具挑战性,尤其在隐蔽毒内容探索场景下。得益于生成式大语言模型(LLM)的发展,我们可直接利用其能力生成用于类似内容探索的有效查询。本文提出 QExplorer,一种基于大语言模型的有毒内容探索查询提取方法。该方法包含两阶段训练流程:指令监督微调(SFT)与使用直接偏好优化(DPO)的偏好对齐,并结合搜索系统反馈构建训练数据集。通过在真实系统上开展一系列离线与在线实验验证有效性。离线实验证明,自动查询提取性能优于多个 LLM 及人类表现;在线部署显示毒内容检出率显著提升。
原文摘要 · Abstract (English)
Automatically extracting effective queries is challenging in information retrieval, especially in toxic content exploration, as such content is likely to be disguised. With the recent achievements in generative Large Language Model (LLM), we are able to leverage the capabilities of LLMs to extract effective queries for similar content exploration directly. This study proposes QExplorer, an approach of large language model based Query Extraction for toxic content Exploration. The QExplorer approach involves a 2-stage training process: instruction Supervised FineTuning (SFT) and preference alignment using Direct Preference Optimization (DPO), as well as the datasets construction with feedback of search system. To verify the effectiveness of QExplorer, a series of offline and online experiments are conducted on our real-world system. The offline empirical results demonstrate that the performance of our automatic query extraction outperforms that of several LLMs and humans. The online deployment shows a significant increase in the detection of toxic items.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。