大模型生成的虚假证据会严重干扰恶意文本检测,需警惕其污染风险。
On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs
- 用大模型伪造或重写证据,误导检测系统
- 生成式污染导致性能下降最高达14.4%
- 现有防御策略有效但难落地,适合安全研究者关注
基于证据的恶意社交文本检测器表现出强大能力,但大语言模型(LLMs)的兴起带来了证据污染的潜在风险。本文探讨了基础污染、以及由LLM重写或生成证据的攻击场景。为缓解负面影响,提出三种防御策略:机器生成文本检测、专家混合模型和参数更新。在四个恶意社交文本检测任务及十个数据集上的实验表明,证据污染显著削弱检测器性能,其中生成策略导致最高14.4%的准确率下降。尽管防御策略能缓解污染,但在实际应用中仍存在局限。进一步分析显示,污染证据具有高可信度(人工与指标评估一致),会破坏模型校准,使期望校准误差上升至21.6%;且可叠加放大危害,尤其对编码器型大模型,准确率下降达21.8%。
原文摘要 · Abstract (English)
Evidence-enhanced detectors present remarkable abilities in identifying malicious social text. However, the rise of large language models (LLMs) brings potential risks of evidence pollution to confuse detectors. This paper explores potential manipulation scenarios including basic pollution, and rephrasing or generating evidence by LLMs. To mitigate the negative impact, we propose three defense strategies from the data and model sides, including machine-generated text detection, a mixture of experts, and parameter updating. Extensive experiments on four malicious social text detection tasks with ten datasets illustrate that evidence pollution significantly compromises detectors, where the generating strategy causes up to a 14.4% performance drop. Meanwhile, the defense strategies could mitigate evidence pollution, but they faced limitations for practical employment. Further analysis illustrates that polluted evidence (i) is of high quality, evaluated by metrics and humans; (ii) would compromise the model calibration, increasing expected calibration error up to 21.6%; and (iii) could be integrated to amplify the negative impact, especially for encoder-based LMs, where the accuracy drops by 21.8%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。