用GPT生成假资料,现有检测器失效,新方法让误判率降到7%以下
Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
- 用GPT生成对抗样本训练检测器,提升抗骗能力
- 检测器对GPT生成资料的误判率从42%-52%降至1%-7%
- 适合需要防伪造的社交平台与内容安全团队
大型语言模型(LLMs)使在领英等平台上创建逼真虚假资料变得更容易,对基于文本的假资料检测器构成重大威胁。本研究评估了现有检测器对LLM生成资料的鲁棒性:尽管在检测人工创建的假资料时表现良好(误接受率6-7%),但对GPT生成的资料却失败严重(误接受率42-52%)。我们提出一种GPT辅助的对抗训练方法,使误接受率恢复至1-7%,同时保持误拒率在0.5-2%的低水平。消融实验表明,结合数值与文本嵌入的检测器最具鲁棒性,其次为仅使用数值嵌入,最后是仅文本嵌入。对提示工程驱动的GPT-4Turbo与人工评估者的补充分析进一步证实了该方法的必要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made it easier to create realistic fake profiles on platforms like LinkedIn. This poses a significant risk for text-based fake profile detectors. In this study, we evaluate the robustness of existing detectors against LLM-generated profiles. While highly effective in detecting manually created fake profiles (False Accept Rate: 6-7%), the existing detectors fail to identify GPT-generated profiles (False Accept Rate: 42-52%). We propose GPT-assisted adversarial training as a countermeasure, restoring the False Accept Rate to between 1-7% without impacting the False Reject Rates (0.5-2%). Ablation studies revealed that detectors trained on combined numerical and textual embeddings exhibit the highest robustness, followed by those using numerical-only embeddings, and lastly those using textual-only embeddings. Complementary analysis on the ability of prompt-based GPT-4Turbo and human evaluators affirms the need for robust automated detectors such as the one proposed in this study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。