用对抗攻击让谎言看起来更真,能骗过人和机器判断。
Effective faking of verbal deception detection with target-aligned adversarial attacks
- 用大模型改写谎言,使其符合目标判断者的特点
- 对齐目标时,人和机器的识别准确率降至随机水平
- 研究提醒警惕可被轻松利用的伪造风险,适合安全与伦理研究者
通过分析语言进行欺骗检测是结合人工与机器学习判断的潜在方向。自动化对抗攻击可通过重写欺骗性陈述使其显得真实,构成严重威胁。本研究使用包含243个真实与262个虚构自传故事的数据集,测试人类和机器学习模型在欺骗检测任务中的表现。实验1中,人类基于细节程度启发式或直接判断原始/对抗修改后的陈述,同时使用微调语言模型和简单n-gram模型进行判断。实验2中,操控攻击的目标一致性,即是否针对人类或机器评估者定制修改内容。结果显示:当对抗修改与目标一致时,人类判断(d=-0.07、d=-0.04)与机器判断(51%准确率)均降至随机水平;若未对齐目标,则人类启发式判断(d=0.30、d=0.36)和机器预测(63%-78%)显著优于随机。结论表明,易获取的语言模型可有效帮助任何人伪造欺骗检测结果,无论针对人类还是机器。抵御对抗修改的鲁棒性依赖于目标对齐。研究建议未来应结合对抗攻击设计推动欺骗检测研究发展。
原文摘要 · Abstract (English)
Background: Deception detection through analysing language is a promising avenue using both human judgments and automated machine learning judgments. For both forms of credibility assessment, automated adversarial attacks that rewrite deceptive statements to appear truthful pose a serious threat. Methods: We used a dataset of 243 truthful and 262 fabricated autobiographical stories in a deception detection task for humans and machine learning models. A large language model was tasked to rewrite deceptive statements so that they appear truthful. In Study 1, humans who made a deception judgment or used the detailedness heuristic and two machine learning models (a fine-tuned language model and a simple n-gram model) judged original or adversarial modifications of deceptive statements. In Study 2, we manipulated the target alignment of the modifications, i.e. tailoring the attack to whether the statements would be assessed by humans or computer models. Results: When adversarial modifications were aligned with their target, human (d=-0.07 and d=-0.04) and machine judgments (51% accuracy) dropped to the chance level. When the attack was not aligned with the target, both human heuristics judgments (d=0.30 and d=0.36) and machine learning predictions (63-78%) were significantly better than chance. Conclusions: Easily accessible language models can effectively help anyone fake deception detection efforts both by humans and machine learning models. Robustness against adversarial modifications for humans and machines depends on that target alignment. We close with suggestions on advancing deception research with adversarial attack designs and techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。