用进化算法生成多轮欺骗性提问,通过嵌入空间几何特征实现高效检测。
Evolving and Detecting Multi-Turn Deception using Geometric Signatures

- 通过多目标遗传优化生成逼真的多轮欺骗性问题集。
- 几何特征组合检测召回率达0.89,三轮场景下F1为0.74-0.86。
- 适合部署于安全系统,无需复杂训练即可透明筛查欺骗意图。
大型语言模型的安全防御通常针对单轮提示进行训练与评估,但真实攻击常表现为间接的多轮探测。为此,我们提出统一管道,通过多目标遗传提示优化与协同进化突变算子生成逼真的多轮欺骗性问题集。人类研究验证了该数据集,发现早期生成版本欺骗性最强,且存在遵循过滤和顺序效应等实际约束。基于此数据,我们利用嵌入空间中简单的可解释几何信号(角度覆盖、距离比、线性度)结合轻量级前馈分类器,成功检测出试图获取禁止信息的欺骗行为。三个几何特征辅以成对相似性统计,构建出紧凑预测模型,在基础、重述及截断(三轮)场景中均保持0.89的高召回率,测试时F1为0.74–0.86。结果支持核心假设:多轮欺骗意图在嵌入空间中留下稳定几何痕迹,可实现无需昂贵端到端训练的轻量级、透明化筛查。我们进一步讨论了负责任使用、局限性及构建更大更多样人类评估数据集的路径。主要贡献在于多目标演化提示生成框架,工程应用则是可部署的轻量几何检测系统用于大模型安全基础设施。
原文摘要 · Abstract (English)
Safety defenses for large language models (LLMs) are typically trained and evaluated on single-turn prompts, yet real attacks often unfold as indirect, multi-turn probing. To defend against this more nuanced form of deception, we present a unified pipeline that generates realistic multi-turn deceptive question sets via multi-objective genetic prompt optimization with co-evolving mutation operators. We validate this dataset through a human study, which also revealed that early generations yielded the most convincing deception and practical constraints such as adherence filtering and ordering effects. Using this data, we were able to detect deceptive attempts to access prohibited information using simple, explainable geometric signals in embedding space coupled with a lightweight feed-forward classifier. Three geometric features (angular coverage, distance ratio, and linearity) augmented with pairwise similarity statistics led to a compact predictive model that achieved consistently high recall (0.89) across base, reworded, and truncated (three-turn) scenarios, with test-time F1 ranging from 0.74-0.86. The results support a central hypothesis that multi-turn deceptive intent leaves a stable geometric footprint that enables lightweight, transparent screening without expensive end-to-end training. We further discuss responsible uses, limitations, and paths toward larger, more diverse human-evaluated datasets. The primary contribution to artificial intelligence is the multi-objective evolutionary framework for prompt generation, and the engineering application is the deployment of a lightweight geometric detection system for LLM safety infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。