提出对话风险累积框架,检测多轮对话中逐渐浮现的安全隐患
Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework

- 通过跟踪语义偏移、敏感信息积累和合规意愿变化三类轨迹信号
- 在1200个八轮会话上实现高精度风险检测,平均提前3轮发现潜在危害
- 适合需要长期对话安全防护的系统开发者与安全研究员
当前大模型安全防护大多孤立评估每轮输入输出,忽略了多轮对话中渐进式风险累积现象——即看似无害的对话逐步积累成有害意图、碎片化拼凑违规指令、敏感信息反复披露导致风险升高。本文提出会话层风险累积(CRA)框架,追踪三类轨迹信号:与会话锚点的语义偏移、基于提取实体的敏感度加权信息累积图、以及反映合规意愿上升的梯度信号。评分方面,提供无监督凸融合用于归因分析,并设计紧凑的CRA-Net DA模型,通过家族对抗目标训练以减少长度和主题覆盖偏差。为评测该框架,发布CRA-Bench v0.1(1200个八轮会话,涵盖三类威胁并配有主题匹配的良性对照),v0.2(LLM重述版本以降低模板痕迹),以及扩展至五类的2000个会话(新增角色提示与上下文填充)。引入原生轨迹评估协议,包含会话级划分、混合集阈值校准、轨迹AUROC、检测延迟、校准误报率、自助置信区间、留一家族诊断压力测试及合成到真实迁移检验。主要验证集中在CRA-Bench内部分布会话评分及人类迁移子集表现。
原文摘要 · Abstract (English)
Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk Accumulation (CRA): gradual intent drift, fragmented assembly of prohibited instructions, and sensitivity build-up from repeated disclosures. We propose a session-layer CRA Framework that tracks three trajectory signals: semantic drift from a session anchor, a sensitivity-weighted information accumulation graph over extracted entities, and a compliance-gradient signal capturing increasing willingness to comply. For scoring, we provide (i) an unsupervised convex fusion for attribution and ablations, and (ii) CRA-Net DA, a compact learned trajectory model trained with family-adversarial objectives to reduce length and topic-coverage confounds. To benchmark CRA, we release CRA-Bench v0.1 (1,200 eight-turn sessions across three threat families with topic-matched benign twins), CRA-Bench v0.2 (LLM-paraphrased variants to reduce template artifacts), and an extended 5-family set (2,000 sessions adding persona priming and context stuffing). We introduce a trajectory-native evaluation protocol with session-level splits, mixed-set threshold calibration, Trajectory AUROC, turns-to-detection, calibrated false-positive metrics, bootstrap confidence intervals, leave-one-family-out diagnostic stress tests, and synthetic-to-human transfer checks. Claims focus on within-distribution session scoring on CRA-Bench and human-transfer subsets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。