通过轨迹级预测,提前发现大模型交互中的潜在安全风险。
Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

- 构建双尺度视角,融合短期对话与长期上下文提取风险证据。
- 可提前2.41轮预测88.3%的安全事故,误报率仅12.3%。
- 适合需要主动防范长程恶意行为的AI系统安全团队使用。
随着大语言模型从独立助手演变为自主代理,保障其安全需从单轮风险评估转向理解长时程轨迹中的风险演化。在多轮交互中,恶意意图可能被分散于看似无害的对话回合,逐步重构并最终导致安全失败。现有防护机制仍以被动检测为主,难以预判潜在风险演变。为此,本文提出Recast框架,将安全防护从轮次级违规检测推进至轨迹级风险预测。该框架首先通过双尺度轨迹视图,从短期对话进展和长期历史上下文中检索风险相关证据;接着建模风险的组合演化,捕捉当前风险状态及其时间动态;最后利用因果时序编码器学习隐式风险演化模式,并预测未来风险爆发轮次的分布。在7类风险上的实验表明,Recast能以平均2.41轮的提前量预测88.3%的安全事故,同时保持12.3%的误报率,验证了轨迹级预测在事前识别新兴风险的有效性。
原文摘要 · Abstract (English)
As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories. In multi-turn interactions, malicious intent can be decomposed across seemingly harmless turns and gradually reconstructed through interaction trajectories, eventually resulting in safety failures. Existing safeguards remain largely reactive, detecting manifested violations while lacking the ability to predict latent risk evolution and enable preemptive prevention. To address this limitation, we propose Recast, a safety risk forecasting framework that advances LLM safeguarding beyond turn-level violation detection to trajectory-level risk prediction. Recast first retrieves risk-relevant evidence from both short-term dialogue progression and long-term historical context via a dual-scale trajectory view. It then models compositional risk evolution by capturing the current risk configuration and its temporal dynamics. Finally, a causal temporal encoder learns latent risk evolution patterns and predicts the distribution of future risk emergence turns. Extensive experiments across 7 risk categories show that Recast predicts 88.3% of future safety failures with an average lead time of 2.41 turns, while maintaining a false alarm rate of 12.3%, showcasing the effectiveness of trajectory-level forecasting in identifying emerging risks before safety violations occur.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。