测试大模型能否提前预测不道德行为,提升AI安全预警能力
PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
- 构建1000组正反向动作轨迹,评估模型预判能力
- 强模型在部分轨迹下预测准确率不足60%,仍具挑战
- 适合关注AI伦理、安全监控的研究者与开发者
大型语言模型(LLMs)正被用于执行多步动作以达成目标的自主代理。现有安全研究主要关注完整动作轨迹中不道德行为的检测,但这种模式本质上是事后追责。本文提出一种关键但被忽视的安全任务——预测性监控:仅基于部分动作轨迹,模型能否在显性行为发生前推断出其将导致不道德结果?为此,我们构建了PreActBench基准,包含1000对涵盖五个领域的伦理与非伦理动作轨迹。我们使用前缀预见性F1(Prefix Foresight F1)指标,在不同轨迹长度下评估多种LLM、安全防护模型及潜在探测方法。结果显示,尽管人类表现良好,即使强模型在部分轨迹上的预测准确率也低于60%,凸显了未来风险推理在大模型安全中的迫切需求。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective. While existing safety research has focused on detecting unethical behavior from complete trajectories, this paradigm is fundamentally retrospective: it identifies harm only after it has already occurred. In this work, we study a critical yet overlooked safety task, which we term Predictive Monitoring: given only a partial action trajectory, can a model infer whether it will culminate in an unethical action before the overt action is executed? To support this task, we present PreActBench, a benchmark of 1,000 paired ethical and unethical action trajectories spanning five domains. We evaluate a range of LLMs, safety guardrail models, and latent probing methods across varying fractions of the action trajectory using our Prefix Foresight F1 metric. Results show that while humans achieve promising performance, predictive monitoring remains challenging even for strong models, highlighting the need for future-oriented risk reasoning in LLM safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。