arXiv:2510.02712cs.CLcs.AI2025-10被引 6

用生存分析方法评估大模型对话鲁棒性,发现语义突变最危险。

Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks

  • 将对话失效建模为时间事件,结合生存分析与语义漂移特征
  • 语义突变使不一致风险飙升,累计漂移反而降低风险
  • 可实时预警对话崩溃,适合部署在实际对话系统中

大型语言模型(LLMs)在对话AI中取得突破,但其在多轮对话中的鲁棒性仍不清晰。现有评估聚焦静态基准与单轮测试,难以捕捉真实交互中的时序退化现象。本文对9个先进LLM在MT-Consistency基准上的36,951轮对话进行大规模生存分析,将失败建模为时间到事件过程。结合Cox比例风险、加速失效时间(AFT)与随机生存森林模型,使用简单语义漂移特征。结果表明:突发的提示间语义漂移显著提升不一致风险,而累积漂移却意外具有保护作用,暗示幸存对话存在适应机制。包含模型-漂移交互的AFT模型在区分度与校准性上最优;比例风险检验显示关键漂移变量存在系统性违背,解释了Cox模型在此场景的局限。最后,轻量级AFT模型可转化为逐轮风险监控器,在首次不一致前数轮预警,且误报率可控。研究确立生存分析在多轮鲁棒性评估中的有效性,并为对话AI系统设计实用防护提供支持。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have revolutionized conversational AI, yet their robustness in extended multi-turn dialogues remains poorly understood. Existing evaluation frameworks focus on static benchmarks and single-turn assessments, failing to capture the temporal dynamics of conversational degradation that characterize real-world interactions. In this work, we present a large-scale survival analysis of conversational robustness, modeling failure as a time-to-event process over 36,951 turns from 9 state-of-the-art LLMs on the MT-Consistency benchmark. Our framework combines Cox proportional hazards, Accelerated Failure Time (AFT), and Random Survival Forest models with simple semantic drift features. We find that abrupt prompt-to-prompt semantic drift sharply increases the hazard of inconsistency, whereas cumulative drift is counterintuitively \emph{protective}, suggesting adaptation in conversations that survive multiple shifts. AFT models with model-drift interactions achieve the best combination of discrimination and calibration, and proportional hazards checks reveal systematic violations for key drift covariates, explaining the limitations of Cox-style modeling in this setting. Finally, we show that a lightweight AFT model can be turned into a turn-level risk monitor that flags most failing conversations several turns before the first inconsistent answer while keeping false alerts modest. These results establish survival analysis as a powerful paradigm for evaluating multi-turn robustness and for designing practical safeguards for conversational AI systems.

对话鲁棒性生存分析大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。