多轮对话中,模型回答真实性随对话长度变化,暴露单轮测试无法发现的漏洞。
Is Length Really A Liability? An Evaluation of Multi-turn LLM Conversations using BoolQ
- 设计多轮对话实验,测试不同长度与引导方式下的模型表现。
- 三种模型在长对话中均出现真实性下降,且影响模式各异。
- 适合关注实际应用中模型安全性的研究者和开发者参考。
当前大模型评估多依赖单轮提示,但难以捕捉真实场景中出现的对话动态与潜在危害。本研究通过在 BoolQ 数据集上设置不同长度和引导条件的多轮对话,评估了三种大模型的回应真实性。结果表明,模型在多轮对话中表现出特定于自身结构的脆弱性,这些缺陷在单轮测试中完全不可见。我们观察到对话长度与引导方式对模型表现有显著影响,揭示了静态评估的根本局限——只有在多轮对话环境下,部署相关的风险漏洞才可能被识别。
原文摘要 · Abstract (English)
Single-prompt evaluations dominate current LLM benchmarking, yet they fail to capture the conversational dynamics where real-world harm occurs. In this study, we examined whether conversation length affects response veracity by evaluating LLM performance on the BoolQ dataset under varying length and scaffolding conditions. Our results across three distinct LLMs revealed model-specific vulnerabilities that are invisible under single-turn testing. The length-dependent and scaffold-specific effects we observed demonstrate a fundamental limitation of static evaluations, as deployment-relevant vulnerabilities could only be spotted in a multi-turn conversational setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。