arXiv:2606.06055cs.AI2026-06

测试对话模型在不主动要求时是否会泄露敏感历史记录

HUSH-Bench: Measuring Memory-Use Boundaries for Sensitive History in Conversational Agents

论文配图:HUSH-Bench: Measuring Memory-Use Boundaries for Sensitive History in Conversational Agents
图 1 · 摘自论文原文
  • 构建2400个配对提示,评估模型在无明确请求下是否无意引用敏感历史
  • 检索式记忆使敏感信息泄露率升至23%~30%,部分模型泄露评分达83.0
  • 明确要求使用历史可提升记忆调用率27%~41%,适合关注隐私安全的研究者

长期记忆有助于对话代理在会话间保持连贯性,但相关性与当前回合的使用许可仍是独立决策。本文研究在保守策略下,敏感历史如何在当前回合提供使用理由时被触发。提出HUSH-Bench基准,包含2,400个良性提示,每个配对历史中含一条标记的敏感披露及对应的无记忆参照。该基准衡量未请求的历史整合程度(UIS,0–100分,越高越差),记录敏感信息是否被生成,并包含仅在是否请求使用上下文上不同的配对提示。在四种模型上评估了无记忆、全上下文及三种基于检索的记忆设置。记忆访问使一个模型的UIS从接近零上升至8.9–26.6,其余三个模型则达51.3–83.0。检索系统在23.0%–30.3%情况下暴露了标记的敏感披露,相关敏感条目或摘要仍可访问,且三模型持续显示高UIS。在四个生成器中,明确邀请使目标记忆调用率提升27.0%–41.3%,帮助性评分稳定,平均过度引用率上升。

原文摘要 · Abstract (English)

Long-term memory helps conversational agents maintain continuity across sessions, while relevance and current-turn warrant remain distinct decisions. We study this boundary under a stated conservative policy in which sensitive history shapes a response when the current turn supplies a reason to use it. We introduce HUSH-Bench, a controlled benchmark of 2,400 benign prompts paired with histories containing one marked sensitive disclosure and matched no-memory references. HUSH-Bench measures unsolicited history integration with the Unsolicited History Integration Score (UIS; 0--100, higher is worse), records whether the marked disclosure reaches the generator, and includes paired prompts that differ only in whether the user asks the assistant to use earlier context. We evaluate four models under no-memory, full-context, and three retrieval-based memory settings. Memory access raises UIS from near zero to 8.9--26.6 for one model and 51.3--83.0 for the other three. Retrieval systems expose the marked disclosure in 23.0\%--30.3\% of cases, while related sensitive entries or summaries remain available and three models continue to show high UIS. Across four generators, an explicit invitation increases target-memory uptake scores by 27.0--41.3; measured helpfulness remains stable while mean over-scope rises. These results motivate treating memory storage, retrieval, warrant, and per-turn scope as separate design decisions.

对话系统隐私安全记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。