提出新框架,让个性化语言系统不再只靠回忆,而是可控承诺。
Recall Isn't Enough: Bounding Commitments in Personalized Language Systems
- 用限定证据激活机制,结合类型覆盖和后果债务控制记忆使用。
- 在360个测试中实现零失败,可用率仅0.49-0.60,远超基线。
- 适合需高可靠性、低冗余的个性化对话系统开发者。
长上下文与记忆系统通常将个性化视为召回问题。但实际中,系统常在承诺阶段出错:将模糊提示转为刚性约束、忽略罕见证据、遗忘后续责任或在不可行时仍作答。本文提出契约约束证据激活(CBEA)与词典承诺验证(LCV)。CBEA通过类型覆盖、尾部见证和后果债务,激活有限证据集;LCV在生成前验证结构化承诺,对不可行状态引导修复、放弃或重订契约。在360个测试用例及三种生成后端中,CBEA+LCV在0.49-0.60可用率下实现验证范围内的零失败,而原始与长上下文基线需0.003-0.092才能达到零失败。影子探测器显示:CBEA+LCV仅召回0.012未编译可见事实,原始模型召回0.53。结果为可控操作点:显式承诺控制与74-75%更低的平均输入负载,非全内存主导。
原文摘要 · Abstract (English)
Long-context and memory systems usually treat personalization as a recall problem. In practice, many failures occur later, when a system commits: it turns noisy hints into hard constraints, drops rare witnesses, forgets downstream obligations, or answers despite infeasibility. We introduce Contract-Bounded Evidence Activation (CBEA) with Lexicographic Commitment Validation (LCV). CBEA activates a bounded evidence set using typed coverage, tail witnesses, and consequence debt; LCV validates structured commitments before prose and routes infeasible states to repair, abstention, or recontract. Across 360 fixtures and three generation backends, CBEA+LCV reaches zero failures within validator scope at 0.49-0.60 availability over attempted runs. Raw and long-context baselines with the same LCV gate reach zero only at 0.003-0.092. A shadow oracle diagnostic marks the limit: CBEA+LCV recalls 0.012 of uncompiled visible facts, while raw recalls 0.53. The result is a bounded operating point: explicit commitment control and 74-75% lower median input payload, not universal memory dominance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。