大模型常忘格式要求,尤其多重任务时更易出错。
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
- 用认知心理学中的前瞻记忆框架测试模型对格式指令的遵守情况。
- 多重任务下格式合规率下降2%-21%,末端约束最多跌50%。
- 加醒目提示可恢复90%-100%合规,适合需精准输出的场景。
大型语言模型在同时执行复杂任务时,常无法满足格式要求。本文基于认知心理学中的前瞻记忆理论,设计了一个控制实验,将可验证的格式约束与逐步增加复杂度的基准任务结合。在三个模型族和超过8000个提示中,任务并发导致合规率下降2%-21%。不同类型的约束表现差异显著:末端约束(要求在响应末尾执行动作)最脆弱,最高下降50%;而规避类约束相对稳健。通过增强提示显眼性(明确指令+末尾提醒),可大幅恢复合规性,多数情况下恢复至90%-100%。干扰效应具有双向性:格式约束也会降低任务准确性,例如某模型在GSM8K上的准确率从93%降至27%。在叠加实验中,随着约束数量增加,联合合规性急剧下降。所有结果均使用确定性程序化检查器验证,未采用LLM作为评判者,数据来自公开数据集。
原文摘要 · Abstract (English)
Large language models often fail to satisfy formatting instructions when they must simultaneously perform demanding tasks. We study this behaviour through a prospective memory inspired lens from cognitive psychology, using a controlled paradigm that combines verifiable formatting constraints with benchmark tasks of increasing complexity. Across three model families and over 8,000 prompts, compliance drops by 2-21% under concurrent task load. Vulnerability is highly type-dependent: terminal constraints (requiring action at the response boundary) degrade most, with drops up to 50%, while avoidance constraints remain comparatively robust. A salience-enhanced format (explicit instruction framing plus a trailing reminder) recovers much of the lost compliance, restoring performance to 90-100% in many settings. Interference is bidirectional: formatting constraints can also reduce task accuracy, with one model's GSM8K accuracy dropping from 93% to 27%. In additional stacking experiments, joint compliance declines sharply as constraints accumulate. All results use deterministic programmatic checkers without an LLM-as-judge component on publicly available datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。