大模型在长期委托任务中会严重损坏文档,平均25%内容出错。
LLMs Corrupt Your Documents When You Delegate

- 构建52个专业领域长流程任务评估模型委托可靠性。
- 前沿模型如GPT 5.4、Claude 4.6 Opus仍导致25%内容被破坏。
- 文档越长、干扰文件越多,错误越严重,工具调用无改善作用。
大型语言模型(LLMs)正重塑知识工作,委托式交互(如氛围编程)成为新范式。信任是委托的核心——期望模型忠实执行任务且不引入错误。我们提出DELEGATE-52,模拟跨52个专业领域的长周期文档编辑任务,涵盖编码、晶体学、乐谱等。对19个主流模型的大规模实验显示,当前模型在委托过程中会破坏文档:即使顶尖模型(Gemini 3.1 Pro、Claude 4.6 Opus、GPT 5.4)在长流程末期也平均造成25%的内容损坏,其他模型表现更差。额外实验表明,使用智能体工具无法提升性能,且文档长度、交互时长或干扰文件的存在会加剧错误。分析显示,当前模型不可靠:它们引入稀疏但严重的错误,随长时间交互不断累积。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding). Delegation requires trust - the expectation that the LLM will faithfully execute the task without introducing errors into documents. We introduce DELEGATE-52 to study the readiness of AI systems in delegated workflows. DELEGATE-52 simulates long delegated workflows that require in-depth document editing across 52 professional domains, such as coding, crystallography, and music notation. Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows, with other models failing more severely. Additional experiments reveal that agentic tool use does not improve performance on DELEGATE-52, and that degradation severity is exacerbated by document size, length of interaction, or presence of distractor files. Our analysis shows that current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding over long interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。