arXiv:2607.11503cs.CL2026-07

通过可迭代优化的技能循环,提升长文生成质量

GEIS: A Generation-Evaluation-Improvement Loop of Agent Skills for Long-Form Article Generation

论文配图:GEIS: A Generation-Evaluation-Improvement Loop of Agent Skills for Long-Form Article Generation
图 1 · 摘自论文原文
  • 构建命名化、声明式的技能循环,实现写作、评估与改进的闭环
  • 在20个维基专题文章上,质量评分从82.90提升至86.95,内容质量显著改善
  • 适合需要持续优化长文生成能力的研究者与工程团队

长篇文本生成对大语言模型仍具挑战,因其需处理长上下文、长指令和长输出。现有多智能体流程如STORM通过角色专业化提升信息覆盖率,但其能力常耦合于提示与固定流程,难以审查、复用或迭代改进。本文提出GEIS(生成-评估-改进循环的智能体技能),一种用于维基风格长文生成的命名且声明式技能循环。在Tasi Harness中实现并评估,包含文章写作、浏览器证据与图像收集、图表渲染、基于PDF的成对评估及规则级技能改进等技能。核心写作技能遵循请求、规划、草稿、审计、精炼、交付流程;成对评估生成结构化质量报告;改进技能将重复发现转化为永久性修复补丁。在20个维基特色文章主题上评估,相同生成后端下,相比Tasi Harness默认写作者提升8.0分(满分100),优于STORM在结构质量与内容质量两个维度的表现。在20主题改进实验中,修补后的写作者平均分由82.90升至86.95,其中17个主题得分提升,增益主要来自内容质量。结果表明,长文生成可从固定流程重构为可检查、模块化且评估驱动的改进循环。

原文摘要 · Abstract (English)

Long-form article generation remains difficult for large language models because it combines long context, long instructions, and long outputs. Existing multi-agent pipelines such as STORM improve information coverage by simulating role-specialized agents, but their capabilities are often entangled in prompts and fixed procedures, making them hard to inspect, reuse, or iteratively improve. This paper presents GEIS (Generation-Evaluation-Improvement loop of agent Skills), a loop of named and declarative skills for Wikipedia-style long-form article generation. Implemented and evaluated in Tasi Harness, GEIS composes skills for article writing, browser-based evidence and image collection, diagram rendering, PDF-aware pairwise evaluation, and rule-level skill improvement. Its core writing skill follows Request, Plan, Draft, Audit, Refine, and Deliver; the pairwise evaluation skill produces structured quality reports; and the improvement skill maps recurrent findings into permanent patches to the writing skill in our 20-topic experiment. We evaluate GEIS on 20 Wikipedia Featured Article topics. Under the same generation backend, GEIS improves over the Tasi Harness default writer by 8.0 points on a 100-point PDF quality rubric and outperforms STORM on the two comparable writing dimensions, structural quality and content quality. In the 20-topic improvement experiment, the patched writing skill raises the average score from 82.90 to 86.95, with 17 out of 20 topics improved and the gain mainly coming from content quality. These results show that long-form generation can be reframed from a fixed workflow into an inspectable, modular, and evaluation-guided improvement loop.

长文生成智能体评估改进维基

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。