arXiv:2503.19711cs.CLcs.AI2025-03被引 8

让大模型当合作者写文章,测试其自主改进能力

Writing as a testbed for open ended agents

  • 让大模型自动提出并执行文本优化建议
  • 三款模型在持续改写中表现差异明显
  • 适合研究自主生成与人机协作的学者

开放性任务对大语言模型极具挑战,因其解空间巨大且成功标准主观。写作兼具广阔解空间和主观评价特性,是理想测试场景。本文研究大模型作为协作写作者的潜力,考察其自主提出并实施文本改进的能力。聚焦Gemini 1.5 Pro、Claude 3.5 Sonnet和GPT-4o三款模型,分析其行动多样性、与人类意图对齐程度及迭代优化能力对整体表现的影响。本工作建立了一个评估自主写作代理的基准框架,更广泛地揭示了构建能在多样化开放领域中表现优异系统所面临的根本挑战与潜在解决方案。

原文摘要 · Abstract (English)

Open-ended tasks are particularly challenging for LLMs due to the vast solution space, demanding both expansive exploration and adaptable strategies, especially when success lacks a clear, objective definition. Writing, with its vast solution space and subjective evaluation criteria, provides a compelling testbed for studying such problems. In this paper, we investigate the potential of LLMs to act as collaborative co-writers, capable of suggesting and implementing text improvements autonomously. We analyse three prominent LLMs - Gemini 1.5 Pro, Claude 3.5 Sonnet, and GPT-4o - focusing on how their action diversity, human alignment, and iterative improvement capabilities impact overall performance. This work establishes a framework for benchmarking autonomous writing agents and, more broadly, highlights fundamental challenges and potential solutions for building systems capable of excelling in diverse open-ended domains.

大模型自主写作人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。