arXiv:2510.24345cs.CLcs.AI2025-10EMNLP被引 3

构建可验证的长文本生成评估基准,解决真实场景与可测性矛盾。

LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability

  • 基于真实场景设定可验证目标,自动生成查询与约束
  • 23个大模型在64K输入下表现显著下降,验证长文本瓶颈
  • 支持定制化长度,适配研究者与产品团队测试需求

长文本生成在事实性、信息量和准确性方面仍是大语言模型的重大挑战。现有长文本生成基准或采用难以验证的真实问题,或使用简化合成数据忽略真实复杂性。本文提出LongWeave,结合约束验证评估(CoV-Eval)实现真实性和可验证性的平衡:先定义真实场景中的可验证目标,再系统生成对应查询、文本材料与约束条件。该方法确保任务兼具现实意义与客观评估基础,能严格检验模型在满足复杂现实约束下的能力。LongWeave支持高达64K/8K token的输入输出长度,覆盖七类任务。对23个大模型的评估显示,随着真实复杂度和输出长度增加,即使是先进模型也面临显著挑战。

原文摘要 · Abstract (English)

Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies. In this paper, we introduce \textbf{LongWeave}, which balances real-world and verifiable assessment with Constraint-Verifier Evaluation (CoV-Eval). CoV-Eval constructs tasks by first defining verifiable targets within real-world scenarios, then systematically generating corresponding queries, textual materials, and constraints based on these targets. This ensures that tasks are both realistic and objectively assessable, enabling rigorous assessment of model capabilities in meeting complex real-world constraints. LongWeave supports customizable input/output lengths (up to 64K/8K tokens) across seven distinct tasks. Evaluation on 23 LLMs shows that even state-of-the-art models encounter significant challenges in long-form generation as real-world complexity and output length increase.

长文本生成评估基准可验证性LLM评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。