arXiv:2604.19071cs.CL2026-04ACL

提出树状写作评估框架,更真实地衡量大模型写作能力

HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing

论文配图:HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
图 1 · 摘自论文原文
  • 用树状结构建模写作各维度权重,解决评分偏差问题
  • 在1302条指令上达成0.93的人类判断相关性
  • 对文本扰动鲁棒,适合中文长文生成评估

大语言模型写作能力的评估仍面临挑战,因其写作技能具有多维特性,且现有度量方法存在局限。传统基于参考文本的指标或现代大模型作为评判者的方法,难以有效评估千字级、开放式写作任务。为此,我们提出树状写作(ToW)框架,通过显式建模子特征的聚合权重,解决大模型评判中常出现的隐含不一致问题。同时构建了涵盖12种文体、1302条指令的中文大规模写作评测基准HowToBench,覆盖情境补全、提纲引导写作与开放式生成三类任务。ToW成功缓解了评分偏倚,在人类判断上达到0.93的皮尔逊相关系数。此外,我们发现重叠度指标和主流大模型评判方法易受文本扰动影响,而ToW具备鲁棒性。还揭示在提纲任务中输入长度与内容得分呈负相关,说明单纯堆叠输入信息无法提升质量。

原文摘要 · Abstract (English)

Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills and the limitations of existing metrics. LLM's performance in thousand-words level and open-ended writing is inadequately assessed by traditional reference-based metrics or modern LLM-as-a-judge methods. We propose Tree-of-Writing (ToW), to resolve the implicit inconsistency often found when LLM-as-a-judge aggregates all sub-features in text evaluation. ToW incorporates a tree-structured workflow by explicitly modeling the aggregation weights of sub-features. We also present HowToBench, a large-scale Chinese writing benchmark encompassing 12 genres and 1302 instructions across three task categories: contextual completion, outline-guided writing, and open-ended generation. ToW successfully mitigates the biases, achieving a 0.93 Pearson correlation with human judgments. Furthermore, we detect that both overlap-based text generation metrics and popular LLM-as-a-judge practices are vulnerable to textual disturbances, while ToW is robust to them. We also uncover a negative correlation between input length and content-related scores in the Guide task, showcasing that it cannot be simply improved by input-side information piling.

大模型评估写作生成中文基准树状结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。