arXiv:2506.04772cs.CL2025-06ACL被引 6

用大模型+领域指标混合评估科学写作修改效果

Identifying Reliable Evaluation Metrics for Scientific Text Revision

  • 用人工标注验证修改质量,对比多种评估方法
  • 大模型能评指令遵循性,但难判内容正确性
  • 结合大模型与专用指标可更可靠评估修改

科学写作修改的评估仍具挑战,传统指标如ROUGE和BERTScore侧重相似性而非实质性改进。本文通过人工标注研究不同修改的质量,考察无参考文本评估方法及大模型作为评判者的能力,分析其在有/无标准参考时的表现。结果表明,大模型擅长判断指令遵循情况,但在内容正确性上表现不佳;而领域特定指标提供互补信息。混合使用大模型评分与任务相关指标,能最可靠地评估修改质量。

原文摘要 · Abstract (English)

Evaluating text revision in scientific writing remains a challenge, as traditional metrics such as ROUGE and BERTScore primarily focus on similarity rather than capturing meaningful improvements. In this work, we analyse and identify the limitations of these metrics and explore alternative evaluation methods that better align with human judgments. We first conduct a manual annotation study to assess the quality of different revisions. Then, we investigate reference-free evaluation metrics from related NLP domains. Additionally, we examine LLM-as-a-judge approaches, analysing their ability to assess revisions with and without a gold reference. Our results show that LLMs effectively assess instruction-following but struggle with correctness, while domain-specific metrics provide complementary insights. We find that a hybrid approach combining LLM-as-a-judge evaluation and task-specific metrics offers the most reliable assessment of revision quality.

文本修订评估指标大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。