arXiv:2605.29857cs.LG2026-05

从写作批注中提炼可复用的评分标准,让AI更懂专家评审逻辑。

Feedback-to-Rubrics: Can We Learn Expert Criteria from Inline Comments?

论文配图:Feedback-to-Rubrics: Can We Learn Expert Criteria from Inline Comments?
图 1 · 摘自论文原文
  • 通过分析批注与预测的差异,迭代优化自然语言评分标准。
  • 在真实和可控场景中,能有效预测批注并支持自动修改。
  • 适合需要统一评审标准的教育、出版或AI辅助写作场景。

大型语言模型(LLMs)在写作与评审支持中应用日益广泛,但其有效性依赖于上下文相关的评判标准,如专家偏好或组织特定规范,这些标准常为隐性、未记录且难以直接获取。本文提出一种从人工撰写或由LLM生成的稿件所积累的内联批注中学习可复用的自然语言评分标准的新问题设定。我们的方法从批注中推断评分标准,并通过观察评分标准条件下的预测与参考批注之间的逐条不一致,迭代优化评分标准。我们在真实评审场景及具有参考评分标准的受控环境中评估该方法。结果表明,内联批注可被提炼为可复用的评分标准,支持批注预测、评分标准理解与自动内容修订。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for writing and review support, but their usefulness depends on context-dependent criteria, such as expert preferences or organization-specific conventions, that are often tacit, undocumented, and difficult to elicit directly. We propose a problem setting for learning reusable natural-language rubrics from accumulated inline comments on artifacts such as human-written or LLM-generated drafts. Our method infers rubrics from these comments and iteratively refines them by observing comment-wise mismatches between rubric-conditioned predictions and reference comments. We evaluate the proposed method in real-world review settings and in controlled settings with reference rubrics. These results show that inline comments can be distilled into reusable rubrics that support comment prediction, rubric understanding, and automatic artifact revision.

评分标准AI评审批注分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。