让评估直接指导答案优化,生成可执行的改进建议。
CriticGen: Generation-Aware Evaluation as Actionable Feedback

- 根据具体回答生成个性化评价维度和标准
- 评分与改进建议相关联,提升反馈可操作性
- 适合需要高质量反馈的模型迭代与教学场景
现有大语言模型评估方法粗粒度且脱离生成过程,给出的解释泛化,无法提供有效改进依据。我们提出CriticGen,一种细粒度、生成感知的评估框架,将评估转化为可执行的控制机制以改进答案。该框架在主观、客观及自衍生约束等高层类别下,生成针对具体样本的评价维度与打分标准,作为动态评分量表,联合输出得分、理由、可执行的优化建议与优化后的答案。这种量表引导的优化流程使模型能诊断问题并进行精准修正。实验表明,细粒度评估需具备实例特异性和可操作性。CriticGen生成更高质量量表,使相关性/覆盖度从3.33/4.03提升至3.97/4.24;评分相关性最优,皮尔逊系数达0.9556,斯皮尔曼系数为0.9560;基于标准的理由与建议的F1值由0.6369/0.5994提升至0.7554/0.7900。关键的是,其反馈可有效提升答案质量,73.17%的答案得到改善,非退化率高达93.28%。
原文摘要 · Abstract (English)
Current evaluation methods for large language models are coarse-grained and decoupled from generation, producing generic explanations that fail to provide actionable feedback for model improvement. We propose CriticGen, a fine-grained, generation-aware evaluation framework that turns evaluation into actionable control for answer improvement. CriticGen first generates sample-specific evaluation dimensions and scoring criteria under high-level categories such as subjective, objective, and self-derived constraints. These criteria then serve as a dynamic rubric for jointly producing a score, a reason, an executable refinement suggestion, and a refined answer. This rubric-conditioned refinement process enables models to diagnose flaws and perform targeted answer improvement. Experimental results show that fine-grained evaluation should be both instance-specific and actionable. CriticGen induces higher-quality rubrics, improving relevance/coverage from 3.33/4.03 to 3.97/4.24. CriticGen also achieves the best score correlations, with 0.9556 Pearson and 0.9560 Spearman, and raises the F1 of criterion-grounded reasons and executable suggestions from 0.6369/0.5994 to 0.7554/0.7900. Crucially, its feedback translates into reliable answer improvement, improving 73.17% of answers with a 93.28% non-degradation rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。