arXiv:2510.02337cs.CLcs.AI2025-10

CRACQ可多维度评估文本质量,比大模型评分更稳定可解释。

CRACQ: A Multi-Dimensional Approach To Automated Document Assessment

  • 基于五个维度构建可解释的自动评估框架
  • 在500份合成申请书上训练,表现优于直接用大模型打分
  • 适合需要细致分析文本质量的研究与评审场景

本文提出CRACQ,一种针对文档五项特质——连贯性、严谨性、适当性、完整性与质量——的多维评估框架。受基于特质的自动作文评分(AES)启发,CRACQ将评估范围从作文拓展至多种机器生成文本,提供基于量规、可解释的自动化评估方法。不同于单一评分方式,CRACQ融合语言、语义与结构信号,实现整体与分项维度的综合判断。该模型在500份合成资助提案上训练,并与大模型作为裁判(LLM-as-a-judge)进行对比,进一步在真实强弱文本上测试。初步结果显示,相较于直接使用大模型评估,CRACQ在各维度上的判断更具稳定性与可解释性,但可靠性与领域适应性仍存挑战。

原文摘要 · Abstract (English)

This paper presents CRACQ, a multi-dimensional evaluation framework tailored to evaluate documents across f i v e specific traits: Coherence, Rigor, Appropriateness, Completeness, and Quality. Building on insights from traitbased Automated Essay Scoring (AES), CRACQ expands its fo-cus beyond essays to encompass diverse forms of machine-generated text, providing a rubricdriven and interpretable methodology for automated evaluation. Unlike singlescore approaches, CRACQ integrates linguistic, semantic, and structural signals into a cumulative assessment, enabling both holistic and trait-level analysis. Trained on 500 synthetic grant pro-posals, CRACQ was benchmarked against an LLM-as-a-judge and further tested on both strong and weak real applications. Preliminary results in-dicate that CRACQ produces more stable and interpretable trait-level judgments than direct LLM evaluation, though challenges in reliability and domain scope remain

文本评估多维度可解释性自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。