用专家知识动态评估AI在专业任务中的表现,兼顾稳定与灵活。
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
- 分两层评估:先用预设技能确保标准一致,再逐条分析推理过程。
- 在BizBench上提升评估稳定性,发现传统方法忽略的关键失败模式。
- 可迁移至医疗、金融等10个领域,贴合专家评分标准。
在开放性专业任务中评估智能体面临严谨性与灵活性的矛盾:静态评分标准虽可复现但难以适配多样策略,而基于大模型的评判方式则存在不稳定和偏见问题。人类专家通过结合领域知识与动态的逐条评估来应对这一挑战。受此启发,我们提出JADE,一种双层评估框架。第一层将专家知识编码为预定义的评估技能集,确保评价标准稳定;第二层针对具体报告进行逐条评估,灵活判断不同推理路径,并通过证据依赖性控制机制,排除基于被证伪前提得出的结论。在BizBench上的实验表明,JADE显著提升了评估稳定性,并揭示了整体式大模型评估所遗漏的关键代理失效模式。进一步实验证明,JADE与专家制定的评分标准高度对齐,且能有效迁移至HealthBench和DR.BENCH,覆盖医疗及十大专业领域。代码与数据已开源。
原文摘要 · Abstract (English)
Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail to accommodate diverse valid response strategies, while LLM-as-a-judge approaches adapt to individual responses yet suffer from instability and bias. Human experts address this dilemma by combining domain-grounded principles with dynamic, claim-level assessment. Inspired by this process, we propose JADE, a two-layer evaluation framework. Layer 1 encodes expert knowledge as a predefined set of evaluation skills, providing stable evaluation criteria. Layer 2 performs report-specific, claim-level evaluation to flexibly assess diverse reasoning strategies, with evidence-dependency gating to invalidate conclusions built on refuted claims. Experiments on BizBench show that JADE improves evaluation stability and reveals critical agent failure modes missed by holistic LLM-based evaluators. We further demonstrate strong alignment with expert-authored rubrics and effective transfer to HealthBench and DR.BENCH, covering medical and 10-domain professional evaluation settings. Code and data are available at https://github.com/smiling-world/JADE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。