构建鲁棒的科学问答评估框架,解决大模型评价中的乐观偏差问题。
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
- 采用细粒度评分标准与强化学习结合的方法
- 在多领域科学问答数据集上实现无成本、可扩展评估
- 适合需要可靠模型评估的AI研究与科学推理场景
大型语言模型(LLMs)驱动现代搜索引擎的科学问答,但其评估鲁棒性尚未充分探索。我们提出YESciEval,一个开源框架,通过细粒度评分标准与强化学习结合,缓解LLM评估者中的乐观偏差。我们发布了多学科科学问答数据集,包含对抗性变体,并提供多个LLM的评估分数。该方法独立于专有模型和人工反馈,实现可扩展、零成本评估。本工作推进了可靠的LLM-as-a-judge模型,支持人工智能对齐,促进科学探究所需的稳健、透明评估。
原文摘要 · Abstract (English)
Large Language Models (LLMs) drive scientific question-answering on modern search engines, yet their evaluation robustness remains underexplored. We introduce YESciEval, an open-source framework that combines fine-grained rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators. We release multidisciplinary scienceQ&A datasets, including adversarial variants, with evaluation scores from multiple LLMs. Independent of proprietary models and human feedback, our approach enables scalable, cost-free evaluation. By advancing reliable LLM-as-a-judge models, this work supports AI alignment and fosters robust, transparent evaluation essential for scientific inquiry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。