arXiv:2508.08277cs.CLcs.LG2025-08被引 3

用外部数据源构建客观评估大模型的新框架,避免主观判断。

Objective Metrics for Evaluating Large Language Models Using External Data Sources

  • 基于多学期文本材料构建主观指标,反向评估大模型输出
  • 通过基准测试与结构化流程实现可复现、低偏差的评分
  • 适合教育、科研等需高可靠性评估的场景

评估大语言模型性能是一项关键但具挑战性的任务,尤其在避免主观评价方面。本文提出一种利用不同学期课程文本材料中提取的主观指标,来评估大模型在各类任务中的表现。通过使用明确的基准测试、事实性数据集和结构化评估流程,该方法确保了测量的一致性、可复现性以及低偏差。框架强调评分过程的自动化与透明化,减少对人工解读的依赖,同时保持与真实应用场景的一致性。该方法克服了传统主观评估的局限,为教育、科学及其他高风险领域提供了可扩展的性能评估方案。

原文摘要 · Abstract (English)

Evaluating the performance of Large Language Models (LLMs) is a critical yet challenging task, particularly when aiming to avoid subjective assessments. This paper proposes a framework for leveraging subjective metrics derived from the class textual materials across different semesters to assess LLM outputs across various tasks. By utilizing well-defined benchmarks, factual datasets, and structured evaluation pipelines, the approach ensures consistent, reproducible, and bias-minimized measurements. The framework emphasizes automation and transparency in scoring, reducing reliance on human interpretation while ensuring alignment with real-world applications. This method addresses the limitations of subjective evaluation methods, providing a scalable solution for performance assessment in educational, scientific, and other high-stakes domains.

大模型评估客观指标外部数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。