arXiv:2508.18646cs.AIcs.CL2025-08被引 2

用四维能力模型诊断大模型短板,让评估从分数竞争转向根源分析。

Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap

  • 构建智力、专业、情感与价值观四维评估框架,映射训练全流程。
  • 分析200+基准测试趋势,揭示技术指标与真实应用间的脱节问题。
  • 适合关注模型可解释性、伦理对齐及长期部署的开发者与研究者。

尽管大语言模型快速演进,其评估仍存在基准分数与实际应用价值严重脱节的问题。现有评估体系碎片化,过度关注孤立技术指标,忽视模型发展全周期与社会影响等关键维度。本文提出一种具身化评估框架,将模型能力重新定义为四个维度:智商(IQ)、职业商(PQ)、情商(EQ)与价值观商(VQ),并建立模块化评估架构。通过元分析200多个公开基准测试,验证该框架在诊断模型缺陷方面的有效性。研究揭示了当前评估体系的核心挑战,并指明未来发展方向。该工作为打造技术可靠、情境适配且伦理合规的大模型提供战略指引。相关资源已开源:https://github.com/onejune2018/Awesome-LLM-Eval。

原文摘要 · Abstract (English)

Despite their rapid advancement, large language models (LLMs) suffer from a critical disconnect between benchmark scores and real-world utility. Current evaluation remains fragmented, prioritizing isolated technical metrics over the holistic, developmental, and societal aspects essential for deployment. Rather than serving merely as a descriptive catalog, this work establishes a diagnostic ontology that causally maps evaluation dimensions to the canonical LLM training pipeline, transforming evaluation from static ranking into a diagnostic tool for root-cause analysis. In this paper, we introduce an anthropomorphic evaluation framework that re-conceptualizes LLM capabilities through a four-dimensional lens: Intelligence Quotient (IQ), Professional Quotient (PQ), Emotional Quotient (EQ), and Value-oriented Quotient (VQ). We operationalize these concepts through a modular evaluation architecture and validate the framework's diagnostic claims through meta-analysis of public benchmark trends. Analyzing over 200 benchmarks, we synthesize key challenges and future directions. This work offers a strategic compass for developing LLMs that are not only technically proficient but also contextually relevant and ethically sound. A curated repository is available at: https://github.com/onejune2018/Awesome-LLM-Eval.

大模型评估框架设计伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。