arXiv:2607.06940cs.CLcs.AI2026-07

提出多维度评分体系,更全面评估大模型回答质量。

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

  • 设计包含准确性、简洁性等五维度的评分框架
  • 在TruthfulQA上发现主流模型复合得分达0.6104
  • 适合关注模型真实性和可解释性的研究者

大语言模型在语言任务中的卓越表现凸显了对其回答质量进行全面评估的紧迫性。现有方法多局限于单一维度,难以捕捉模型能力的全貌。本研究提出一种多因素评分范式,融合准确性、简洁性、事实一致性、可读性和连贯性,并配备图形化界面以可视化评估结果。在TruthfulQA数据集上的评估显示,主流大模型在推理任务中表现突出(复合得分最高达0.6104),但在处理复杂事实和模糊信息时仍存在普遍局限。该框架突破传统指标的狭隘视角,为揭示模型潜力与缺陷提供了透明、可扩展的路径。尽管目前聚焦英文任务,其应用前景已延伸至多语言领域。本工作为知识工程与模型优化开辟了新路径。

原文摘要 · Abstract (English)

The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs' strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.

大模型评估多维度评分TruthfulQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。