用六边形框架综合评估学者,兼顾研究质量与可验证行为。
HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment

- 分内在与外在两层证据,从多维度评估学者
- 六维度评估中与人工标准一致率提升,外部数据显轨迹信号
- 结果可解释可审计,适合学术评审与人才选拔
学者评估在教师招聘、经费分配、职称晋升和人才发现中具有基础作用。现有方法多依赖引文指标和声誉代理,而近期基于大语言模型的方法主要聚焦单篇论文评价,未能全面评估学者。我们提出HexEval——一种以证据驱动的六边形框架,将学者评估视为需结合内在研究质量与外部可验证学术行为的推理问题。该框架包含两个互补的证据层:内在层对匿名代表性成果从研究严谨性、方法创新性和科学贡献三个维度进行评估;外在层则通过来自GitHub、Lens、OpenAlex等公开可验证来源的异构数据,刻画知识转化、研究连贯性和学术影响力。不同于生成模糊总分,HexEval保留中间证据、各维度理由和验证信号,实现可解释、可审计的学者画像。跨六个维度的实验表明,结构化校准提升了内在质量评估的绝对一致性,外部模块有效恢复了整体发展轨迹和等级影响信号。结果支持基于异构学术证据的证据驱动推理作为可审计的AI辅助评估范式,同时揭示了公共学术数据在覆盖范围与归属准确性上的局限。
原文摘要 · Abstract (English)
Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing scholar assessment methods predominantly rely on bibliometric indicators and reputation proxies, while recent large language model (LLM)-based approaches mainly focus on evaluating individual research papers rather than comprehensively assessing scholars. We argue that scholar assessment should be formulated as an evidence-driven reasoning problem that jointly considers intrinsic research quality and externally verifiable scholarly behavior. To this end, we propose HexEval, an evidence-driven hexagonal framework for multidimensional scholar assessment. HexEval explicitly organizes scholar assessment into two complementary evidence layers. The intrinsic layer evaluates anonymized representative works along three dimensions, namely research rigor, methodological innovation, and scientific contribution, whereas the external layer characterizes scholars through knowledge translation, research coherence, and academic impact using heterogeneous evidence collected from GitHub, Lens, OpenAlex, and other publicly verifiable sources. Instead of producing opaque aggregate scores, HexEval preserves intermediate evidence, dimension-specific rationales, and verification signals throughout the evaluation process, enabling interpretable and auditable scholar profiles. Experiments across all six dimensions show dimension-dependent agreement with human or external reference criteria: structured calibration improves absolute agreement for intrinsic quality, while the external modules recover broad trajectory and ordinal impact signals. These results support evidence-driven reasoning over heterogeneous scholarly evidence as a promising paradigm for auditable AI-assisted scholar assessment, while exposing the coverage and attribution limitations of public scholarly data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。