用AI辅助梳理生成式AI评估中的模糊概念,让评价更清晰可测。
AI-Assisted Systematization for Evaluating GenAI Systems

- 提出概念规格书(concept spec)和验证表,将抽象概念转为可测量条目
- 用AI生成的两个概念规格在内容有效性和信息可恢复性上表现良好
- 适合从事生成式AI评估、评测框架设计的研究者和工程师参考
评估生成式人工智能(GenAI)系统面临挑战,因其评估目标如“推理”“公平性”“创造力”等多为宽泛且存在争议的概念。若这些概念未明确定义,就难以明确测量什么或如何解读结果。这反映出一个缺失环节:系统化——即从宽泛背景概念过渡到可测量的结构化表述。为缓解系统化过程认知负担重、资源消耗大的问题,本文探究是否可用AI辅助完成该过程。为此,我们提出一种结构化概念表示形式——概念规格书(concept spec),以及配套验证工作表。进而开发两种AI辅助系统化方法:直接零样本法与模拟人工流程的多智能体法。利用这两种方法对“基于仇恨的修辞”和“数字共情”两个概念生成概念规格书,并在内容有效性与信息可恢复性方面进行评估,结果表明生成结果具备良好质量。
原文摘要 · Abstract (English)
Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as "reasoning," "fairness," or "creativity." When these concepts are left underspecified, it becomes unclear what should be measured or how evaluation results should be interpreted. This problem reflects a missing step: systematization, that is, moving from a broad background concept to an explicit, structured account of the concept in measurable terms. To help address the fact that systematization is cognitively demanding and resource-intensive, we investigate whether AI assistance can support this process. To enable AI-assisted systematization and assess its quality, we introduce a structured representation of a systematized concept, a concept spec, and a validation worksheet. We then develop two AI-assisted systematizers: a direct, zero-shot approach and a multi-agent approach that more closely mirrors manual systematization approaches from existing literature. We use these systematizers to produce concept specs for two concepts -- hate-based rhetoric and digital empathy -- and evaluate resulting concept specs on content validity and information recoverability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。