提出新评估方法,让模型评价更贴近真实使用场景。
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
- 设计可扩展的人类评估流程,模拟实际使用中的分类任务。
- 验证大模型代理可媲美人类标注,统计上无显著差异。
- 适合需要真实反馈的模型评估者,尤其关注实用性。
主题模型与文档聚类的评估通常依赖与人类偏好不符的自动指标,或需难以扩展的专家标注。我们设计了一种可扩展的人类评估协议及相应的自动化替代方案,反映从业者在真实使用中的模型应用方式。标注者(或基于LLM的代理)审查分配给某一主题或聚类的文本,推断该组类别,并将此类别应用于其他文档。利用该协议,我们在两个数据集上收集了多种主题模型输出的大量众包标注。随后用这些标注验证自动化代理,发现最佳的LLM代理与人工标注者在统计上无法区分,因此可在自动化评估中作为合理替代。代码、网页界面和数据见https://github.com/ahoho/proxann。
原文摘要 · Abstract (English)
Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding automated approximation that reflect practitioners' real-world usage of models. Annotators -- or an LLM-based proxy -- review text items assigned to a topic or cluster, infer a category for the group, then apply that category to other documents. Using this protocol, we collect extensive crowdworker annotations of outputs from a diverse set of topic models on two datasets. We then use these annotations to validate automated proxies, finding that the best LLM proxies are statistically indistinguishable from a human annotator and can therefore serve as a reasonable substitute in automated evaluations. Package, web interface, and data are at https://github.com/ahoho/proxann
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。