调研12位从业者,发现现有模型偏见评估工具难用
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
- 通过访谈发现评估工具与实际需求错位
- 多数工具因不匹配需求或使用障碍无法落地
- 建议从测量理论出发优化工具设计与推广
自然语言处理领域已公开发布多种用于衡量大语言模型系统所造成表征危害的工具,包括数据集、度量指标、软件工具等。本文通过对12位负责评估大语言模型系统的从业者进行半结构化访谈,探究现有工具在满足实践需求方面的有效性。研究发现,从业者往往无法有效使用这些公开工具。我们识别出两类挑战:一是部分工具本身与从业者关注的测量目标不一致,存在概念错位;二是即便工具有效,也常因实际操作和组织层面的障碍而难以采纳。基于测量理论与实用测量原则,本文提出改进建议,以更好地契合从业者的真实需求。
原文摘要 · Abstract (English)
The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instruments have taken the form of datasets, metrics, tools, and more. In this paper, we examine the extent to which such instruments meet the needs of practitioners tasked with evaluating LLM-based systems. Via semi-structured interviews with 12 such practitioners, we find that practitioners are often unable to use publicly available instruments for measuring representational harms. We identify two types of challenges. In some cases, instruments are not useful because they do not meaningfully measure what practitioners seek to measure or are otherwise misaligned with practitioner needs. In other cases, instruments - even useful instruments - are not used by practitioners due to practical and institutional barriers impeding their uptake. Drawing on measurement theory and pragmatic measurement, we provide recommendations for addressing these challenges to better meet practitioner needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。