构建以人工评估为核心的生成文本偏见检测框架,解决真实场景下偏见评估难题。
Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text
- 基于人工洞察定义偏见,实现可自动化处理的评估流程。
- 突破多选题限制,提出自由文本偏见分类方法。
- 发现基准数据集中的问题模板,提升评估可靠性。
大语言模型评估本就复杂,实际部署中更因任务提示和经验背景的交互而加剧难度。大规模偏见评估常依赖短上下文、固定选项的基准,虽易快速评估,但当模型部署环境变化时,其有效性可能丧失。大规模人工评估常被视为不可行且成本过高。本文呈现我们开发一个以人工洞察为核心、支持自由文本响应偏见评估的半自动化框架。我们讨论如何制定可操作的偏见定义以推动流程自动化,并提出超越多选题的偏见分类方法。此外,我们指出人工评估如何帮助发现偏见基准中的问题模板。
原文摘要 · Abstract (English)
LLM evaluation is challenging even the case of base models. In real world deployments, evaluation is further complicated by the interplay of task specific prompts and experiential context. At scale, bias evaluation is often based on short context, fixed choice benchmarks that can be rapidly evaluated, however, these can lose validity when the LLMs' deployed context differs. Large scale human evaluation is often seen as too intractable and costly. Here we present our journey towards developing a semi-automated bias evaluation framework for free text responses that has human insights at its core. We discuss how we developed an operational definition of bias that helped us automate our pipeline and a methodology for classifying bias beyond multiple choice. We additionally comment on how human evaluation helped us uncover problematic templates in a bias benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。