检验仇恨言论模型是否真按预设定义运行
DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?
- 用三步流程将用户定义的仇恨言论标准注入模型
- 在六个主流数据集上发现模型行为与定义存在偏差
- 可定位训练流程中导致偏差的关键环节
构建预测模型时,常难以确保部署模型的行为符合应用需求。以仇恨言论检测为例,研究者虽有明确的定义,但数据构建与模型训练过程中的采样偏差、标注偏差和模型误设,可能导致实际行为与预期不符。为解决此问题,本文提出DefVerify:一个三步流程,首先编码用户指定的仇恨言论定义,其次量化模型对定义的反映程度,最后识别工作流中失效环节。该方法应用于六个主流仇恨言论基准数据集,揭示了定义与模型行为间的显著差距。
原文摘要 · Abstract (English)
When building a predictive model, it is often difficult to ensure that application-specific requirements are encoded by the model that will eventually be deployed. Consider researchers working on hate speech detection. They will have an idea of what is considered hate speech, but building a model that reflects their view accurately requires preserving those ideals throughout the workflow of data set construction and model training. Complications such as sampling bias, annotation bias, and model misspecification almost always arise, possibly resulting in a gap between the application specification and the model's actual behavior upon deployment. To address this issue for hate speech detection, we propose DefVerify: a 3-step procedure that (i) encodes a user-specified definition of hate speech, (ii) quantifies to what extent the model reflects the intended definition, and (iii) tries to identify the point of failure in the workflow. We use DefVerify to find gaps between definition and model behavior when applied to six popular hate speech benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。