让标签反映人类价值观,提升主观任务模型的对齐与准确性
Labels have Human Values: Value Calibration of Subjective Tasks
- 通过标注者理由、专家价值分类或社会文化特征聚类,识别不同人类价值观
- 在多个数据集上显著优于忽略价值观结构的基线模型,提升判断区分度和校准性
- 适合需要对齐人类多元价值的生成式AI、内容安全与偏好学习场景
构建面向主观任务的NLP系统需确保其与不同人类价值观对齐。本文提出多校准主观任务学习框架(MC-STL),通过标注者理由相似性、专家价值分类体系或标注者社会文化特征三种方式,将标注聚类为可识别的人类价值群体,并为每个群体学习特定嵌入以校准预测。我们在序数、二元及偏好学习等多种主观任务设置中验证了该方法,涵盖有毒聊天机器人对话、攻击性社交媒体内容及人类偏好对齐等多数据集。结果表明,MC-STL始终优于忽略标注潜在价值观结构的基线模型,在判别力、价值特异性校准及冲突感知指标上均有显著提升。
原文摘要 · Abstract (English)
Building NLP systems for subjective tasks requires one to ensure their alignment to contrasting human values. We propose the MultiCalibrated Subjective Task Learner framework (MC-STL), which clusters annotations into identifiable human value clusters by three approaches (similarity of annotator rationales, expert-value taxonomies or rater's sociocultural descriptors) and calibrates predictions for each value cluster by learning cluster-specific embeddings. We demonstrate MC-STL on several subjective learning settings, including ordinal, binary, and preference learning predictions, and evaluate it on multiple datasets covering toxic chatbot conversations, offensive social media posts, and human preference alignment. The results show that MC-STL consistently outperforms the baselines that ignore the latent value structure of the annotations, delivering gains in discrimination, value-specific calibration, and disagreement-aware metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。