用拆解法提升人类对齐大模型的反馈质量,让判断更准更快。
DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
- 将长文本拆成独立观点逐条对比,降低认知负担。
- 实验显示反馈准确率平均提升5%,尤其在不确定时效果显著。
- 适合需要高质量人工标注的AI对齐研究者使用。
人类偏好广泛用于通过强化学习从人类反馈(RLHF)对齐大语言模型。然而,现有用户界面要求标注者比较长段落文本,当内容冗长或陌生时,认知负担大。本文提出分解原则作为改进人类反馈质量的方法:将文本拆分为独立论点,而非直接比较整段回复。基于此,我们构建了新界面DxHF,通过展示拆解后的论点、视觉化呈现论点与对话的相关性并链接相似论点,帮助用户快速浏览关键信息,识别差异以做出更优判断。技术评估表明,分解方法普遍提高反馈准确性,尤其在用户存在不确定性时。160名参与者参与的众包研究显示,使用DxHF使反馈准确率平均提升5%,但平均耗时增加18秒。研究结果凸显人机交互设计在提升人机对齐中的潜力。
原文摘要 · Abstract (English)
Human preferences are widely used to align large language models (LLMs) through methods such as reinforcement learning from human feedback (RLHF). However, the current user interfaces require annotators to compare text paragraphs, which is cognitively challenging when the texts are long or unfamiliar. This paper contributes by studying the decomposition principle as an approach to improving the quality of human feedback for LLM alignment. This approach breaks down the text into individual claims instead of directly comparing two long-form text responses. Based on the principle, we build a novel user interface DxHF. It enhances the comparison process by showing decomposed claims, visually encoding the relevance of claims to the conversation and linking similar claims. This allows users to skim through key information and identify differences for better and quicker judgment. Our technical evaluation shows evidence that decomposition generally improves feedback accuracy regarding the ground truth, particularly for users with uncertainty. A crowdsourcing study with 160 participants indicates that using DxHF improves feedback accuracy by an average of 5%, although it increases the average feedback time by 18 seconds. Notably, accuracy is significantly higher in situations where users have less certainty. The finding of the study highlights the potential of HCI as an effective method for improving human-AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。