arXiv:2410.01957cs.CL2024-10被引 9

提出数据驱动的AI对齐新方向,强调高质量数据的关键作用。

Challenges and Future Directions of Data-Centric AI Alignment

  • 聚焦数据质量与代表性,推动从算法主导转向数据驱动的对齐方法。
  • 发现人类反馈存在不可靠性、时间漂移和情境依赖等多重问题。
  • 适合关注AI伦理、数据治理的研究者与实践者参考。

随着AI系统能力增强与影响力扩大,确保其与人类价值观、偏好和目标对齐已成为关键研究课题。当前对齐方法多集中于算法与损失函数设计,却常低估数据的核心作用。本文倡导转向数据中心的对齐范式,强调提升对齐过程中所用数据的质量与代表性。通过定性分析,本文揭示了基于人类与AI反馈在数据中心对齐框架下的多重挑战:人类反馈存在可靠性不足、随时间演变、受上下文影响等问题;而基于AI的反馈因模型自身局限,难以准确捕捉人类价值观。为此,本文提出未来研究方向,包括改进反馈收集流程、发展稳健的数据清洗方法以及建立严格的反馈验证机制,呼吁针对这些关键领域展开深入研究,以弥补当前在数据中心对齐实践中的理解与改进空白。

原文摘要 · Abstract (English)

As AI systems become increasingly capable and influential, ensuring their alignment with human values, preferences, and goals has become a critical research focus. Current alignment methods primarily focus on designing algorithms and loss functions but often underestimate the crucial role of data. This paper advocates for a shift towards data-centric AI alignment, emphasizing the need to enhance the quality and representativeness of data used in aligning AI systems. In this position paper, we highlight key challenges associated with both human-based and AI-based feedback within the data-centric alignment framework. Through qualitative analysis, we identify multiple sources of unreliability in human feedback, as well as problems related to temporal drift, context dependence, and AI-based feedback failing to capture human values due to inherent model limitations. We propose future research directions, including improved feedback collection practices, robust data-cleaning methodologies, and rigorous feedback verification processes. We call for future research into these critical directions to ensure, addressing gaps that persist in understanding and improving data-centric alignment practices.

AI对齐数据质量人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。