提出兼顾多元观点的可复现机器学习评估框架
Thesis Proposal: Toward a Human-Centered and Perspective-Aware Framework for Reproducible ML Evaluation and AI Alignment

- 构建考虑人类分歧的评估体系,避免单一共识掩盖少数意见
- 通过多视角反馈提升模型评价的透明度与可复现性
- 适合关注AI对齐、内容审核等主观性强领域的研究者
人类在人工智能开发的各个阶段——从数据收集、整理到模型训练与评估——都发挥着关键作用。然而,人们常常存在意见分歧,甚至随时间推移自我矛盾。在人工智能安全、内容审核或情感分析等主观性较强的领域,这种分歧尤为显著。当前的大语言模型评估方法普遍依赖多数投票聚合标签以代表共识,从而掩盖了少数观点。忽视人类分歧会加剧人工智能的可复现性危机。同时,人类反馈对确保AI系统与人类价值观对齐至关重要。为建立可信的AI系统,必须使其反映多样化的价值与视角。本论文提案提出一个以人为本、具备视角感知能力的可复现机器学习评估与对齐框架。
原文摘要 · Abstract (English)
Humans play a vital role at every stage of AI development, from data collection and curation to model development and evaluation. However, humans often disagree with each other and sometimes with themselves over time. It is essential to take disagreement into account when building human-centered AI systems, especially in domains where it is prevalent, such as AI safety, content moderation, or sentiment analysis. Disagreement often arises from subjective human opinion and can vary with one's identity, beliefs, and social environment. Despite this, current LLM evaluation approaches frequently rely on aggregating labels (often via plurality voting) to represent consensus, thereby obscuring minority perspectives. By failing to account for human disagreement, these evaluation methods contribute to the reproducibility crisis in AI. Human feedback is also crucial for ensuring that AI systems align with human values. For these systems to be trustworthy, it is critical to ensure that they reflect diverse human values and perspectives. In this thesis proposal, we present a human-centered and perspective-aware framework for reproducible ML evaluation and AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。