arXiv:2507.17477cs.AI2025-07被引 1

无需人工标注,自动提升大模型对齐人类意图的能力

An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models

  • 通过多响应与不确定性评估,自动识别需改进的输出
  • 分三阶段优化模型,从保守到探索逐步提升对齐效果
  • 在安全性、真实性等任务上超越现有方法,适合训练新模型

大语言模型在指令遵循和通用推理方面已取得显著进展,但实现高质量的人类意图与安全规范对齐,且无需人工标注,仍是根本性挑战。本文提出一种不确定性驱动的自适应自对齐框架(UDASA),可完全自动化地提升模型对齐能力。该框架首先为每个输入生成多个响应,并从语义、事实性和价值对齐三个维度量化输出不确定性。基于不确定性得分,构建偏好对并按不确定性差异将训练样本分为保守、适度和探索三类,模型按阶段逐步优化。我们还进行了初步研究,验证核心设计假设并提供充分实证支持。实验结果表明,UDASA在危害性规避、帮助性、真实性及受控情感生成等多个任务上均优于现有对齐方法,显著提升模型表现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable progress in instruction following and general-purpose reasoning. However, achieving high-quality alignment with human intent and safety norms without human annotations remains a fundamental challenge. In this work, we propose an Uncertainty-Driven Adaptive Self-Alignment (UDASA) framework designed to improve LLM alignment in a fully automated manner. UDASA first generates multiple responses for each input and quantifies output uncertainty across three dimensions: semantics, factuality, and value alignment. Based on these uncertainty scores, the framework constructs preference pairs and categorizes training samples into three stages, conservative, moderate, and exploratory, according to their uncertainty difference. The model is then optimized progressively across these stages. In addition, we conduct a series of preliminary studies to validate the core design assumptions and provide strong empirical motivation for the proposed framework. Experimental results show that UDASA outperforms existing alignment methods across multiple tasks, including harmlessness, helpfulness, truthfulness, and controlled sentiment generation, significantly improving model performance.

大模型对齐自适应优化不确定性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。