让大模型像人一样做决策,优先满足核心目标并守住底线。
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
- 推理时采用满意策略,主目标优化+次目标阈值约束
- 在PKU-SafeRLHF上帮助性得分领先22.3%,同时确保无害性达标
- 适合追求安全与实用平衡的模型部署场景
由于偏好反馈具有固有的多维特性,将大语言模型与人类对齐极具挑战。现有方法通常将其视为多目标优化问题,却忽略了人类实际决策方式。研究表明,人类决策遵循‘满意’策略——在优化主要目标的同时,确保其他目标达到可接受阈值。为弥合这一差距并实现满意对齐,我们提出SITAlign:一种推理时框架,通过最大化主目标并满足次级标准的阈值约束来应对多维度对齐问题。我们从理论上推导了该满意推理对齐方法的次优性边界。在多个基准上的大量实验验证了SITAlign的有效性。例如,在以帮助性为主目标、无害性为阈值约束的PKU-SafeRLHF数据集上,SITAlign在帮助性奖励的GPT-4胜率平局率上比当前最优多目标解码策略高出22.3%,同时严格遵守无害性阈值。
原文摘要 · Abstract (English)
Aligning large language models with humans is challenging due to the inherently multifaceted nature of preference feedback. While existing approaches typically frame this as a multi-objective optimization problem, they often overlook how humans actually make decisions. Research on bounded rationality suggests that human decision making follows satisficing strategies-optimizing primary objectives while ensuring others meet acceptable thresholds. To bridge this gap and operationalize the notion of satisficing alignment, we propose SITAlign: an inference time framework that addresses the multifaceted nature of alignment by maximizing a primary objective while satisfying threshold-based constraints on secondary criteria. We provide theoretical insights by deriving sub-optimality bounds of our satisficing based inference alignment approach. We empirically validate SITAlign's performance through extensive experimentation on multiple benchmarks. For instance, on the PKU-SafeRLHF dataset with the primary objective of maximizing helpfulness while ensuring a threshold on harmlessness, SITAlign outperforms the state-of-the-art multi objective decoding strategy by a margin of 22.3% in terms of GPT-4 win-tie rate for helpfulness reward while adhering to the threshold on harmlessness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。