arXiv:2605.03900cs.AI2026-05

让大模型在复杂场景中智能权衡多个目标,避免盲目追求单一指标。

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems

  • 基于上下文动态选择并协调多个目标,如帮助性、真实性、安全性等。
  • 提出可分解的目标表示与层级约束机制,支持多目标冲突处理。
  • 适合高风险决策、个性化服务等需要综合考量的AI应用领域。

前沿AI系统在代码生成、数学推理、游戏和单元测试驱动任务等目标明确、稳定且可验证的场景中表现优异,但在科学辅助、长周期代理、高风险建议、个性化和工具使用等开放性场景中仍不可靠,因其面临目标模糊、依赖上下文、延迟显现或部分可观测等问题。我们指出,这些问题不仅是规模或能力不足所致,更是目标选择失败:系统优化了局部可见信号,却忽略了应主导交互的核心目标。为此,我们提出「上下文多目标优化」框架,要求系统在不同情境下同时考虑多种目标,如帮助性、真实性、安全性、隐私保护、校准性、非操纵性、用户偏好、可逆性及利益相关方影响,并判断哪些是活跃目标、软性偏好或硬性/准硬性约束。这些目标并非穷尽清单,不同领域和部署场景可能激活不同维度及冲突解决策略。该框架将AI行为建模为对候选动作、目标估计、活跃约束、利益相关方、不确定性及冲突解决方式的上下文依赖选择规则。我们进一步提出实现路径:包括目标的解耦表示、上下文到目标的路由机制、层级约束、反思式策略推理、可控个性化、工具使用控制、诊断评估、审计与部署后修订。

原文摘要 · Abstract (English)

Frontier AI systems perform best in settings with clear, stable, and verifiable objectives, such as code generation, mathematical reasoning, games, and unit-test-driven tasks. They remain less reliable in open-ended settings, including scientific assistance, long-horizon agents, high-stakes advice, personalization, and tool use, where the relevant objective is ambiguous, context-dependent, delayed, or only partially observable. We argue that many such failures are not merely failures of scale or capability, but failures of objective selection: the system optimizes a locally visible signal while missing which objectives should govern the interaction. We formulate this problem as \emph{contextual multi-objective optimization}. In this setting, systems must consider multiple, context-dependent objectives, such as helpfulness, truthfulness, safety, privacy, calibration, non-manipulation, user preference, reversibility, and stakeholder impact, while determining which objectives are active, which are soft preferences, and which must function as hard or quasi-hard constraints. These examples are not intended as an exhaustive taxonomy: different domains and deployment settings may activate different objective dimensions and different conflict-resolution procedures. Our framework models AI behavior as a context-dependent choice rule over candidate actions, objective estimates, active constraints, stakeholders, uncertainty, and conflict-resolution procedures. We outline an implementation pathway based on decomposed objective representations, context-to-objective routing, hierarchical constraints, deliberative policy reasoning, controlled personalization, tool-use control, diagnostic evaluation, auditing, and post-deployment revision.

多目标优化上下文感知大模型安全目标对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。