arXiv:2607.09706cs.AIcs.GT2026-07

让语言模型做出更稳健的决策,通过不确定性建模避免盲目依赖单一数值。

YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate

论文配图:YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate
图 1 · 摘自论文原文
  • 用带不确定性的命题图表示问题,融合结构先验与假设溯源
  • 在41,188条真实决策上,回测表现优于现状34%,减少优化器诅咒
  • 提供可审计的决策追溯和后悔值边界,适合高风险决策场景

语言模型将自然语言情境转化为数值计划,现有主流方法(NL4Opt、OptiMUS、ORLM、OR-LLM-Agent)通常固定单一目标并采用点估计系数,仅求解一次。但涉及真实预算、资源或临床关注的决策中,这种确定性假设极易失效:每个数值都是潜在假设,仅在假设完全正确时最优的方案本质上脆弱且类比计算。YUKTI 改变了自动形式化的目标,其表示为带有类型命题的图结构,关系携带形状先验、系数不确定性及来源信息。该框架在各阶段使用精确、非线性或进化求解器;通过分布式帕累托传递耦合阶段;引入假设鲁棒帕累托前沿(ARPF),对假设(包括结构性epsilon污染)进行重采样,评估每项行动的存活频率(rho)。我们证明了rho是决策后悔的精确因子,并增强了可审计追溯性;在无基准数据时构建了符合实际的测试数据集(SRJANA)。验证表明:在受控误设下,鲁棒妥协方案使均值与尾部后悔降低超90%;在受监管商业决策中,优化在合法空间内进行并以欧元量化下行风险;在包含41,188条真实决策的公共数据集上,样本外回测优于现有做法34%,优于朴素点规则4%,同时缓解优化器诅咒。求解器为标准工具,不宣称基准领先。对比实验显示,即使给定正确数值并单目标优化,大模型仍产生约47倍于YUKTI的持有期后悔——大模型是形式化者,而非求解者。当存在长程因果耦合时,前向传递失效,需转为后向归纳因果策略。

原文摘要 · Abstract (English)

Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-valued coefficients, then solve once. For decisions that allocate real budget, effort, or clinical attention, that confidence is the failure mode: every objectified number is an assumption, and a plan optimal only if the guesses are exactly right is fragile -- mimicry of computation. YUKTI changes the target of autoformulation. Its representation is a typed-proposition graph whose relationships carry shape priors, coefficient uncertainty, and provenance. YUKTI routes each stage to an exact, nonlinear, or evolutionary solver; couples stages by a distributional Pareto hand-off; and introduces Assumption-Robust Pareto Frontiers (ARPF), resampling assumptions (including structural epsilon-contamination) to score how often each action survives (rho). We prove a bound making rho an exact factor of decision regret, add auditable traceability, and synthesize a benchmark-faithful data foundation when none exists (SRJANA). We validate three ways: under controlled misspecification the robust compromise cuts mean and tail regret by over 90% versus a naive point plan; on a regulated commercial decision we optimize inside a lawful action space and price the downside in euros; and on a real public dataset of 41,188 decisions an out-of-sample backtest beats the logged status quo by 34% and a naive point rule by 4% while reducing the optimizer's curse. The solvers are standard; we claim no benchmark-SOTA win. A head-to-head shows an LLM given the correct numbers, and single-objective optimization, both incur about 47x the held-out regret of YUKTI -- an LLM is a formulator, not a solver. Under long-range causal coupling, the forward hand-off becomes unsound, locating where it must become a backward-induction causal policy.

决策系统不确定性建模鲁棒优化语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。