构建复杂动作空间的条件决策评估基准,测试大模型真实场景下的决策能力。
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space

- 动作由变量分配构成,受变量、上下文和分配层级的显式条件约束
- 采用基于预言机的评估方式,同时衡量决策质量与条件遵守程度
- 适用于高风险领域中大模型决策支持能力的真实测评
大型语言模型因其上下文理解与推理能力,被广泛探索为高风险领域的决策辅助工具。然而,现有决策评估基准依赖两个简化假设:动作从预定义有限集合中选择,且未将限制动作可行性的显式条件纳入决策过程。这些假设无法捕捉现实世界动作的组合结构及对其有效性起约束作用的显式条件。为解决上述局限,我们提出 CONDESION-BENCH,一个用于评估大模型在组合动作空间中条件决策能力的基准。在 CONDESION-BENCH 中,动作被定义为对决策变量的分配,并受到变量级、上下文级和分配级的显式条件约束。通过基于预言机的评估方法,同时评估决策质量与条件遵守情况,我们提供了对大模型作为决策辅助工具更严格的评估标准。
原文摘要 · Abstract (English)
Large language models have been widely explored as decision-support tools in high-stakes domains due to their contextual understanding and reasoning capabilities. However, existing decision-making benchmarks rely on two simplifying assumptions: actions are selected from a finite set of pre-defined candidates, and explicit conditions restricting action feasibility are not incorporated into the decision-making process. These assumptions fail to capture the compositional structure of real-world actions and the explicit conditions that constrain their validity. To address these limitations, we introduce CONDESION-BENCH, a benchmark designed to evaluate conditional decision-making in compositional action space. In CONDESION-BENCH, actions are defined as allocations to decision variables and are restricted by explicit conditions at the variable, contextual, and allocation levels. By employing oracle-based evaluation of both decision quality and condition adherence, we provide a more rigorous assessment of LLMs as decision-support tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。