构建零售领域政策模糊下的决策评估基准,揭示大模型在真实场景中的分歧
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

- 基于真实零售政策模糊性设计多轮对话任务
- 顶尖模型对同一问题产生根本性分歧,验证模糊性挑战
- 适合评估大模型在现实业务中的决策一致性与合规性
基于大语言模型的智能体正被广泛应用于现实领域中例行但关键的任务,其行为受制于本质上模糊的行业政策,这些政策存在多种合理解释。然而现有评估基准大多假设政策清晰明确,导致实际评估存在显著空白。本文提出DRIP-R,一个系统性利用真实零售政策模糊性的基准,构建出无唯一正确解的决策场景。该基准包含经筛选的政策模糊退货案例与真实客户画像,支持全双工对话及工具调用,并配备多评委评估框架,涵盖政策遵循度、对话质量、行为一致性与解决质量。实验表明,前沿模型在相同模糊场景下存在根本性分歧,证实模糊性是大模型决策面临的真正且系统性挑战。
原文摘要 · Abstract (English)
LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations. Despite the prevalence of such ambiguities in practice, existing agent benchmarks largely assume unambiguous, well-specified policies, leaving a critical evaluation gap. We introduce DRIP-R, a benchmark that systematically exploits real-world retail policy ambiguities to construct scenarios in which no single correct resolution exists. DRIP-R comprises a curated set of policy-ambiguous return scenarios paired with a realistic customer personas, a full-duplex conversational simulation with tool-calling capabilities and a multi-judge evaluation framework covering policy adherence, dialogue quality, behavioral alignment, and resolution quality. Our experiments show that frontier models fundamentally disagree on identical policy-ambiguous scenarios, confirming that ambiguity poses a genuine and systematic challenge to LLM decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。