arXiv:2601.01825cs.CL2026-01

构建供应链推理评测基准,揭示大模型在商品规则上的短板。

CSCBench: A PVC Diagnostic Benchmark for Commodity Supply Chain Reasoning

  • 基于流程、品类、认知三维度设计评测框架
  • 大模型在流程与认知上表现好,品类规则理解差
  • 特别暴露货运协议等复杂规则的缺陷,适合工业界评估

大语言模型在通用基准上表现优异,但在受制度规则和可行性约束的大宗商品供应链(CSC)领域能力仍不明确。供应链决策涉及流程阶段(如规划、采购、交付)、品类特有规则(如合同条款、交货等级)及推理深度(从检索到多步分析与决策选择)。我们提出CSCBench,一个包含2300+单选题的供应链推理评测集,基于PVC三维评估框架(流程、品类、认知)。流程轴对齐SCOR+Enable;品类轴在真实物资-信息-金融耦合约束下,依据权威交易所指南与行业报告建模;认知轴遵循布卢姆修订版分类法。在直接提示设置下评估主流LLM,发现其在流程与认知轴表现良好,但在品类轴显著退化,尤其在货运协议任务上。CSCBench为诊断和提升大模型在高风险供应链领域的推理能力提供精准标尺。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success in general benchmarks, yet their competence in commodity supply chains (CSCs) -- a domain governed by institutional rule systems and feasibility constraints -- remains under-explored. CSC decisions are shaped jointly by process stages (e.g., planning, procurement, delivery), variety-specific rules (e.g., contract specifications and delivery grades), and reasoning depth (from retrieval to multi-step analysis and decision selection). We introduce CSCBench, a 2.3K+ single-choice benchmark for CSC reasoning, instantiated through our PVC 3D Evaluation Framework (Process, Variety, and Cognition). The Process axis aligns tasks with SCOR+Enable; the Variety axis operationalizes commodity-specific rule systems under coupled material-information-financial constraints, grounded in authoritative exchange guidebooks/rulebooks and industry reports; and the Cognition axis follows Bloom's revised taxonomy. Evaluating representative LLMs under a direct prompting setting, we observe strong performance on the Process and Cognition axes but substantial degradation on the Variety axis, especially on Freight Agreements. CSCBench provides a diagnostic yardstick for measuring and improving LLM capabilities in this high-stakes domain.

供应链大模型评测规则推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。