arXiv:2601.00121cs.AIcs.HC2026-01被引 3

让大模型当智能助手,帮中小企业优化库存,成本降三成。

Ask, Clarify, Optimize: Human-LLM Agent Collaboration for Smarter Inventory Control

  • 大模型只负责理解自然语言和解释结果,计算由专业算法完成。
  • 相比直接用大模型求解,总库存成本降低32.1%。
  • 适合不懂运筹学的管理者使用,提升决策效率。

库存管理对缺乏专业知识的小中型企业仍是挑战。本文研究大语言模型(LLMs)是否能弥补这一差距。结果显示,将大模型作为端到端求解器会因‘幻觉’产生显著性能损失,源于其无法进行基于事实的随机推理。为此,我们提出一种混合代理框架,严格分离语义推理与数学计算:大模型充当智能接口,从自然语言中提取参数并解读结果,同时自动调用严谨算法构建优化引擎。为评估该交互系统在真实管理对话中的模糊性与不一致性,我们引入‘人类模仿者’——一个微调后的‘数字孪生’经理,实现可扩展、可复现的压力测试。实证分析表明,该混合框架相比以GPT-4o为端到端求解器的基线,使总库存成本降低32.1%。此外,仅提供完美真实信息仍无法提升GPT-4o性能,证实瓶颈本质是计算而非信息。研究结论指出,大模型并非运筹学替代品,而是让非专家也能访问严谨求解策略的自然语言接口。

原文摘要 · Abstract (English)

Inventory management remains a challenge for many small and medium-sized businesses that lack the expertise to deploy advanced optimization methods. This paper investigates whether Large Language Models (LLMs) can help bridge this gap. We show that employing LLMs as direct, end-to-end solvers incurs a significant "hallucination tax": a performance gap arising from the model's inability to perform grounded stochastic reasoning. To address this, we propose a hybrid agentic framework that strictly decouples semantic reasoning from mathematical calculation. In this architecture, the LLM functions as an intelligent interface, eliciting parameters from natural language and interpreting results while automatically calling rigorous algorithms to build the optimization engine. To evaluate this interactive system against the ambiguity and inconsistency of real-world managerial dialogue, we introduce the Human Imitator, a fine-tuned "digital twin" of a boundedly rational manager that enables scalable, reproducible stress-testing. Our empirical analysis reveals that the hybrid agentic framework reduces total inventory costs by 32.1% relative to an interactive baseline using GPT-4o as an end-to-end solver. Moreover, we find that providing perfect ground-truth information alone is insufficient to improve GPT-4o's performance, confirming that the bottleneck is fundamentally computational rather than informational. Our results position LLMs not as replacements for operations research, but as natural-language interfaces that make rigorous, solver-based policies accessible to non-experts.

库存优化大模型应用人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。