测试大模型在业务数据探索中的可靠性,发现多数表现不可靠。
Business Utility of Large Language Models as Exploratory Data Analysis Agents

- 用供应链模拟构建评测基准,考察模型从间接数据推断问题的能力
- 多数配置平均表现尚可但重复性差,仅GPT-5.4高推理版达标
- 提出'业务效用'指标,综合衡量准确性与稳定性,适合企业落地评估
大型语言模型(LLMs)在分析工作流中日益普及,但在业务场景下作为探索性数据分析(EDA)代理的适用性仍不确定。一个可部署的EDA代理不仅需要良好平均性能,还需足够可重复性以建立信任。本文在基于代理的供应链仿真中构建了可控、贴近业务的基准测试,任务是从间接运营痕迹中推理导致质量低下和下游销售损失的供应商-产品组合,而非依赖显式标签。评估了来自八个模型家族的十五种模型变体配置,在四种不同条件(数据表示、提示清晰度、信号强度)下进行,每条件五条轨迹。输出结果通过杰卡德指数与确定性真实值对比评分,并采用结合均值分数(ms)、变异系数(CV)、跨条件显著性检验及本文提出的‘业务效用’(风险调整后单一操作度量)的框架进行评估。结果显示,大多数配置即使平均得分尚可,也不够可靠,无法用于自主EDA。GPT-5.4搭配极高推理努力时表现最优,实验平均ms为0.8748,实验平均业务效用达0.6952;其余最佳配置在可变性折扣后效用大幅下降。研究建议:对EDA代理的评估应将平均质量、可重复性和条件敏感性视为操作可信度的互补维度。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in analytical workflows, but their suitability as exploratory data analysis (EDA) agents in business settings remains uncertain. In practice, a deployable EDA agent must provide not only useful average performance but also sufficient repeatability to support trust in its outputs. We evaluate this requirement in a controlled, business-relevant benchmark built on an agent-based supply chain simulation. The task is to identify supplier-product combinations responsible for low quality and downstream sales loss by reasoning from indirect operational traces rather than from explicit labels. Fifteen model-variant configurations from eight model families were evaluated under four experimental conditions that varied data representation, prompt clarity, and signal strength, with five trajectories per condition. Outputs were scored against deterministic ground truth using the Jaccard index and assessed through a framework that combines mean score (ms), coefficient of variation (CV), exploratory cross-condition significance tests, and Business utility, a risk-adjusted metric that we propose to summarise quality and repeatability in a single operational measure. The results show that most configurations are not reliable enough for autonomous EDA use, even when their average scores appear acceptable. GPT-5.4 with extra-high reasoning effort achieved the strongest overall profile, with an experiment-averaged ms of 0.8748 and an experiment-averaged Business utility of 0.6952, while the next-best configurations lost substantially more utility after variability discounting. Our findings suggest that evaluation of EDA agents should treat average quality, repeatability, and condition sensitivity as complementary dimensions of operational trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。