arXiv:2507.21790econ.EMcs.AI2025-07被引 1

LLM能辅助构建和估算选择模型,但效果依赖提示策略与模型类型。

Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities

  • 用思维链提示让大模型生成合理效用函数,提升建模质量。
  • 部分模型(如GPT-o3)可自动生成代码并正确估算自身模型。
  • 仅提供数据字典比全量数据更利于模型推理,适合初学者使用。

大型语言模型(LLMs)在多个领域广泛应用,但在离散选择建模中的潜力仍待探索。本文系统评估了七种主流大模型(ChatGPT、Claude、DeepSeek、Gemini、Gemma、Llama、Mistral)在十二个版本下的表现,涵盖五种实验配置:建模目标(仅建议或建议并估计)、提示策略(零样本与思维链)、信息可用性(完整数据集或仅数据字典)。所有由模型提出的规格均被实现、估计并评估,基于拟合优度、行为合理性及模型复杂度。结果表明,专有模型在结构化提示下可生成有效且行为合理的效用表达式;而开放权重模型(如Llama、Gemma)难以产出有意义规格。有趣的是,部分模型在仅有数据字典时表现更佳,提示限制数据访问可能促进内部推理。其中,GPT-o3在代理模式下可执行自生成代码,成功估算自身模型。研究揭示了大模型作为建模助手的潜力与局限,为选择建模流程中整合工具提供了实用指导。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are becoming widely used to support various workflows across different disciplines, yet their potential in discrete choice modelling remains relatively unexplored. This work examines the potential of LLMs as assistive agents in the specification and, where technically feasible, estimation of Multinomial Logit models. We implement a systematic experimental framework involving twelve versions of seven leading LLMs (ChatGPT, Claude, DeepSeek, Gemini, Gemma, Llama, and Mistral) evaluated under five experimental configurations. These configurations vary along three dimensions: (i) modelling goal (suggesting vs. suggesting and estimating MNL models); (ii) prompting strategy (Zero-Shot vs. Chain-of-Thoughts (CoT)); and (iii) information availability (full dataset vs. data dictionary summarising variable names and types). Each specification suggested by the LLMs is implemented, estimated, and evaluated based on goodness-of-fit metrics, behavioural plausibility, and model complexity. Our findings reveal that proprietary LLMs can generate valid and behaviourally sound utility specifications, particularly when guided by structured prompts (CoT). Open-weight models such as Llama and Gemma struggled to produce meaningful specifications. Notably, some LLMs performed better when provided with just data dictionary, suggesting that limiting raw data access may enhance internal reasoning capabilities. Among all LLMs, GPT o3, operating in an agentic setting, was uniquely capable of correctly estimating its own specifications by executing self-generated code. Overall, the results demonstrate both the promise and current limitations of LLMs as assistive agents in discrete choice modelling, not only for model specification but also for supporting modelling decision and estimation, and provide practical guidance for integrating these tools into choice modellers' workflows.

大模型选择建模提示工程智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。