用大模型模拟人类能源选择,准确率达77%。
Can Large Language Models Simulate Human Responses? A Case Study of Stated Preference Experiments in the Context of Heating-related Choices
- 用大模型生成用户对供暖方案的假想选择,测试其真实度。
- 推理型模型DeepSeek-R1准确率最高,达77%,优于其他模型。
- 大模型偏好节能选项,适合研究能源政策与用户行为。
陈述偏好(SP)调查是研究个体在假设或未来情境中权衡决策的关键方法,尤其在低碳技术、分布式可再生能源和需求侧响应等能源转型场景中至关重要。但传统问卷成本高、耗时长,且易受受访者疲劳和伦理限制影响。大型语言模型(LLMs)展现出生成类人文本的强大能力,引发其在调查研究中的应用兴趣。本研究系统评估多个大模型(LLaMA 3.1、Mistral、GPT-3.5、DeepSeek-R1)在能源相关SP调查中模拟消费者选择的能力,涵盖提示设计、上下文学习(ICL)、思维链(CoT)推理、模型类型、与传统选择模型结合及潜在偏差等因素。结果显示,云端模型并未持续优于本地小模型;推理型模型DeepSeek-R1平均准确率达77%,在准确性、因素识别和选择分布一致性方面表现最佳。各模型均呈现对燃气锅炉和不改造选项的系统性偏见,更倾向节能替代方案。结果表明,过往选择是最重要的输入因子,而过长或复杂提示会令模型失焦,降低准确率。
原文摘要 · Abstract (English)
Stated preference (SP) surveys are a key method to research how individuals make trade-offs in hypothetical, also futuristic, scenarios. In energy context this includes key decarbonisation enablement contexts, such as low-carbon technologies, distributed renewable energy generation, and demand-side response [1,2]. However, they tend to be costly, time-consuming, and can be affected by respondent fatigue and ethical constraints. Large language models (LLMs) have demonstrated remarkable capabilities in generating human-like textual responses, prompting growing interest in their application to survey research. This study investigates the use of LLMs to simulate consumer choices in energy-related SP surveys and explores their integration into data analysis workflows. A series of test scenarios were designed to systematically assess the simulation performance of several LLMs (LLaMA 3.1, Mistral, GPT-3.5 and DeepSeek-R1) at both individual and aggregated levels, considering contexts factors such as prompt design, in-context learning (ICL), chain-of-thought (CoT) reasoning, LLM types, integration with traditional choice models, and potential biases. Cloud-based LLMs do not consistently outperform smaller local models. In this study, the reasoning model DeepSeek-R1 achieves the highest average accuracy (77%) and outperforms non-reasoning LLMs in accuracy, factor identification, and choice distribution alignment. Across models, systematic biases are observed against the gas boiler and no-retrofit options, with a preference for more energy-efficient alternatives. The findings suggest that previous SP choices are the most effective input factor, while longer prompts with additional factors and varied formats can cause LLMs to lose focus, reducing accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。