用理论解构造数据,让大模型更懂群体投资决策。
InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior
- 用简单投资问题的理论解生成训练数据,替代真实用户数据。
- 训练收敛速度比真实数据快,学习效率更高。
- 适合研究行为金融与大模型对齐的学者和从业者。
在行为金融领域,将大语言模型(LLMs)与群体羊群行为下的投资者决策过程对齐是一项关键挑战,其核心限制在于监督微调(SFT)所需的真实用户数据稀缺。尽管SFT可缩小模型输出与人类行为模式之间的差距,但其对海量真实数据的依赖带来高昂的数据收集成本和隐私风险。本文提出InvestAlign框架,通过利用相似且简单的最优投资问题的理论解来构建高质量SFT数据集,而非复杂场景。理论分析表明,使用InvestAlign生成数据训练的LLM参数收敛速度优于真实用户数据,显示出更高的学习效率。此外,我们开发了基于InvestAlign微调的InvestAgent,该代理在简单与复杂投资问题中均表现出与真实用户数据更接近的对齐效果,显著优于预微调模型。这表明InvestAlign是一种有潜力解决复杂最优投资问题并实现大模型与投资者决策对齐的新方法。代码已公开于https://github.com/thu-social-network-research-group/InvestAlign。
原文摘要 · Abstract (English)
Aligning Large Language Models (LLMs) with investor decision-making processes under herd behavior is a critical challenge in behavioral finance, which grapples with a fundamental limitation: the scarcity of real-user data needed for Supervised Fine-Tuning (SFT). While SFT can bridge the gap between LLM outputs and human behavioral patterns, its reliance on massive authentic data imposes substantial collection costs and privacy risks. We propose InvestAlign, a novel framework that constructs high-quality SFT datasets by leveraging theoretical solutions to similar and simple optimal investment problems rather than complex scenarios. Our theoretical analysis demonstrates that training LLMs with InvestAlign-generated data achieves faster parameter convergence than using real-user data, suggesting superior learning efficiency. Furthermore, we develop InvestAgent, an LLM agent fine-tuned with InvestAlign, which demonstrates significantly closer alignment to real-user data than pre-SFT models in both simple and complex investment problems. This highlights our proposed InvestAlign as a promising approach with the potential to address complex optimal investment problems and align LLMs with investor decision-making processes under herd behavior. Our code is publicly available at https://github.com/thu-social-network-research-group/InvestAlign.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。