用个性化智能交易机器人模拟市场,预测价格波动分布。
Persona-Trained Monte Carlo: Estimating Market-Outcome Distributions via Swarms of Persona-Conditioned Neural Policy Bots in a Limit Order Book
- 用个性化的神经网络交易员组成群体,通过反复模拟市场交互
- 通过多轮随机个性抽样和外部冲击,生成真实价格路径的统计分布
- 适合研究市场风险、交易策略或系统性风险的学者与从业者
我们提出人格化训练蒙特卡洛(PTMC),一种通过反复模拟大量人格化神经策略交易机器人在限价订单簿中的交互,来估计市场结果统计分布的方法。每轮模拟中,多个共享同一训练策略网络但具有不同个性参数的机器人参与连续双拍卖,生成一条价格路径作为一次蒙特卡洛采样。通过独立抽取个性群体重复该过程,形成样本集合以估计目标市场统计量。随机性来自个性抽样、运行内动作采样及可选外生冲击,不仅限于价格本身。我们区分了PTMC与经典蒙特卡洛、手工编码模型、单智能体强化学习及基于大语言模型的生成代理等范式。为支持设计,我们综述跨学科基础——包括计算经济学、市场微观结构、行为金融、深度强化学习、生成/大模型代理、新闻驱动交易、系统性风险、经济物理学和博弈论,并将各文献与策略网络、训练数据或验证协议的设计选择对应。我们形式化了PTMC估计器及其收敛性质,指定候选机器人架构与训练目标,并提出四级验证方法:典型事实匹配、微观结构与代理层级检验、以及与零智能基线的历史压力测试对比。框架尚未实现,但贡献包括形式化估计器、跨学科设计依据与验证路线图,并提出开放研究问题。
原文摘要 · Abstract (English)
We propose Persona-Trained Monte Carlo (PTMC), a method for estimating distributions of market-outcome statistics by repeatedly simulating limit-order-book interaction among swarms of persona-conditioned neural-policy trading bots. Each run instantiates many bots sharing one trained policy network but conditioned on heterogeneous, individually sampled persona parameters drawn from a learned trader-heterogeneity distribution; the bots interact in a continuous double auction, and the resulting price path is one Monte Carlo sample. Repeating this over independent persona-population draws yields an ensemble from which a target market statistic is estimated. Randomness enters through persona draws, within-run action sampling, and optional exogenous shocks, not solely through price as in classical Monte Carlo. We distinguish PTMC from adjacent paradigms, including classical Monte Carlo, hand-coded agent-based models, single-agent reinforcement learning, and large-language-model-based generative agents. To justify the design, we survey cross-disciplinary foundations -- agent-based computational economics, market microstructure, behavioral finance, deep reinforcement learning, generative/LLM-based agents, news-driven trading, systemic risk, econophysics, and game theory -- connecting each literature to a specific design choice in the policy network, training data, or validation protocol. We formalize the PTMC estimator and its convergence properties, specify a candidate bot architecture and training objective, and propose a four-level validation methodology: stylized-fact matching, microstructure- and agent-level checks, and historical stress-test comparison against a zero-intelligence baseline. The framework is proposed but not implemented: we contribute a formal estimator, a cross-disciplinary design justification, and a validation roadmap, and conclude with open research questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。