通过观察订单簿变化,提升市场做市的预测能力与收益。
Online Market Making and the Value of Observing the Order Book
- 设计新算法利用订单簿未成交时的供需信息进行学习。
- 在随机和均值回归场景下,实现高概率 $O(\ oot{2}{T})$ 的后悔界。
- 适用于真实交易场景,尤其适合追求低风险做市的机构投资者。
我们研究了一种在线市场做市问题:学习者需持续发布单一资产的买卖报价,与持有私有价值评估的交易者交互。不同于以往完全隐藏反馈的在线学习模型,本文引入一种受动作影响的反馈机制——当交易发生时,交易者估值仍不可见;当无交易时,可获得有关供需的有用信息。我们证明这一额外信息从根本上改变了问题的学习性。在独立同分布市场价的随机设定下,提出基于消除的算法,在无需估值分布光滑性假设的前提下,实现了高概率 $O(\sqrt{T})$ 的后悔上界。进一步扩展至一类广义的均值回归价格过程,包括局部自回归动态和基于累计偏离均值的全局漂移条件,同样建立高概率 $O(\sqrt{T})$ 后悔界,依赖于一个独立感兴趣的新型集中不等式。在对抗性设定(对手价格)下,设计了探索-扰动算法,保证期望后悔为 $O(T^{2/3})$。结果量化了观察订单簿在在线做市中的价值,表明即使有限的、依赖动作的反馈,也能显著优于标准老虎机反馈模型的后悔性能。
原文摘要 · Abstract (English)
We study an online market-making problem in which a learner sequentially posts bid and ask prices for a single asset while interacting with traders holding private valuations. Unlike existing online learning formulations that assume fully censored feedback, we introduce an action-dependent feedback model inspired by real limit order books: when a trade occurs, the trader's valuation remains hidden, whereas when no trade occurs, informative feedback about supply and demand is revealed. We show that this additional information fundamentally changes the learnability of the problem. In the stochastic setting with i.i.d. market prices, we propose an elimination-based algorithm that achieves $O(\sqrt T)$ regret with high probability, without requiring any smoothness assumptions on the distribution of trader valuations. We then extend this result to a broad class of mean-reverting price processes by considering both local, autoregressive dynamics and a weaker global drift condition based on cumulative deviations from the mean. Under either assumption, we establish high-probability $O(\sqrt T)$ regret bounds, relying on a new concentration inequality of independent interest. Finally, in the adversarial setting with oblivious prices, we design an explore-then-perturb algorithm that guarantees $O(T^{2/3})$ regret in expectation. Our results quantify the value of observing the order book in online market making and demonstrate that even limited, action-dependent feedback can substantially improve regret guarantees compared to standard bandit feedback models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。