broker在不确定市场中通过上下文优化定价,实现近最优交易收益。
A Tight Regret Analysis of Non-Parametric Repeated Contextual Brokerage
- 基于上下文信息动态调整报价,利用相似上下文预测市场价。
- 在全反馈与有限反馈下均达到紧致的后悔界,表现接近理想情况。
- 证明了理想策略至少能实现完全知情策略一半的交易收益,适合在线决策研究者。
我们研究了一种上下文化的重复经纪问题。每次交互中,两名拥有私有价值评估的交易者根据经纪人提出的、受上下文信息影响的价格决定买卖。经纪人的目标是最大化交易者的净效用(即交易收益),通过最小化相对于已知交易者价值分布的代理(oracle)的后悔值来实现。我们假设交易者的价值是未知物品当前市场价格的零均值扰动,且该价格可任意变化;同时,相似的上下文应对应相似的市场价格。我们分析了两种反馈机制:全反馈(每次交互后披露交易者的真实价值)和有限反馈(仅披露交易尝试结果)。针对两种反馈,我们提出了实现紧致后悔界的算法。此外,我们进一步证明了一个紧致的1/2-近似结果:已知价值分布的代理所能实现的交易收益,至少为完全知晓实际交易者价值实现的全能代理的一半。
原文摘要 · Abstract (English)
We study a contextual version of the repeated brokerage problem. In each interaction, two traders with private valuations for an item seek to buy or sell based on the learner's-a broker-proposed price, which is informed by some contextual information. The broker's goal is to maximize the traders' net utility-also known as the gain from trade-by minimizing regret compared to an oracle with perfect knowledge of traders' valuation distributions. We assume that traders' valuations are zero-mean perturbations of the unknown item's current market value-which can change arbitrarily from one interaction to the next-and that similar contexts will correspond to similar market prices. We analyze two feedback settings: full-feedback, where after each interaction the traders' valuations are revealed to the broker, and limited-feedback, where only transaction attempts are revealed. For both feedback types, we propose algorithms achieving tight regret bounds. We further strengthen our performance guarantees by providing a tight 1/2-approximation result showing that the oracle that knows the traders' valuation distributions achieves at least 1/2 of the gain from trade of the omniscient oracle that knows in advance the actual realized traders' valuations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。