算法预测会改变未来数据,需区分历史与部署风险。
Algometrics: Forecasting Under Algorithmic Feedback
- 提出'算法计量学'框架,分析预测模型影响自身数据生成的过程。
- 历史表现好的模型,在算法拥挤时可能部署风险更高。
- 随机或干预动作可识别短期反馈,提供风险估计的有限样本界。
在算法市场中,预测模型成为其要预测的数据生成过程的一部分。一旦模型输出转化为交易、分配、执行计划或风险控制,就会改变其评估所依赖的未来数据。本文提出‘算法计量学’框架,用于分析受预测算法影响的时间序列演化。该框架区分了被动预测下的历史风险与算法驱动行为带来的部署风险。证明了三个结论:第一,仅凭历史数据无法识别部署风险——即使在一阶线性反馈模型中,无限多个算法中介环境也可能产生相同的历史规律,却对应不同的部署风险;第二,模型在历史上的排名可能因算法拥挤而反转,即历史误差低的模型在算法广泛采用后部署误差反而更高;第三,随机或工具化动作可识别短时线性反馈,并推导出部署风险估计的有限样本边界。研究建议算法市场的时序基准应同时报告反馈敏感性与预测准确性。
原文摘要 · Abstract (English)
In algorithmic markets, predictive models become part of the data-generating process they aim to forecast. Once their outputs are converted into trades, allocations, execution schedules, or risk controls, they change the future data on which they are evaluated. I introduce algometrics, a framework for time series whose evolution depends on the predictive algorithms forecasting them. The framework distinguishes historical risk, measured under passive forecasting, from deployment risk, measured when forecasts drive actions. I prove three results. First, deployment risk is not identifiable from passive historical data alone: even in a one-step linear feedback model, infinitely many algorithm-mediated environments induce the same historical law while implying different deployment risks for the same forecaster. Second, historical model rankings can invert under crowding, so a predictor with lower passive error can have higher deployment error once similar algorithms are adopted. Third, randomized or instrumented actions identify short-horizon linear feedback, and I derive a finite-sample bound for deployment-risk estimation. These results suggest that time-series benchmarks in algorithmic markets should report feedback sensitivity alongside predictive accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。