用连续时间样条重构噪声数据,提升决策模型在嘈杂信号中的学习能力
Kinematic Tokenization: Optimization-Based Continuous-Time Tokens for Learnable Decision Policies in Noisy Time Series
- 基于优化的样条重构法,将噪声观测转为位置速度加速度等连续系数
- 在多资产日频数据上,连续令牌使策略稳定且不退化为完全避险
- 适合需要理性回避风险的金融决策等高噪声场景
Transformer 专为离散标记设计,但许多真实信号是通过噪声采样获得的连续过程。离散标记(原始值、片段、有限差分)在信噪比低时易失效,尤其当下游目标对错误有不对称惩罚,理性引导放弃决策。本文提出运动学标记(Kinematic Tokenization),一种基于优化的连续时间表示:从噪声测量中重建显式样条,并对局部样条系数(位置、速度、加速度、加速度变化率)进行标记。该方法应用于资产价格与交易量曲线的金融时间序列数据。在多资产日频股票测试集上,采用风险规避型非对称分类目标作为学习性压力测试。在此目标下,多个离散基线退化为吸收态现金策略(清算均衡),而连续样条标记保持校准的非平凡动作分布和稳定策略。结果表明,显式的连续时间标记可提升噪声时间序列中选择性决策策略的学习性和校准性。
原文摘要 · Abstract (English)
Transformers are designed for discrete tokens, yet many real-world signals are continuous processes observed through noisy sampling. Discrete tokenizations (raw values, patches, finite differences) can be brittle in low signal-to-noise regimes, especially when downstream objectives impose asymmetric penalties that rationally encourage abstention. We introduce Kinematic Tokenization, an optimization-based continuous-time representation that reconstructs an explicit spline from noisy measurements and tokenizes local spline coefficients (position, velocity, acceleration, jerk). This is applied to financial time series data in the form of asset prices in conjunction with trading volume profiles. Across a multi-asset daily-equity testbed, we use a risk-averse asymmetric classification objective as a stress test for learnability. Under this objective, several discrete baselines collapse to an absorbing cash policy (the Liquidation Equilibrium), whereas the continuous spline tokens sustain calibrated, non-trivial action distributions and stable policies. These results suggest that explicit continuous-time tokens can improve the learnability and calibration of selective decision policies in noisy time series under abstention-inducing losses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。