arXiv:2409.03762q-fin.STcs.LG2024-09被引 1

用混合模型提取市场特征,提升预测收益。

Combining supervised and unsupervised learning methods to predict financial market movements

  • 结合线性模型与高斯混合模型提取新特征
  • KNN和随机森林平均回报高于随机策略
  • 适用于量化交易研究者与金融算法开发者

交易员买卖资产的决策依赖多种分析,需具备识别可获利模式的专业能力。本文利用线性模型与高斯混合模型(GMM),从比特币、佩佩币和纳斯达克等新兴与成熟金融市场中提取新型特征,旨在发现盈利机会。基于约六个月的分钟级数据(每市场前59分钟数据用于特征提取,预测未来一小时走势),对比了新型特征与常用特征的表现。评估了随机森林(RF)与K近邻(KNN)等机器学习方法的分类性能,并以等概率随机选择作为基准。采用时间交叉验证,测试集占总时长的40%、30%和20%。结果表明,对时间序列进行过滤有助于算法泛化;其中,经GMM过滤后,KNN与RF的平均回报显著高于随机策略。

原文摘要 · Abstract (English)

The decisions traders make to buy or sell an asset depend on various analyses, with expertise required to identify patterns that can be exploited for profit. In this paper we identify novel features extracted from emergent and well-established financial markets using linear models and Gaussian Mixture Models (GMM) with the aim of finding profitable opportunities. We used approximately six months of data consisting of minute candles from the Bitcoin, Pepecoin, and Nasdaq markets to derive and compare the proposed novel features with commonly used ones. These features were extracted based on the previous 59 minutes for each market and used to identify predictions for the hour ahead. We explored the performance of various machine learning strategies, such as Random Forests (RF) and K-Nearest Neighbours (KNN) to classify market movements. A naive random approach to selecting trading decisions was used as a benchmark, with outcomes assumed to be equally likely. We used a temporal cross-validation approach using test sets of 40%, 30% and 20% of total hours to evaluate the learning algorithms' performances. Our results showed that filtering the time series facilitates algorithms' generalisation. The GMM filtering approach revealed that the KNN and RF algorithms produced higher average returns than the random algorithm.

金融预测特征工程机器学习量化交易

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。