提出新型跨模态注意力机制,解决新闻与行情不同步时的融合难题。
CGCMA: Conditionally-Gated Cross-Modal Attention for Event-Conditioned Asynchronous Fusion
- 分两阶段处理:先用文本找相关行情,再根据时效性控制信息注入
- 在加密货币市场数据上实现最高0.449的夏普比率提升
- 适合处理新闻延迟、信息冲突等真实场景的多模态建模任务
我们研究异步对齐这一关键多模态学习问题,即需将密集的主信号流与零星到达的外部上下文融合,且其价值依赖于到达时机。不同于假设结构同步的标准多模态基准,此设置要求模型显式推理信息的新鲜度与可信度。聚焦事件驱动情形,连续市场状态与延迟的网络情报配对,仅以高频加密货币市场作为带时间戳、高噪声的测试环境。提出CGCMA(条件门控跨模态注意力)模型,核心设计是分离文本引导的定位与滞后感知的信任控制。文本首先对价格序列进行注意力,识别事件相关市场状态;随后条件门控基于模态一致性、网页特征和滞后时间τₗₐ𝗴,调节残差注入,并在外部信息过时或矛盾时回退至单模态预测。引入CMI(Crypto Market Intelligence)异步评估语料库,包含27,914条真实新闻样本,与延迟的网络情报配对。在现有短篇真实新闻语料上,CGCMA在共享零成本阈值交易评估下,于新闻可用时段取得最高平均夏普比率(+0.449 ± 0.257)。额外控制实验表明,该增益非仅由网页标量解释,也非简单新鲜度启发式可恢复。结果支持该问题的有效性,并在压力测试设置中展现显著的异步多模态收益。
原文摘要 · Abstract (English)
We study asynchronous alignment, a first-class multimodal learning setting in which a dense primary stream must be fused with sporadic external context whose value depends on when it arrives. Unlike standard multimodal benchmarks that assume structural synchrony, this setting requires models to reason explicitly about freshness and trust. We focus on the event-conditioned case in which continuous market states are paired with delayed web intelligence, and we use high-frequency cryptocurrency markets only as a timestamped, high-noise stress test for this broader problem. We propose CGCMA (Conditionally-Gated Cross-Modal Attention), whose central design principle is to separate text-conditioned grounding from lag-aware trust control. Text first attends over price sequences to identify event-relevant market states, after which a conditional gate uses modality agreement, web features, and lag $τ_{\mathrm{lag}}$ to regulate residual injection and fall back toward unimodal prediction when external context is stale or contradictory. We introduce CMI (Crypto Market Intelligence), an asynchronous evaluation corpus with 27,914 real-news samples pairing high-frequency price sequences with lagged web intelligence. On the current short real-news corpus, CGCMA attains the highest mean downstream Sharpe ratio ($+0.449 \pm 0.257$) among the evaluated baselines under a shared zero-cost threshold-trading evaluation on news-available bars. Additional controls show that the gain is not explained by web scalars alone and is not recovered by simple freshness heuristics. The resulting evidence supports problem validity and a promising asynchronous multimodal gain on this stress-test setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。