构建可审计的结构化通信协议,实现代理间阿尔法发现的确定性回放与验证。
VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery
- 将自由文本通信改为带类型的因果单播记录,形成可追踪的因果轨迹。
- 在CSI~1000外样本上,唯一实现正中位年化收益与夏普比率,顶20组合夏普达0.71。
- 核心价值在于可验证性与防泄露设计,而非预测准确率,适合量化研究者参考。
代理间(A2A)阿尔法发现因挖掘与评估代理间的重复反馈循环而受阻。当前大模型多代理系统中的信息传递依赖自由形式的自然语言消息,缺乏稳定契约且无法重放。本文首次将通信重构为具有类型、因果可寻址、单播特性的结构化协议,使提交的数据流构成因果轨迹。在此轨迹上,一个四头类型预测器可提前数轮预测两方挖掘者获得的累积指引;事务性验证-跃迁控制器仅在通过四级验证后才提交多轮推测结果,否则回滚至精确前序状态。结构本身是关键贡献,其价值不在于预测精度。控制性消融实验显示,同等信息量的自由文本通道可达相同预测命中率。类型化设计提供了可模式校验、确定性重放、并从构造上防止预测泄露给评估者的审计能力。在CSI~1000外样本测试中,本方法单次运行是八种方法(七基准+自身)中唯一在因子层面保持正中位年化收益与夏普比率的结果,尽管所有方法(包括本方法)的超额中位收益仍为负。其开发选定的前20组合在优化期内划分的样本上实现0.71的中位外样本夏普比率。报告为描述性结果,未扣除成本,全文明确说明其局限性,尤其未分离跃迁机制与继承搜索底座的影响,留待未来工作。
原文摘要 · Abstract (English)
Agent-to-agent (A2A) alpha discovery is slowed by repeated feedback cycles between mining and evaluation agents, whose hand-offs, in contemporary LLM multi-agent systems, are free-form natural-language messages that carry no stable contract and cannot be replayed. We first restructure this communication as a structured agent-to-agent protocol of \emph{typed, causally addressable, unicast records}, so that the committed stream forms a causal trajectory. On that trajectory a single predictor with four typed heads forecasts the accumulated guidance the two miners would receive several cycles ahead; a transactional verify--leap controller then commits a multi-cycle speculative outcome only when it passes a four-level gate, and otherwise rolls back to the exact prior state. Structure is the enabling contribution, and its value is not accuracy. A controlled ablation shows an equal-information free-text channel reaches the same predictor hit rate. What typing provides is a state that can be schema-checked, replayed deterministically, and prevented by construction from leaking a forecast to an evaluator: auditability by construction, not an empirically stress-tested guarantee. On a CSI~1000 out-of-sample holdout, our single run is the only one among eight methods (seven baselines and ours) to hold a positive median annualized return and Sharpe at the factor level, though the median return \emph{in excess} of the benchmark stays negative for every method including ours; its development-selected top-20 portfolios reach a $0.71$ median holdout Sharpe, selected on a split inside the optimization horizon. We report these single-run results descriptively, gross of costs, and are explicit about their limits throughout; in particular we do not isolate the effect of the leap machinery from the inherited search substrate, which we leave to future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。