用多视角变换器+测试时微调,提升机器解视觉谜题能力
Multi-Perspective Transformers in ARC-AGI-2 Challenge

- 构建多视角变换器模型,融合不同角度的推理线索
- 在训练集上达96.1%准确率,评估集21.7%准确率
- 适合研究少样本泛化与动态推理的AI学者
ARC-AGI-2是一个衡量机器从少量示例中泛化、理解符号意义并在不同情境灵活应用规则的人类直觉型视觉谜题基准。本文介绍我们使用TinyLM解决该基准的方法,包括在测试时进行微调,如测试时训练(Test-Time-Training, TTT)和专家产品(Products of Experts, POE)。模型在训练集上达到96.1%的准确率,在评估集上为21.7%。
原文摘要 · Abstract (English)
ARC-AGI-2 is a benchmark of human-intuitive visual puzzles that measures a machine's ability to generalize from limited examples, interpret symbolic meaning, and flexibly apply rules in varying contexts. In this paper, we discuss our approach to solving the ARC-AGI-2 puzzles with TinyLM, with additional fine-tuning at test time, including Test-Time-Training (TTT) and Products of Experts (POE). Our model achieves 96.1% accuracy on the training set and 21.7% accuracy on the evaluation set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。