arXiv:2507.14520cs.AI2025-07EMNLP

让下棋语言模型同时看棋盘,提升理解与稳定性。

What if Othello-Playing Language Models Could See?

  • 结合走法序列与棋盘图像进行多模态训练。
  • 多模态模型在规则理解任务上表现更好,抗干扰更强。
  • 不同架构模型趋向共享内部表征,适合研究认知机制。

语言模型常面临符号接地问题。尽管有人认为无需引入其他模态即可解决,但更多人推测具身学习效率更高。我们在奥赛罗(Othello)这一规则明确、结构简化世界中探索此问题,构建了基于走法序列与棋盘图像联合训练的多模态模型VISOTHELLO。通过奥赛罗规则理解任务,检验多模态学习是否优于纯文本方法。进一步评估在语义无关扰动下的鲁棒性,并分析跨模态对齐的一致性。结果表明,多模态训练不仅提升性能与鲁棒性,还促使不同模型架构收敛至共享的内部表示。该研究为理解模型世界认知提供了可控且可解释的测试平台。

原文摘要 · Abstract (English)

Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that grounded learning is more efficient. We explore this question in Othello, a simplified, rule-based world that offers a controlled and interpretable testbed for studying world understanding. Building on prior work, we introduce VISOTHELLO, a multi-modal model trained jointly on move sequences and board images. Using the Othello rule understanding task, we examine whether multi-modal learning provides advantages over text-only approaches. We further evaluate robustness under semantically irrelevant perturbations and analyze the consistency of cross-modal alignment. Our results suggest that multi-modal training not only improves performance and robustness but also promotes convergence toward shared internal representations across different model architectures.

多模态语言模型符号接地奥赛罗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。