测试大模型破解未编码字形的能力,发现视觉与描述方法各有优劣。
Reasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
- 用字形分词构建多模态谜题数据集,模拟真实破译场景。
- GPT-4o等模型在无Unicode支持的罕见字形上表现有限,视觉模型更优。
- 适合对AI语言破译、跨文明沟通感兴趣的学者和开发者。
我们研究了大型视觉语言模型(LVLM)和大型语言模型(LLM)在破解未纳入Unicode的罕见字形方面的表现。为此,我们提出一种新方法,构建包含此类字形的多模态语言谜题数据集,并采用字形分词技术。针对不同模型,我们设计了图像法(Picture Method)用于LVLM,描述法(Description Method)用于LLM,以应对挑战。实验使用GPT-4o、Gemini和Claude 3.5 Sonnet等主流模型进行评估。结果表明,当前AI方法在语言破译中存在显著局限性,尤其是缺乏Unicode编码时模型性能下降明显;同时,通过文字描述建模视觉语言符号仍具挑战。本研究深化了对AI在语言破译中潜力的理解,也指出了未来研究方向。
原文摘要 · Abstract (English)
We explore the capabilities of LVLMs and LLMs in deciphering rare scripts not encoded in Unicode. We introduce a novel approach to construct a multimodal dataset of linguistic puzzles involving such scripts, utilizing a tokenization method for language glyphs. Our methods include the Picture Method for LVLMs and the Description Method for LLMs, enabling these models to tackle these challenges. We conduct experiments using prominent models, GPT-4o, Gemini, and Claude 3.5 Sonnet, on linguistic puzzles. Our findings reveal the strengths and limitations of current AI methods in linguistic decipherment, highlighting the impact of Unicode encoding on model performance and the challenges of modeling visual language tokens through descriptions. Our study advances understanding of AI's potential in linguistic decipherment and underscores the need for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。