用类比构造谜题,测试大模型的推理覆盖能力。
Riddle Quest : The Enigma of Words
- 构建四步流程:事实生成、属性筛选、谜面创作、答案验证。
- 模型常只猜到主要答案,遗漏其他合理解释。
- 适合研究语言模型对歧义和多义性的理解能力。
谜题是通过间接、比喻或戏谑线索描述事物或概念的简短语言谜题,需解谜者解读提示、识别模式并推断出答案。本文提出一个简单的类比式谜题生成与评估流水线,包含三元组生成器、语义映射器、风格化生成器和验证器。验证器用于收集谜题可能指向的所有答案,以此研究大语言模型能否恢复不同谜题类型下的完整答案集。案例研究表明,尽管模型常能猜中主要目标答案,但往往遗漏其他有效解释,凸显了谜题作为轻量级工具在检验语言模型推理覆盖范围与歧义处理能力方面的价值。
原文摘要 · Abstract (English)
Riddles are concise linguistic puzzles that describe an object or idea through indirect, figurative, or playful clues. They are a longstanding form of creative expression, requiring the solver to interpret hints, recognize patterns, and draw inferences to identify the answers. In this work, we introduce a simple pipeline for creating and evaluating analogy-based riddles. The system includes a triples creator that builds structured facts about a concept, a semantic mapper that selects attributes useful for analogy, a stylized generator that turns them into riddle clues, and a validator that collects all possible answers the riddle could point to. We use this validator to study whether large language models can recover the full answer set for different riddle types. Our case study shows that while models often guess the main intended answer, they frequently miss other valid interpretations. This highlights the value of riddles as a lightweight tool for examining reasoning coverage and ambiguity handling in language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。