大模型证明词语类比可用平行四边形模型解释,优于人类表现。
Large Language Models provide support for the parallelogram theory of analogy
- 用大模型生成类比,比人类更符合向量平行四边形结构。
- 模型类比正确率更高,且与嵌入空间中的平行关系更强。
- 适合对类比推理、语言模型机制感兴趣的读者。
四词类比(A:B::C:D)传统上被建模为几何上的平行四边形:B-A+C 应得 D。近期研究认为该模型无法捕捉人类类比生成方式,局部相似性启发法更具解释力(Peterson 等,2020)。我们对比了人类与大语言模型(LLM)在 Peterson 等人(2020)数据集上的类比补全表现。结果发现,LLM 生成的类比被普遍评价为更优,且在分布嵌入空间中更符合平行四边形结构。关键在于,模型优势源于更强的平行对齐能力,而非对局部相似性的敏感度提升。进一步地,将 GloVe 模型微调以更好满足平行四边形约束后,其最佳候选答案更接近人类与 LLM 实际选择,并显著提升对人类评分的预测能力。整体表明,平行四边形模型仍可作为词语类比的有效解释框架。
原文摘要 · Abstract (English)
Four-term word analogies (A:B::C:D) are classically modeled geometrically as parallelograms: adding the vector B-A+C produces D. Recent work suggests that this model poorly captures how humans produce analogies, with simple local-similarity heuristics often providing a better account (Peterson et al., 2020). But does the parallelogram model fail because it is a bad model of analogical relations, or because people are not very good at generating relation-preserving analogies? We compared human and large language model (LLM) analogy completions on the set of problems from Peterson et al. (2020). We find that LLM-generated analogies are reliably judged as better than human-generated ones, and are also more consistent with parallelograms in a distributional embedding space. Crucially, we show that the improvement over human analogies is driven by greater parallelogram alignment and reduced reliance on accessible words rather than enhanced sensitivity to local similarity. Finally, fine-tuning GloVe to better satisfy the parallelogram constraint makes the model's top-ranked candidates more likely to be the completions humans and LLMs actually produced, and improves its prediction of human ratings. Overall, these results provide support for the parallelogram model of word analogies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。