arXiv:2609.03775cs.CL2026-09

用大模型基于上下文学习预测语言类型特征,提升可解释性与跨资源表现。

Typological Feature Prediction with Large Language Models: An In-Context Learning Approach

  • 利用语言谱系与地理邻近信息构建上下文提示,增强预测能力。
  • 在多种语言资源水平上均超越基线,低资源语言无性能劣势。
  • 生成的推理理由与输入证据高度一致,推动可解释预测发展。

语言类型特征在多语言自然语言处理中广泛应用,其缺失值预测具有下游应用价值。然而,现有方法缺乏可解释的预测依据,且在不同资源水平和特征类型下的表现尚未充分探索。鉴于大语言模型在元语言推理与提供理由方面的能力,本文研究了基于上下文学习的方法在类型特征预测中的表现,使用来自 URIEL+ 与 Glottolog 的语言数据。发现零样本提示效果不足,但引入谱系与地理邻近证据后,大模型显著优于所有基线,且不损害低资源语言的表现。进一步发现,大多数大模型生成的推理理由与输入证据一致,为可解释的语言类型特征预测提供了新路径。

原文摘要 · Abstract (English)

Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels and feature types remains underexplored. Given LLMs' abilities in meta-linguistic reasoning and in providing rationales, we investigate LLMs' performance in typological feature prediction via an in-context learning approach with linguistic data from URIEL+ and Glottolog. We find that zero-shot prompting is insufficient, but when given phylogenetic and geographic neighbour evidence, LLMs substantially outperform all baselines without disadvantaging low-resource languages. We further find that most LLM rationales are consistent with the provided evidence, offering a step toward explainable typological feature prediction.

语言类型学大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。