arXiv:2603.22497cs.CL2026-03

用密码转换高资源语言,模拟全新语言研究上下文学习。

Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning

  • 将高资源语言通过可逆加密生成虚构新语言
  • 在多种任务上验证现有方法在未知语言上的表现
  • 适合研究上下文学习机制与跨语言泛化能力的学者

随着大语言模型在未见语言上的上下文语言学习(ICLL)研究兴起,这些语言普遍缺乏自然语言处理工具、数据资源和研究者经验,导致进展难以评估,难以开展低成本大规模实验,且现有成果多局限于少数语言与任务。为突破此限制,我们提出Rashid框架:通过可逆加密高资源语言(HRLs),构建真正未见的语言,同时保留对HRL丰富资源的访问,从而实现此前无法探索的ICLL现象研究。利用该框架,我们采用当前最优评估工具与人工分析,评估现有方法性能,探究昂贵资源对ICLL的提升作用,并在机器翻译以外的丰富下游任务中测试ICLL策略。本研究展示了框架带来的新可能性,也提供了关于现有表现与未来方向的实用洞见。

原文摘要 · Abstract (English)

Where there is growing interest in in-context language learning (ICLL) for unseen languages with large language models, such languages usually suffer from the lack of NLP tools, data resources, and researcher expertise. This means that progress is difficult to assess, the field does not allow for cheap large-scale experimentation, and findings on ICLL are often limited to very few languages and tasks. In light of such limitations, we introduce a framework (Rashid), for studying ICLL wherein we reversibly cipher high-resource languages (HRLs) to construct truly unseen languages with access to a wide range of resources available for HRLs, unlocking previously impossible exploration of ICLL phenomena. We use our framework to assess current methods in the field with SOTA evaluation tools and manual analysis, explore the utility of potentially expensive resources in improving ICLL, and test ICLL strategies on rich downstream tasks beyond machine translation. These lines of exploration showcase the possibilities enabled by our framework, as well as providing actionable insights regarding current performance and future directions in ICLL.

上下文学习语言模拟NLP研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。