arXiv:2603.24258cs.CL2026-03中稿 · LREC 2026

跨古埃及语四个阶段的语义对齐,提升历史语言模型性能。

Semantic Alignment across Ancient Egyptian Language Stages via Normalization-Aware Multitask Learning

  • 多任务联合训练,融合音译与音标辅助视图增强对齐。
  • 音标结合KL一致性使跨语支对齐提升,翻译任务收益最大。
  • 为稀疏数据下的历史语言建模提供可复现基线与设计指南。

我们研究了古埃及语四个历史阶段之间的词级语义对齐问题。这些阶段在书写系统和正字法上差异显著,平行数据稀缺。我们联合训练一个紧凑的编码器-解码器模型,使用共享的字节级分词器,在所有四个阶段上同时执行掩码语言建模(MLM)、翻译语言建模(TLM)、序列到序列翻译及词性标注任务,采用任务感知损失并固定权重与基于不确定性的缩放机制。为减少表面差异,引入拉丁转写和IPA音标重构作为辅助视图,并通过KL一致性与嵌入层融合进行整合。使用精心构建的埃及语-英语及内部同源词数据集,通过成对评估指标(如ROC-AUC与三元组准确率)衡量对齐质量。结果表明,翻译任务带来最强增益;结合KL一致性的IPA视图显著改善跨分支对齐效果,而早期融合则作用有限。尽管整体对齐仍受限,但本研究提供了可复现的基线,并为真实条件下历史语言建模提供实用指导,揭示了归一化与任务设计如何影响形态差异大的语境中对齐的定义。

原文摘要 · Abstract (English)

We study word-level semantic alignment across four historical stages of Ancient Egyptian. These stages differ in script and orthography, and parallel data are scarce. We jointly train a compact encoder-decoder model with a shared byte-level tokenizer on all four stages, combining masked language modeling (MLM), translation language modeling (TLM), sequence-to-sequence translation, and part-of-speech tagging under a task-aware loss with fixed weights and uncertainty-based scaling. To reduce surface divergence we add Latin transliteration and IPA reconstruction as auxiliary views. We integrate these views through KL-based consistency and through embedding-level fusion. We evaluate alignment quality using pairwise metrics, specifically ROC-AUC and triplet accuracy, on curated Egyptian-English and intra-Egyptian cognate datasets. Translation yields the strongest gains. IPA with KL consistency improves cross-branch alignment, while early fusion demonstrates limited efficacy. Although the overall alignment remains limited, the findings provide a reproducible baseline and practical guidance for modeling historical languages under real constraints. They also show how normalization and task design shape what counts as alignment in typologically distant settings.

历史语言语义对齐多任务学习古埃及语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。