arXiv:2502.14394cs.CL2025-02AAAI被引 7

用跨领域数据提升葡语变体识别准确率

Enhancing Portuguese Variety Identification with Cross-Domain Approaches

  • 构建跨领域葡语变体数据集PtBrVarId,解决巴西葡语数据主导问题
  • 基于Transformer的分类器在跨域场景下实现高精度识别
  • 开源代码与数据,助力欧洲葡语资源建设

自然语言处理的进展提升了生成模型在多种语言变体中生成连贯文本的能力。以葡萄牙语为例,网络上以巴西葡语语料为主,导致模型存在语言偏见,限制其在巴西以外的应用。为填补这一空白并推动欧洲葡语资源建设,我们开发了跨领域语言变体识别器(LVI),用于区分欧洲葡语与巴西葡语。基于文献综述结果,我们构建了跨领域的PtBrVarId语料库,并研究了基于Transformer的LVI分类器在跨域场景下的有效性。尽管本研究聚焦于两种葡语变体,但方法可扩展至其他语言变体。相关代码、语料和模型均已开源,以促进该任务的进一步研究。

原文摘要 · Abstract (English)

Recent advances in natural language processing have raised expectations for generative models to produce coherent text across diverse language varieties. In the particular case of the Portuguese language, the predominance of Brazilian Portuguese corpora online introduces linguistic biases in these models, limiting their applicability outside of Brazil. To address this gap and promote the creation of European Portuguese resources, we developed a cross-domain language variety identifier (LVI) to discriminate between European and Brazilian Portuguese. Motivated by the findings of our literature review, we compiled the PtBrVarId corpus, a cross-domain LVI dataset, and study the effectiveness of transformer-based LVI classifiers for cross-domain scenarios. Although this research focuses on two Portuguese varieties, our contribution can be extended to other varieties and languages. We open source the code, corpus, and models to foster further research in this task.

语言识别跨域学习葡语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。