arXiv:2605.06276cs.CLcs.AI2026-05ACL被引 1

为低资源方言口语设计了语义分割新方法,提升对话类文本识别准确率。

Linear Semantic Segmentation for Low-Resource Spoken Dialects

  • 基于本地化语义连贯性建模,增强对话语断续的鲁棒性。
  • 在方言口语数据上表现优于主流基线模型,尤其在非新闻类文本中。
  • 适用于其他低资源口语语言,提供可复用的评测基准和方法框架。

语义分割是话语分析的核心组件,但现有模型主要针对高资源书面文本开发与评估,难以应对低资源口语变体。以阿拉伯方言为例,其语法非正式、存在混码现象且话语结构标记弱,给标准分割方法带来挑战。本文构建了一个多语体基准(超1000个样本),涵盖电话对话、混码播客、广播新闻及小说对白等场景,由母语者标注并验证。实验表明,适用于现代标准阿拉伯语新闻的模型在方言口语上性能显著下降。为此,我们提出一种聚焦局部语义连贯性与抗话语断裂的新模型,在非新闻类方言文本上持续超越强基线。该基准与方法可推广至其他低资源口语语言。

原文摘要 · Abstract (English)

Semantic segmentation is a core component of discourse analysis, yet existing models are primarily developed and evaluated on high-resource written text, limiting their effectiveness on low-resource spoken varieties. In particular, dialectal Arabic exhibits informal syntax, code-switching, and weakly marked discourse structure that challenge standard segmentation approaches. In this paper, we introduce a new multi-genre benchmark (more than 1000 samples) for semantic segmentation in conversational Arabic, focusing on dialectal discourse. The benchmark covers transcribed casual telephone conversations, code-switched podcasts, broadcast news, and expressive dialogue from novels, and was annotated and validated by native Arabic annotators. Using this benchmark, we show that segmentation models performing well on MSA news genres degrade on dialectal transcribed speech. We further propose a segmentation model that targets local semantic coherence and robustness to discourse discontinuities, consistently outperforming strong baselines on dialectal non-news genres. The benchmark and approach generalize to other low-resource spoken languages.

语义分割方言识别低资源语言口语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。