arXiv:2509.10644cs.CL2025-09EMNLP被引 6

用用户中心设计改进语言文档的形态分析工具,让技术真正落地。

Interdisciplinary Research in Conversation: A Case Study in Computational Morphology for Language Documentation

  • 引入用户中心设计,重构语言文档中的形态分析研究
  • 实测发现先进模型仍无法满足实际使用需求
  • 适合语言学研究者与NLP交叉方向开发者

计算形态学有望通过形态切分和逐行注释文本(IGT)生成支持语言文档工作。然而,当前研究成果在真实语言记录场景中应用有限。本文指出,这反映出自然语言处理领域研究与实践之间的深层脱节,并警告若不系统融入用户中心设计(UCD),该领域将趋于脱离语境、失去实效。为展示UCD如何重塑研究方向,本文以最先进的多语言IGT生成模型GlossLM为例,开展小规模用户研究,邀请三位语言记录学者参与。结果表明,尽管模型在指标上表现优异,但在实际文档工作中仍无法满足核心可用性需求。这些发现引出关于模型约束、标注标准化、切分策略与个性化的新研究问题。我们认为,以用户为中心不仅可提升工具实用性,更能激发更丰富、更相关的研究方向。

原文摘要 · Abstract (English)

Computational morphology has the potential to support language documentation through tasks like morphological segmentation and the generation of Interlinear Glossed Text (IGT). However, our research outputs have seen limited use in real-world language documentation settings. This position paper situates the disconnect between computational morphology and language documentation within a broader misalignment between research and practice in NLP and argues that the field risks becoming decontextualized and ineffectual without systematic integration of User-Centered Design (UCD). To demonstrate how principles from UCD can reshape the research agenda, we present a case study of GlossLM, a state-of-the-art multilingual IGT generation model. Through a small-scale user study with three documentary linguists, we find that despite strong metric based performance, the system fails to meet core usability needs in real documentation contexts. These insights raise new research questions around model constraints, label standardization, segmentation, and personalization. We argue that centering users not only produces more effective tools, but surfaces richer, more relevant research directions

语言文档用户中心形态学IGT生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。