arXiv:2509.16666cs.CL2025-09EMNLP被引 2

用代理协作分析语言证据,提升母语识别的可靠性。

Robust Native Language Identification through Agentic Decomposition

  • 分角色代理分工处理语言特征,避免依赖表面线索。
  • 在两个数据集上显著降低误导信息影响,提升预测一致性。
  • 适合需要高可信度母语识别的场景,如跨语言研究。

大型语言模型(LLMs)在母语识别(NLI)任务中常依赖姓名、地点和文化刻板印象等表面线索,而非真实反映母语影响的语言模式。为提升鲁棒性,已有方法要求模型忽略这些线索,但本文证明该策略不可靠,模型预测易受误导性提示影响。为此,我们提出一种受司法语言学启发的代理式NLI流程:多个专用代理分别收集并分类多样化的语言证据,最终由一个目标感知的协调代理综合所有证据做出判断。在两个基准数据集上的实验表明,该方法相比标准提示策略,在抵抗误导性上下文线索和保持性能一致性方面均有显著提升。

原文摘要 · Abstract (English)

Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than the underlying linguistic patterns indicative of native language (L1) influence. To improve robustness, previous work has instructed LLMs to disregard such clues. In this work, we demonstrate that such a strategy is unreliable and model predictions can be easily altered by misleading hints. To address this problem, we introduce an agentic NLI pipeline inspired by forensic linguistics, where specialized agents accumulate and categorize diverse linguistic evidence before an independent final overall assessment. In this final assessment, a goal-aware coordinating agent synthesizes all evidence to make the NLI prediction. On two benchmark datasets, our approach significantly enhances NLI robustness against misleading contextual clues and performance consistency compared to standard prompting methods.

母语识别代理系统语言分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。