用大模型辅助标注翻译数据,提升复杂语言对的精准度。
LATA: A Tool for LLM-Assisted Translation Annotation
- 基于提示模板的LLM自动分句对齐,输出严格约束的JSON格式。
- 人机协同工作流支持自定义翻译技术标注,提升标注效率与精度。
- 适合需要高精度翻译语料的研究者,尤其适用于阿拉伯语-英语等复杂语言对。
构建高质量翻译平行语料库已从简单的句子对齐发展为复杂的多层标注任务。这一方法演进对结构差异大的语言对(如阿拉伯语-英语)带来挑战,传统自动化工具常无法捕捉深层语言转换或语义细微差别。本文提出一种新型、基于大语言模型(LLMs)的交互式标注工具,旨在缩小可扩展自动化与专家级人工判断之间的差距。不同于传统统计对齐器,本系统采用基于模板的提示管理器,在严格的JSON输出约束下利用LLMs完成句子分割与对齐。该工具将自动化预处理融入人机协同流程,支持研究者通过独立标注架构精修对齐结果并添加自定义翻译技术标注。借助LLM辅助处理,工具在保持语言学精确性的同时显著提升标注效率,适用于分析专业领域中的复杂翻译现象。
原文摘要 · Abstract (English)
The construction of high-quality parallel corpora for translation research has increasingly evolved from simple sentence alignment to complex, multi-layered annotation tasks. This methodological shift presents significant challenges for structurally divergent language pairs, such as Arabic--English, where standard automated tools frequently fail to capture deep linguistic shifts or semantic nuances. This paper introduces a novel, LLM-assisted interactive tool designed to reduce the gap between scalable automation and the rigorous precision required for expert human judgment. Unlike traditional statistical aligners, our system employs a template-based Prompt Manager that leverages large language models (LLMs) for sentence segmentation and alignment under strict JSON output constraints. In this tool, automated preprocessing integrates into a human-in-the-loop workflow, allowing researchers to refine alignments and apply custom translation technique annotations through a stand-off architecture. By leveraging LLM-assisted processing, the tool balances annotation efficiency with the linguistic precision required to analyze complex translation phenomena in specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。