一站式机器翻译后编辑与数据收集框架,提升效率与质量。
TranslationCorrect: A Unified Framework for Machine Translation Post-Editing with Predictive Error Assistance
- 集成翻译生成、错误预测与交互式编辑界面。
- 用户研究显示效率与满意度显著优于传统方法。
- 输出符合MQM标准的高质量错误标注数据。
机器翻译后编辑与研究数据收集常依赖低效且割裂的工作流程。我们提出TranslationCorrect,一个一体化框架,整合NLLB等模型的翻译生成、XCOMET或LLM API提供的错误预测(含详细推理)以及直观的后编辑界面。该框架基于人机交互原则设计,经用户研究验证可降低认知负荷。对译者而言,支持高效纠错与批量翻译;对研究者而言,可导出符合错误片段标注(ESA)格式的高质量细粒度标注数据,采用受多维质量度量(MQM)启发的错误分类体系。这些数据兼容先进错误检测模型,适用于训练机器翻译或后编辑系统。用户研究表明,TranslationCorrect在翻译效率和用户满意度上均显著优于传统方法。
原文摘要 · Abstract (English)
Machine translation (MT) post-editing and research data collection often rely on inefficient, disconnected workflows. We introduce TranslationCorrect, an integrated framework designed to streamline these tasks. TranslationCorrect combines MT generation using models like NLLB, automated error prediction using models like XCOMET or LLM APIs (providing detailed reasoning), and an intuitive post-editing interface within a single environment. Built with human-computer interaction (HCI) principles in mind to minimize cognitive load, as confirmed by a user study. For translators, it enables them to correct errors and batch translate efficiently. For researchers, TranslationCorrect exports high-quality span-based annotations in the Error Span Annotation (ESA) format, using an error taxonomy inspired by Multidimensional Quality Metrics (MQM). These outputs are compatible with state-of-the-art error detection models and suitable for training MT or post-editing systems. Our user study confirms that TranslationCorrect significantly improves translation efficiency and user satisfaction over traditional annotation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。