多智能体系统用质量维度精准纠错,提升翻译准确性与上下文适配。
MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation
- 分角色智能体按质量维度(如准确、流畅)分工诊断错误
- 在多种语言对上显著优于单智能体基线,尤其在语义准确性和远距离语言对表现突出
- 适合需要高精度、可解释性翻译的场景,如本地化和专业文档
我们提出MAATS,一种基于多维质量度量(MQM)框架的多智能体自动化翻译系统。该系统通过多个专注于不同MQM类别(如准确性、流畅性、风格、术语)的专用智能体进行错误检测与修正,再由合成智能体整合标注结果实现迭代优化。相比依赖自我修正的单智能体方法,MAATS在多种语言对和大语言模型(LLMs)上均取得统计显著的提升,不仅在自动评估指标上表现更优,在人工评估中也更具优势。其在语义准确性、本地化适配以及语言差异较大的翻译任务中尤为突出。定性分析显示,系统具备多层次错误诊断能力,能识别跨视角遗漏,并实现上下文感知的精细化修正。通过将模块化智能体角色与可解释的MQM维度对齐,MAATS缩小了黑箱大模型与人类翻译流程之间的差距,推动翻译重点从表面流畅转向深层语义与上下文一致性。
原文摘要 · Abstract (English)
We present MAATS, a Multi Agent Automated Translation System that leverages the Multidimensional Quality Metrics (MQM) framework as a fine-grained signal for error detection and refinement. MAATS employs multiple specialized AI agents, each focused on a distinct MQM category (e.g., Accuracy, Fluency, Style, Terminology), followed by a synthesis agent that integrates the annotations to iteratively refine translations. This design contrasts with conventional single-agent methods that rely on self-correction. Evaluated across diverse language pairs and Large Language Models (LLMs), MAATS outperforms zero-shot and single-agent baselines with statistically significant gains in both automatic metrics and human assessments. It excels particularly in semantic accuracy, locale adaptation, and linguistically distant language pairs. Qualitative analysis highlights its strengths in multi-layered error diagnosis, omission detection across perspectives, and context-aware refinement. By aligning modular agent roles with interpretable MQM dimensions, MAATS narrows the gap between black-box LLMs and human translation workflows, shifting focus from surface fluency to deeper semantic and contextual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。