用大模型实现无需参考文本的多维度翻译质量评估
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
- 通过提示工程驱动大模型,不依赖参考译文
- 可量化错误类型与编辑难度,支持多维度评分
- 适合研究者、开发者和专业译者快速评估翻译质量
本文提出CATER(综合人工智能辅助翻译编辑率),一种全提示驱动的机器翻译质量评估框架。借助大语言模型(LLMs)与精心设计的提示协议,CATER突破传统依赖参考译文的限制,实现多维度、无参考的评估,涵盖语言准确性、语义保真度、上下文连贯性、风格恰当性和信息完整性。只需提供源文本和目标文本及标准提示,即可由大模型快速识别错误、量化编辑成本,并生成类别级与总体评分。该方法无需预设参考译文或领域专用资源,可通过调整权重和提示灵活适配多种语言、文体与用户需求。其基于大模型的策略能捕捉细微遗漏、幻觉和话语层面偏移等挑战性现象。结合MQM和DQF的理论严谨性与大模型的可扩展性,CATER为全球研究人员、开发者和专业译者提供高效评估工具。框架与示例提示已开源,鼓励社区共建与实证验证。
原文摘要 · Abstract (English)
This paper introduces the Comprehensive AI-assisted Translation Edit Ratio (CATER), a novel and fully prompt-driven framework for evaluating machine translation (MT) quality. Leveraging large language models (LLMs) via a carefully designed prompt-based protocol, CATER expands beyond traditional reference-bound metrics, offering a multidimensional, reference-independent evaluation that addresses linguistic accuracy, semantic fidelity, contextual coherence, stylistic appropriateness, and information completeness. CATER's unique advantage lies in its immediate implementability: by providing the source and target texts along with a standardized prompt, an LLM can rapidly identify errors, quantify edit effort, and produce category-level and overall scores. This approach eliminates the need for pre-computed references or domain-specific resources, enabling instant adaptation to diverse languages, genres, and user priorities through adjustable weights and prompt modifications. CATER's LLM-enabled strategy supports more nuanced assessments, capturing phenomena such as subtle omissions, hallucinations, and discourse-level shifts that increasingly challenge contemporary MT systems. By uniting the conceptual rigor of frameworks like MQM and DQF with the scalability and flexibility of LLM-based evaluation, CATER emerges as a valuable tool for researchers, developers, and professional translators worldwide. The framework and example prompts are openly available, encouraging community-driven refinement and further empirical validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。