arXiv:2607.02381cs.CL2026-07

多智能体协作生成西班牙语易读文本,效果优于传统循环重写。

HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generation

论文配图:HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generation
图 1 · 摘自论文原文
  • 基于LangGraph构建多智能体流程,结合大模型与路由机制实现并行生成。
  • 最优方案SARI达44.05,显著高于基线38.51,证明信号引导路由有效。
  • 适合关注可解释性、易读性生成的自然语言处理研究者。

本文介绍HULAT2-UC3M参与MER-TRANS 2026西班牙语易读翻译共享任务的情况,提交了三个全自动西班牙语生成运行。RUN1和RUN2采用基于LangGraph的多智能体工作流,融合Gemini 2.5 Flash与RigoChat-7B-v2,使用并行生成、内部质量信号、事件-条件-动作路由、可控编辑与可追溯决策;其中RUN1为基准流程,RUN2额外引入词表与词汇资源支持层。RUN3为基于RigoChat的提示工程与LoRA微调的生成-评估-重生成基线。官方排行榜采用BLEU-Orig、BLEU-Gold、SARI与BERTScore评估。开发阶段还分析了语义保真度、可读性、词汇简单性、句法清晰度与事实一致性等内部指标。根据官方SARI评分,RUN1以44.0543分位居第一,其次为RUN2(43.1049)与RUN3(38.5136)。结果表明,在该任务中,信号引导的多智能体路由优于线性重生成基线,但增加词汇支持并未自动提升参考指标。需进一步开展段落级与文档级分析以评估可读性、事实一致性与用户适用性。

原文摘要 · Abstract (English)

This paper describes the participation of HULAT2-UC3M in the Spanish track of MER-TRANS 2026, a shared task on multilingual Easy-to-Read translation. Three fully automatic Spanish runs were submitted. RUN1 and RUN2 used a LangGraph-based multi-agent workflow combining Gemini 2.5 Flash and RigoChat-7B-v2, parallel generation strategies, internal quality signals, Event-Condition-Action routing, controlled editing and traceable decisions. RUN1 used the base workflow, while RUN2 activated an additional lexical-support layer based on a glossary and lexical resources. RUN3 was a RigoChat-based generate-evaluate-regenerate baseline with prompt engineering and LoRA-based adaptation. The official leaderboard reports BLEU-Orig, BLEU-Gold, SARI and BERTScore. During development, additional internal signals were also inspected, including semantic fidelity, readability, lexical simplicity, syntactic clarity and factual consistency. According to official SARI, RUN1 was the best HULAT2 run, with 44.0543 points, followed by RUN2 with 43.1049 and RUN3 with 38.5136. These results indicate that, in this task setting, signal-guided multi-agent routing outperformed the linear regeneration baseline. They also show that adding lexical support did not automatically improve reference-based scores. Further segment-level and document-level analysis are required to assess readability, factual consistency and user-oriented adequacy.

多智能体易读生成自然语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。