arXiv:2603.03311cs.CL2026-03

记录一个运行30年的中英日翻译系统,展示规则引擎如何应对实际场景挑战。

The Logovista English-Japanese Machine Translation System

  • 用手工编写规则+大型词典+加权解析处理语言歧义
  • 系统从1990年代持续运行至2012年,支持大规模真实应用
  • 适合研究传统规则系统长期维护与演化策略的学者

本文记录了Logovista中英日机器翻译系统的架构、开发实践及保存下来的资源。该系统是一个大型显式规则驱动的机器翻译系统,自1990年代初起商业化开发并持续运营至至少2012年。系统结合手工编写的语法规则、包含句法与语义约束的大规模中央词典,以及基于图表的解析与加权解释评分机制,以应对复杂的结构歧义问题。文章重点阐述了系统在真实使用压力下的持续扩展与维护过程,包括回归控制、歧义管理,以及覆盖范围扩大时遇到的限制。与多数仅在研究环境中描述的规则系统不同,Logovista在实际应用中运行了数十年,并不断根据现实需求演进。本文旨在提供技术与历史记录,而非倡导复兴规则系统,同时介绍了已保存的软件与语言资源,供未来研究参考。

原文摘要 · Abstract (English)

This paper documents the architecture, development practices, and preserved artifacts of the Logovista English--Japanese machine translation system, a large, explicitly rule-based MT system that was developed and sold commercially from the early 1990s through at least 2012. The system combined hand-authored grammatical rules, a large central dictionary encoding syntactic and semantic constraints, and chart-based parsing with weighted interpretation scoring to manage extensive structural ambiguity. The account emphasizes how the system was extended and maintained under real-world usage pressures, including regression control, ambiguity management, and the limits encountered as coverage expanded. Unlike many rule-based MT systems described primarily in research settings, Logovista was deployed for decades and evolved continuously in response to practical requirements. The paper is intended as a technical and historical record rather than an argument for reviving rule-based MT, and describes the software and linguistic resources that have been preserved for potential future study.

机器翻译规则系统语言工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。