自动化评估COBOL转Java代码质量,提升现代代码迁移可信度
Quality Evaluation of COBOL to Java Code Transformation
- 结合静态检查与大模型判题,多维度评估转换质量
- 支持持续集成,实现大规模自动评测,减少人工审查
- 适合企业级系统现代化团队使用,助力高质量代码重构
我们提出一种自动化评估系统,用于评估IBM watsonx Code Assistant for Z(WCA4Z)中COBOL到Java代码转换的质量。该系统针对基于大模型的翻译器存在的模型不可解释性与评估复杂性问题,融合静态分析检查器与大模型作为裁判(LLM-as-a-judge, LaaJ)技术,实现可扩展、多角度的评估。系统支持持续集成流程,支持大规模基准测试,并降低对人工评审的依赖。本文描述了系统架构、评估策略及报告机制,为开发者和项目管理者提供可操作的洞察,推动高质量现代代码库的演进。
原文摘要 · Abstract (English)
We present an automated evaluation system for assessing COBOL-to-Java code translation within IBM's watsonx Code Assistant for Z (WCA4Z). The system addresses key challenges in evaluating LLM-based translators, including model opacity and the complexity of translation quality assessment. Our approach combines analytic checkers with LLM-as-a-judge (LaaJ) techniques to deliver scalable, multi-faceted evaluations. The system supports continuous integration workflows, enables large-scale benchmarking, and reduces reliance on manual review. We describe the system architecture, evaluation strategies, and reporting mechanisms that provide actionable insights for developers and project managers, facilitating the evolution of high-quality, modernized codebases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。