首个标准化评估框架,可对比自动化领域建模方法效果
Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark

- 整合45+8条参考模型,构建多层级评估数据集
- 支持自然语言到领域模型的生成效果量化比较
- 符合开放科学规范,适合研究者复现与对比
领域建模在领域驱动设计中至关重要,用于捕捉特定领域中的核心实体及其关系。尽管自动化领域建模取得进展,但缺乏标准化评估基准,制约了不同方法的比较。本文提出一个基准,融合45条来自Golden UML Modelset(Verbruggen等,2025)的记录(由Calamo等,2025年文本转UML项目提供),以及8条来自Chen等(2023a,b)的参考档案,实现对不同复杂度和规模下的自动化建模方法的评估。任务是根据自然语言描述生成对应领域模型,每条输入均有真实参考模型作为标准。采用度量方法比较生成模型与真实模型的差异。为验证基准效用,评估了多种方法,包括启发式规则方法和大模型驱动策略。依据FAIR4RS建议(Chue Hong等,2022),该基准以研究工具形式发布,支持未来研究复用与扩展。
原文摘要 · Abstract (English)
Domain modeling plays an essential role in domain-driven design, capturing essential entities and their relationships within a specific domain. Despite advancements in automated domain modeling, the absence of standardized benchmarks has hindered the comparative assessment of existing approaches. This paper introduces a benchmark designed to address this gap. The benchmark combines the 45-record Golden UML Modelset (Verbruggen et al., 2025) on Zenodo, as distributed by the Text2UML project of Calamo, Mecella, and Snoeck (Calamo et al., 2025), with the 8-record reference archive of Chen et al. (Chen et al., 2023a,b), enabling the evaluation of automated domain modeling approaches across different levels of complexity and scale. Given a natural language description, the task is to generate a corresponding domain model. For each description, a reference domain model is provided as ground truth. A metric is used to compare the generated domain model with the corresponding ground-truth model. To demonstrate the utility of the benchmark, we evaluate multiple automated domain modeling approaches, including heuristic rule-based methods and LLM-driven strategies. In accordance with the FAIR4RS recommendations (Chue Hong et al., 2022), the benchmark is provided as a research artifact to encourage reuse and support future research on automated domain modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。