用自然语言找仿真模型,验证了数据格式和检索策略的关键作用。
How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies
- 用嵌入模型将模型与自然语言对齐,实现语义级搜索
- 开源嵌入模型在召回率@5上达82.3%,性能接近商业方案
- 复杂查询下重排序机制显著提升检索效果,适合系统开发者
在建模与仿真(M&S)领域,发现可复用的仿真模型仍面临根本性挑战。当大量模型共存时,仅凭建模意图定位匹配模型依然困难。近年来,以检索为基础的人工智能技术为解决该问题提供了新路径。本文通过实验研究数据表示方式、基于Transformer的嵌入模型及检索策略对自然语言查询下仿真模型发现的影响。采用标准信息检索指标(如recall@5、nDCG@5)评估多种查询类型的表现。结果表明:数据表示方式显著影响性能;开源嵌入模型可达到82.3%的recall@5,表现优异;重排序方法在复杂查询中尤为关键。本研究为人工智能驱动的模型发现提供了基准,并探讨其在推动模型可组合性与互操作性中的作用。
原文摘要 · Abstract (English)
Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identifying those that align with a given modeling intent remains difficult. Recent advances in Artificial Intelligence (AI), particularly retrieval-based approaches, offer a promising pathway to operate at this semantic layer. In this paper, we present an experimental study investigating the impact of data representation, transformer-based embedding models, and retrieval strategies on the discovery of simulation models using natural language queries. We evaluated performance across multiple query types using standard information retrieval metrics, including recall@5 and nDCG@5. Results show that data representation matters, open-source embedding models can achieve high performance, and reranking methods are important, especially as query complexity increases. This work provides a baseline for AI-driven model discovery and discusses its role in advancing toward AI-driven composability and interoperability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。