arXiv:2603.24774cs.SEcs.AI2026-03中稿 · publication at IEE…被引 1

用测试关系代替真实答案,让大模型系统可测

From Untestable to Testable: Metamorphic Testing in the Age of LLMs

  • 用多个测试用例间的逻辑关系做判断标准
  • 无需真实标签也能发现大模型功能异常
  • 适合验证大模型集成系统的可靠性

本文探讨了集成AI与大语言模型(LLM)功能的软件系统在测试中面临的挑战。大语言模型虽强大但不可靠,且用于测试的带标注真实答案难以规模化。元测试(Metamorphic Testing)通过将多次测试执行之间的关系转化为可执行的测试断言,解决了这一难题。

原文摘要 · Abstract (English)

This article discusses the challenges of testing software systems with increasingly integrated AI and LLM functionalities. LLMs are powerful but unreliable, and labeled ground truth for testing rarely scales. Metamorphic Testing solves this by turning relations among multiple test executions into executable test oracles.

大模型测试元测试AI可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。