arXiv:2509.22751stat.MLcs.AI2025-09被引 1

无真实标签时评估实体中心AI系统,用概率方法测可靠性和鲁棒性。

Variance-Bounded Evaluation of Entity-Centric AI Systems Without Ground Truth: Theory and Measurement

  • 通过约束松弛和蒙特卡洛采样生成合理解释,计算概率分布。
  • 输出评估基于期望表现,用方差惩罚反映系统稳定性。
  • 适合无真实标签的实体链接、数据集成等企业级场景。

当缺乏真实标签时,可靠评估人工智能系统仍是根本挑战,尤其在生成自然语言输出的AI聊天与代理系统中。许多此类系统聚焦于实体中心任务,在企业环境中用于实体链接、数据集成与信息检索,因数据保密而难以验证。学术研究也面临类似问题,尤其在标注标准模糊的专有数据集上。传统评估框架依赖监督学习范式,无法处理无唯一正确答案的情况。本文提出VB-Score,一种无需真实标签的方差有界评估框架,联合衡量有效性与鲁棒性。给定输入,该方法通过约束松弛与蒙特卡洛采样枚举可能解释,并赋予其概率。系统输出根据跨解释的期望成功率评估,同时以方差惩罚衡量鲁棒性。我们提供形式化理论分析,涵盖范围、单调性、稳定性及蒙特卡洛估计的集中界。在具有模糊输入的AI系统案例研究中,验证了VB-Score能揭示传统框架忽略的鲁棒性差异,为标签稀缺领域提供可信的评估基准。

原文摘要 · Abstract (English)

Reliable evaluation of AI systems remains a fundamental challenge when ground truth labels are unavailable, particularly for systems generating natural language outputs like AI chat and agent systems. Many of these AI agents and systems focus on entity-centric tasks. In enterprise contexts, organizations deploy AI systems for entity linking, data integration, and information retrieval where verification against gold standards is often infeasible due to proprietary data constraints. Academic deployments face similar challenges when evaluating AI systems on specialized datasets with ambiguous criteria. Conventional evaluation frameworks, rooted in supervised learning paradigms, fail in such scenarios where single correct answers cannot be defined. We introduce VB-Score, a variance-bounded evaluation framework for entity-centric AI systems that operates without ground truth by jointly measuring effectiveness and robustness. Given system inputs, VB-Score enumerates plausible interpretations through constraint relaxation and Monte Carlo sampling, assigning probabilities that reflect their likelihood. It then evaluates system outputs by their expected success across interpretations, penalized by variance to assess robustness of the system. We provide formal theoretical analysis establishing key properties including range, monotonicity, and stability along with concentration bounds for Monte Carlo estimation. Through case studies on AI systems with ambiguous inputs, we demonstrate that VB-Score reveals robustness differences hidden by conventional evaluation frameworks, offering a principled measurement framework for assessing AI system reliability in label-scarce domains.

AI评估无标签实体链接鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。