提出统一评估模型完整性和责任性的评分框架,让复杂模型不再盲目胜出。
Multi-Dimensional Model Integrity and Responsibility Assessment Index and Scoring Framework
- 构建多维度评估框架,整合可解释性、公平性等五项指标
- 实测显示高精度模型未必整体更负责任,简单模型反而更均衡
- 适合监管场景下负责任的模型选型,尤其医疗金融领域
高风险表格数据领域的人工智能不能仅靠预测性能评估,但当前实践仍孤立评价可解释性、公平性、鲁棒性、隐私和可持续性。本文提出模型完整性与责任评估指数(MIRAI),一个统一评估框架,在受控比较环境下衡量表格模型在上述五个维度的表现,并将其聚合为单一得分。MIRAI通过归一化与方向对齐的维度分数,结合已有指标,实现不同架构和计算配置模型间的直接比较。在医疗、金融及社会经济数据集上的实验表明,更高的预测性能并不意味着更强的整体完整性与责任性;在某些情况下,更简单的模型比复杂的深度表格架构展现出更好的跨维度平衡。MIRAI为受监管环境中负责任的模型选择提供了简洁实用的依据。
原文摘要 · Abstract (English)
Artificial intelligence in high-stakes tabular domains cannot be evaluated by predictive performance alone, yet current practice still assesses explainability, fairness, robustness, privacy, and sustainability mostly in isolation. We propose the Model Integrity and Responsibility Assessment Index (MIRAI), a unified evaluation framework that measures tabular models across these five dimensions under a controlled comparison setting and aggregates them into a single score. MIRAI combines established metrics through normalized and direction-aligned dimension scores, which enables direct comparison across models with different architectural and computational profiles. Experiments on healthcare, financial, and socioeconomic datasets show that higher predictive performance does not necessarily imply better overall integrity and responsibility. In several cases, simpler models achieve a stronger cross-dimensional balance than more complex deep tabular architectures. MIRAI provides a compact and practical basis for responsible model selection in regulated settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。