平衡能力与安全性的模型评估框架,推动大模型负责任发展。
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
- 用距离最优得分的动态距离法综合评估模型表现
- 首轮评测26个主流大模型,发现顶尖模型仍有安全短板
- 适合关注AI安全与性能均衡的研究者和开发者
为弥补这一空白,我们提出Libra-Leaderboard,一个全面评估大模型能力与安全性的平衡框架。该框架结合动态排行榜与交互式LLM竞技场,促进能力与安全的协同优化。不同于传统平均评分方式,Libra-Leaderboard采用距离最优得分的方法计算综合排名,激励模型在各维度间取得平衡,而非片面追求某一项优势。首次发布即评测了来自14家领先机构的26个主流大模型,揭示了即使是当前最先进的模型也存在显著的安全挑战。
原文摘要 · Abstract (English)
To address this gap, we introduce Libra-Leaderboard, a comprehensive framework designed to rank LLMs through a balanced evaluation of performance and safety. Combining a dynamic leaderboard with an interactive LLM arena, Libra-Leaderboard encourages the joint optimization of capability and safety. Unlike traditional approaches that average performance and safety metrics, Libra-Leaderboard uses a distance-to-optimal-score method to calculate the overall rankings. This approach incentivizes models to achieve a balance rather than excelling in one dimension at the expense of some other ones. In the first release, Libra-Leaderboard evaluates 26 mainstream LLMs from 14 leading organizations, identifying critical safety challenges even in state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。