arXiv:2510.13106cs.SEcs.AI2025-10

TRUSTVIS自动化评估大模型可信度,支持交互式可视化分析

TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models

  • 融合扰动方法与多数投票,实现多维度可信度评估
  • 在Vicuna-7b等模型上识别出安全与鲁棒性漏洞
  • 交互界面直观展示结果,适合研究人员优化模型

随着大型语言模型(LLMs)持续革新自然语言处理应用,其安全性与鲁棒性方面的可信度问题仍备受关注。为此,我们提出TRUSTVIS——一个自动化的可信度评估框架,提供对LLM可信度的全面测评。该框架的核心特点是交互式用户界面,能够直观呈现可信度指标的可视化结果。通过集成AutoDAN等知名扰动方法,并结合多种评估方法的多数投票机制,TRUSTVIS不仅保证了评估结果的可靠性,也使复杂的评估过程对用户更易用。对Vicuna-7b、Llama2-7b和GPT-3.5的初步案例研究证明,该框架能有效发现模型的安全与鲁棒性缺陷;交互式界面则支持用户深入探索结果,助力针对性改进。视频链接:https://youtu.be/k1TrBqNVg8g

原文摘要 · Abstract (English)

As Large Language Models (LLMs) continue to revolutionize Natural Language Processing (NLP) applications, critical concerns about their trustworthiness persist, particularly in safety and robustness. To address these challenges, we introduce TRUSTVIS, an automated evaluation framework that provides a comprehensive assessment of LLM trustworthiness. A key feature of our framework is its interactive user interface, designed to offer intuitive visualizations of trustworthiness metrics. By integrating well-known perturbation methods like AutoDAN and employing majority voting across various evaluation methods, TRUSTVIS not only provides reliable results but also makes complex evaluation processes accessible to users. Preliminary case studies on models like Vicuna-7b, Llama2-7b, and GPT-3.5 demonstrate the effectiveness of our framework in identifying safety and robustness vulnerabilities, while the interactive interface allows users to explore results in detail, empowering targeted model improvements. Video Link: https://youtu.be/k1TrBqNVg8g

大模型评估可信度交互可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。