无需模型参数即可比较大模型生成文本相似性,助力快速评估模型差异。
ConSCompF: Consistency-focused Similarity Comparison Framework for Generative Large Language Models
- 基于少量无标签指令数据,通过一致性聚焦框架计算生成文本相似度。
- 在少样本场景下表现稳定,与ROUGE-L等指标具有强相关性。
- 适合研究人员、投资者快速对比模型,识别训练数据或投资风险。
近年来,大型语言模型(LLMs)成为机器学习领域最重要的突破之一。以ChatGPT为代表的基于LLM的人工智能助手持续吸引研究者、投资者和公众关注,推动该产业快速发展。随着新模型不断涌现,区分它们变得愈发困难,亟需新型比较方法。本文提出一致性聚焦相似性比较框架(ConSCompF),用于比较两个生成式大模型的输出文本,并生成反映其响应整体相似程度的相似度分数。该框架主要优势在于仅需少量无标签数据(如聊天机器人指令提示)即可运行,且无需模型开发者披露任何内部信息。为验证其有效性,本文开展了两项实验,旨在识别多个模型间的相似性,并分析ConSCompF生成的相似度分数与其他基准技术(如ROUGE-L)输出差异之间的相关性。此外,还进行了系列少样本模型比较实验,评估ConSCompF在少样本场景下的表现。所提框架可用于构建多模型相似度矩阵,结合主成分分析(PCA)实现有效可视化。ConSCompF的输出可能揭示模型训练数据线索,帮助识别潜在的投资欺诈行为。
原文摘要 · Abstract (English)
Large language models (LLMs) have been one of the most important discoveries in machine learning in recent years. LLM-based artificial intelligence (AI) assistants, such as ChatGPT, have consistently attracted the attention from researchers, investors, and the general public, driving the rapid growth of this industry. With the frequent introduction of new LLMs to the market, it becomes increasingly difficult to differentiate between them, creating a demand for new LLM comparison methods. In this research, the Consistency-focused Similarity Comparison Framework (ConSCompF) for generative large language models is proposed. It compares texts generated by two LLMs and produces a similarity score, indicating the overall degree of similarity between their responses. The main advantage of this framework is that it can operate on a small number of unlabeled data, such as chatbot instruction prompts, and does not require LLM developers to disclose any information about their product. To evaluate the efficacy of ConSCompF, two experiments aimed at identifying similarities between multiple LLMs are conducted. Additionally, these experiments examine the correlation between the similarity scores generated by ConSCompF and the differences in the outputs produced by other benchmarking techniques, such as ROUGE-L. Finally, a series of few-shot LLM comparison experiments is conducted to evaluate the performance of ConSCompF in a few-shot LLM comparison scenario. The proposed framework can be used for calculating similarity matrices of multiple LLMs, which can be effectively visualized using principal component analysis (PCA). The ConSCompF output may provide useful insights into data that might have been used during LLM training and help detect possible investment fraud attempts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。