arXiv:2605.20193cs.CLcs.AI2026-05中稿 · publish in 12th In…

用多轮提示验证提升低比特量化模型在定性分析中的准确性。

Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification

论文配图:Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
图 1 · 摘自论文原文
  • 设计多轮提示验证机制,逐步剔除不可靠内容,减少幻觉。
  • 8比特模型最接近人工标注标准,4比特经优化后稳定性提升。
  • 适合资源有限但需高可靠性的定性研究团队使用。

量化大语言模型(LLMs)因速度快、资源消耗少,常用于定性分析。本研究考察了8比特、4比特、3比特和2比特不同量化级别及类型对LLaMA-3.1(8B)在定性分析中的影响。基于82份访谈转录文本,专家与非专家语料显示,低比特模型易产生幻觉和结果不稳定,尤其在处理术语模糊的非专家语言时。为此,提出一种量化感知的多轮提示验证方法:通过受控步骤引导模型,剔除不可靠内容,验证后传递至下一轮,提升准确性。以人工编码(NVivo)与半精度(BF16)LLaMA输出为依据,构建主题提取与频率分析的金标准(GSGT)。结果显示,8比特模型最接近GSGT;4比特模型虽准确率下降,但应用该方法后趋于稳定;3比特与2比特因压缩过重性能下降,但仍可通过优化提示设计与验证得到改善。同比特位下,不同量化类型表现差异显著。整体表明,该方法使低资源模型更稳定、准确,适用于低成本定性研究。

原文摘要 · Abstract (English)

Quantized Large Language Models (LLMs) are used more often in qualitative analysis because they run fast and need fewer computing resources. This study examines how different lower bits quantization levels (8-bit, 4-bit, 3-bit, and 2-bit) and quantization types affect the performance of LLaMA-3.1 (8B) on qualitative analysis. The study uses expert and non-expert responses from 82 interview transcripts. Low-bit models often produce higher levels of hallucinations and unstable results, especially when reading non-expert language with unclear terms. To improve performance, we propose a quantization-aware multi-pass prompt verification method. This method guides the model through controlled steps that reduce hallucinations. It removes unreliable content and passes the results to the next transcript after verification, improving accuracy. To validate performance, human coders analyzed transcripts using NVivo and BF16 LLaMA. BF16 LLaMA-3.1 produced high-precision output but had semantic drift and hallucination. These errors were corrected manually. The corrected BF16 output and NVivo human coding were combined to create a gold-standard ground truth (GSGT) for thematic extraction and frequency analysis. The results show that 8-bit models stay closest to the GSGT. The 4-bit models lose accuracy but become stable when the proposed method is applied. The 3-bit and 2-bit models drop in performance because of heavy compression, but they improve with the proposed prompt design and verification. The study also finds that models at the same bit level behave differently depending on quantization type. Overall, the method helps low-resource LLMs become more stable, accurate, and suitable for qualitative research at lower cost.

量化大模型定性分析提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。