用确定性框架提升大模型文本分类精度与可复现性
Leveraging Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) Towards Text Categorization Using LLMs

- 构建分层结构+信噪比评分,筛选高价值语义特征
- 在多个数据集上降低分类熵,提升聚类完整性
- 适合需要稳定、可复现文本分析的工业级应用
大型语言模型(LLMs)在企业级文本分类等可靠分析中常受注意力机制随机性及对噪声敏感性影响,导致精度与可复现性下降。本文提出加权句法与语义上下文评估摘要(wSSAS),一种确定性框架,用于保障大规模混乱数据集的数据完整性。该框架采用两阶段验证:首先将原始文本组织为包含主题、故事和聚类的分层分类结构;随后利用信噪比(SNR)优先筛选高价值语义特征,确保模型注意力聚焦于代表性数据点。通过将评分机制融入摘要之摘要(SoS)架构,有效隔离关键信息并抑制聚合过程中的背景噪声。在Gemini 2.0 Flash Lite下,针对Google商业评论、Amazon产品评论和Goodreads书籍评论等多个数据集的实验表明,wSSAS显著提升聚类完整性和分类准确率,降低分类熵,提供基于高精度、确定性流程的大规模文本分类可复现路径。
原文摘要 · Abstract (English)
The use of Large Language Models (LLMs) for reliable, enterprise-grade analytics such as text categorization is often hindered by the stochastic nature of attention mechanisms and sensitivity to noise that compromise their analytical precision and reproducibility. To address these technical frictions, this paper introduces the Weighted Syntactic and Semantic Context Assessment Summary (wSSAS), a deterministic framework designed to enforce data integrity on large-scale, chaotic datasets. We propose a two-phased validation framework that first organizes raw text into a hierarchical classification structure containing Themes, Stories, and Clusters. It then leverages a Signal-to-Noise Ratio (SNR) to prioritize high-value semantic features, ensuring the model's attention remains focused on the most representative data points. By incorporating this scoring mechanism into a Summary-of-Summaries (SoS) architecture, the framework effectively isolates essential information and mitigates background noise during data aggregation. Experimental results using Gemini 2.0 Flash Lite across diverse datasets - including Google Business reviews, Amazon Product reviews, and Goodreads Book reviews - demonstrate that wSSAS significantly improves clustering integrity and categorization accuracy. Our findings indicate that wSSAS reduces categorization entropy and provides a reproducible pathway for improving LLM based summaries based on a high-precision, deterministic process for large-scale text categorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。