arXiv:2608.07439cs.CLquant-ph2026-08

用大模型重写金融句子,让量子语法分析更高效准确

An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis

论文配图:An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis
图 1 · 摘自论文原文
  • 用大模型控制重写金融句,简化结构以适配量子语法分析
  • 最优方案使量子比特和门数减少70%以上,准确率达55.0%
  • 提示词设计与过滤策略影响效果,适合做量子自然语言处理的开发者

量子自然语言处理(QNLP)提供了一种语法感知的文本建模框架,其中分布组合范畴(DisCoCat)是其理论基础之一。以往研究发现,金融情感分析中使用DisCoCat存在解析器敏感、仿真成本高及难以处理长句等问题。本文研究一种基于大模型的预处理流程,通过受控重写将中等复杂度的金融情感句压缩、简化或分解为解析兼容、电路高效的变体,同时保留情感语义。对比不同提示策略、语言模型和过滤配置,与Stein等人提出的仅限低复杂度输入的基准相比,在电路层面,最优压缩方案使平均量子比特和门数减少超过70%。在多次训练中,GPT-4.1-mini配合提示B取得最高均值准确率0.550±0.035,优于基准的0.521±0.050。更大的训练集不必然提升性能,训练集规模与准确率呈中度负相关(Pearson r = -0.446)。结果表明,大模型辅助重写可使部分中等复杂度输入适用于当前的DisCoCat设置,同时凸显提示设计、过滤机制与电路感知预处理的重要性。

原文摘要 · Abstract (English)

Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat) is one of its theoretically grounded formulations. Prior work on financial sentiment analysis has identified practical limitations of DisCoCat, including parser sensitivity, high simulation cost, and difficulty handling longer sentences. We study an LLM-assisted preprocessing workflow that uses controlled rewriting to compress, simplify, or decompose moderate-complexity financial sentiment sentences into parser-compatible, circuit-efficient variants while preserving sentiment-bearing meaning. We compare prompting strategies, language models, and filtering configurations with the low-complexity-only DisCoCat baseline of Stein et al. At the circuit level, the strongest compression variants reduce average qubit and gate counts by more than 70 percent relative to the raw moderate-complexity subset. Across repeated training runs, GPT-4.1-mini with Prompt B achieves the highest observed mean accuracy, $0.550 \pm 0.035$, compared with $0.521 \pm 0.050$ for the baseline. Larger training splits do not necessarily improve downstream performance; across evaluated configurations, training-split size has a moderately negative association with accuracy (Pearson $r=-0.446$). These results provide exploratory evidence that LLM-assisted rewriting can make some moderate-complexity inputs usable within the evaluated DisCoCat configuration, while highlighting prompt design, filtering, and circuit-aware preprocessing as considerations for more scalable QNLP-based financial sentiment analysis.

量子自然语言金融情感分析大模型预处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。