arXiv:2601.19214cs.CLcs.AI2026-01Conference of the …

用混合模型从用户评论中精准提取可操作改进建议

A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews

  • 结合罗伯特分类器与指令微调大模型,提升建议识别召回率
  • 在酒店和餐饮数据集上准确率优于基线方法,聚类结果更清晰
  • 适合需要精细化用户反馈分析的企业决策场景

从客户评论中提取可操作建议对运营决策至关重要,但这些指令常隐藏在混合意图的非结构化文本中。现有方法要么分类含建议句子,要么生成高层摘要,却很少精准提取企业所需的改进指令。本文评估一种混合流水线:先用高召回率的RoBERTa分类器(基于精确率-召回率代理训练)减少无法恢复的假负例,再通过受控的指令微调大模型完成建议提取、分类、聚类与摘要。在真实世界的酒店与食品数据集上,该混合系统在提取准确率与聚类一致性上均优于仅用提示、规则或分类器的基线方法。人工评估进一步证实,生成的建议与摘要清晰、忠实且可解释。总体表明,混合推理架构在细粒度可操作建议挖掘中取得显著提升,同时揭示了领域适配与高效本地部署的挑战。

原文摘要 · Abstract (English)

Extracting actionable suggestions from customer reviews is essential for operational decision-making, yet these directives are often embedded within mixed-intent, unstructured text. Existing approaches either classify suggestion-bearing sentences or generate high-level summaries, but rarely isolate the precise improvement instructions businesses need. We evaluate a hybrid pipeline combining a high-recall RoBERTa classifier trained with a precision-recall surrogate to reduce unrecoverable false negatives with a controlled, instruction-tuned LLM for suggestion extraction, categorization, clustering, and summarization. Across real-world hospitality and food datasets, the hybrid system outperforms prompt-only, rule-based, and classifier-only baselines in extraction accuracy and cluster coherence. Human evaluations further confirm that the resulting suggestions and summaries are clear, faithful, and interpretable. Overall, our results show that hybrid reasoning architectures achieve meaningful improvements fine-grained actionable suggestion mining while highlighting challenges in domain adaptation and efficient local deployment.

自然语言处理客户反馈大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。