arXiv:2606.29857cs.LGmath.AT2026-06

用拓扑方法增强聊天机器人输入数据,提升性能且无需扩充数据集。

Comparing Chatbot Performance Enhanced with Persistent Homology

  • 用持久同调对原始数据做向量化,生成增强特征。
  • 部分场景下性能显著提升,且几乎不增加计算成本。
  • 适合隐私敏感场景,尤其本地化小数据训练时使用。

聊天机器人在多个领域广泛应用,尤其在心理健康支持方面。训练通常依赖大规模数据集,但特定领域数据可能稀缺,且涉及患者隐私,无法在共享服务器上使用个人数据进行训练。因此,仅通过增强输入数据来提升聊天机器人性能,而无需扩大数据集,具有重要意义。本文利用从原始数据中提取的持久同调(PH)向量化方法增强输入数据,并在多个指标下对比了多种聊天模型在有无PH增强情况下的表现。实验表明,尽管某些情况下提升不明显,但在部分场景中,该方法能带来显著性能改善,且几乎不增加计算开销。

原文摘要 · Abstract (English)

Chatbots have become increasingly prevalent across various domains, offering automated assistance in many areas, especially mental health support. The training is done using extremely large datasets, which are sometimes not available in very specific domains. Moreover, it would sometimes be ideal to train the chatbot with personal information about the patients, which, of course, cannot be done on shared servers since it would violate patient confidentiality. Hence, being able to improve the performance of a chatbot, possibly trained locally and on a restricted dataset, without having to increase the dataset itself, would be extremely beneficial. In this work, we will enhance the input datasets using persistent homology (PH) vectorizations computed from the raw datasets themselves. Then we will compare, across several metrics, the performance of multiple chatbot models with or without the PH enhancement. Our experiments suggest that, while at times the PH enhancement is not particularly beneficial, it sometimes brings remarkable advantages for virtually no cost.

聊天机器人拓扑分析数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。