STEMM Institute Press
Science, Technology, Engineering, Management and Medicine
Design and Implementation of a Chinese Semantic Matching Model Based on Siamese LSTM
DOI: https://doi.org/10.62517/jike.202604317
Author(s)
Lin Liu1, Wanyue Liu2, Xiaoying Liu1
Affiliation(s)
1School of Artificial Intelligence and Big Data, Henan University of Technology, Zhengzhou, China 2iFLYTEK Co., Ltd., Hefei, China
Abstract
To address the limited semantic discrimination of term-frequency-based methods and the high deployment cost of large pretrained models, this paper develops a Chinese semantic matching model that combines a Siamese architecture, bidirectional long short-term memory networks, and attention pooling. Two weight-sharing encoders project a sentence pair into the same semantic space. Contextual representations are extracted by a BiLSTM, salient tokens are aggregated by an attention mechanism, and the two sentence vectors are compared through concatenation, absolute difference, and element-wise product before binary classification. Experiments on the LCQMC corpus show that the proposed model achieves 82.48% accuracy, 76.84% precision, 92.98% recall, 84.14% F1-score, and 92.79% AUC. Compared with the TF-IDF baseline, accuracy and F1-score improve by 26.38 and 15.78 percentage points, respectively. With about 4.8 million parameters, the model uses only approximately one twenty-third of the parameters of BERT-base. The model is further integrated into the knowledge retrieval pipeline of the Zhixue learning platform to rank candidate passages and return Top-5 results, demonstrating a favorable trade-off among accuracy, inference efficiency, and deployability.
Keywords
Chinese Text Matching; Siamese Model; BiLSTM; Attention; Knowledge Retrieval
References
[1]Manning C D, Raghavan P, Schütze H. Introduction to Information Retrieval. Cambridge: Cambridge University Press, 2008. [2]Devlin J, Chang M W, Lee K, et al. BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT. Minneapolis: ACL, 2019: 4171-4186. [3]Mueller J, Thyagarajan A. Siamese recurrent architectures for learning sentence similarity. Proceedings of the 30th AAAI Conference on Artificial Intelligence. Phoenix: AAAI Press, 2016: 2786-2792. [4]Liu X, Chen Q, Deng C, et al. LCQMC: A large-scale Chinese question matching corpus. Proceedings of the 27th International Conference on Computational Linguistics. Santa Fe: ACL, 2018: 1952-1962. [5]Hochreiter S, Schmidhuber J. Long short-term memory. Neural Computation, 1997, 9(8): 1735-1780. [6]Bahdanau D, Cho K, Bengio Y. Neural machine translation by jointly learning to align and translate. International Conference on Learning Representations. San Diego, 2015. [7]Meng J X, Shan H T, Wan J J, et al. BSLA: An improved Siamese-LSTM text similarity model. Computer Engineering and Applications, 2022, 58(23): 178-185. (in Chinese). [8]Li Y L, Zhou Y P. Text similarity matching based on a Siamese network and combined character-word vectors. Computer Systems & Applications, 2022, 31(10): 295-302. (in Chinese). [9]Muennighoff N, Tazi N, Magne L, et al. MTEB: Massive text embedding benchmark. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, 2023: 2014-2037. doi:10.18653/v1/2023.eacl-main.148. [10]Xiao S, Liu Z, Zhang P, et al. C-Pack: Packed resources for general Chinese embeddings. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024: 641-649. doi:10.1145/3626772.3657878. [11]Chen J, Xiao S, Zhang P, et al. M3-Embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. Findings of the Association for Computational Linguistics: ACL 2024, 2024: 2318-2335. doi:10.18653/v1/2024.findings-acl.137. [12]Wang L, Yang N, Huang X, et al. Improving text embeddings with large language models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2024: 11897-11916. doi:10.18653/v1/2024.acl-long.642. [13]Kingma D P, Ba J. Adam: A method for stochastic optimization. International Conference on Learning Representations. San Diego, 2015
Copyright @ 2020-2035 STEMM Institute Press All Rights Reserved