Digital Islamic Humanities

Digital Islamic Humanities

Monolingual and cross-lingual semantic similarity detection of Arabic texts using deep learning

Document Type : Original Article

Authors
1 PhD student in Artificial Intelligence - Faculty of Computer Engineering, Iran University of Science and Technology (IUST)
2 Professor at the Faculty of Computer Engineering, Iran University of Science and Technology (IUST)
Abstract
Semantic Textual Similarity (STS) is a crucial subfield of Natural Language Processing (NLP) that has garnered significant attention in recent years. STS aims to compute the degree of semantic similarity between two textual documents, paragraphs, or sentences, either monolingually or cross-lingually. In this paper, we focus on calculating the semantic similarity between two sentences in Arabic and Arabic-English cross-lingual pairs. Given the prevalence of Arabic texts in Islamic literature, this research has numerous practical applications. The semantic similarity between two sentences can be determined using their semantic vectors. To compute this similarity, we first need to represent each sentence as a vector. 
In this study, word vectors were extracted using pre-trained embeddings on Arabic texts from Twitter and Wikipedia, employing two well-known word embedding techniques: CBOW and Skip-Gram. Additionally, transformer-based models such as paraphrase-xlm-roberta were utilized for cross-lingual semantic similarity calculation between Arabic and English. To evaluate and train the model, we used data from the Semantic Textual Similarity Conference of 2017, which includes Arabic-Arabic and Arabic-English sentence pairs. A deep neural network model, specifically a Siamese network with an LSTM layer, was trained. LSTM enables the network to learn long-term dependencies. Siamese networks, despite their simplicity, yield satisfactory results, while transformer-based models demonstrate strong cross-lingual learning capabilities.
In the final layer of the network, the cosine similarity between the vectors of the two input sentences is used to determine their degree of similarity. The results indicate that the proposed method achieves a Pearson correlation of 83.4% for Arabic-Arabic sentence pairs and 82% for Arabic-English pairs, outperforming other existing approaches.

Highlights

  1. Agirre, Eneko, et al. "Semeval-2012 task 6: A pilot on semantic textual similarity." * SEM 2012: The First Joint Conference on Lexical and Computational Semantics–Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012). 2012.
  2. Agirre, Eneko, et al. "Semeval-2014 task 10: Multilingual semantic textual similarity." Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014). 2014.
  3. Agirre, Eneko, et al. "* SEM 2013 shared task: Semantic textual similarity." Second Joint Conference on Lexical and Computational Semantics (* SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity. 2013.
  4. Bar D., Biemann C., Gurevych I., and Zesch T. Ukp: Computing semantic textual similarity by combining multiple content similarity measures. Proceedings of the 6th International Workshop on Semantic Evaluation, in conjunction with the 1st Joint Conference on Lexical and Computational Semantics, 2012.
  5. Bjerva, Johannes, and Robert Östling. "Cross-lingual learning of semantic textual similarity with multilingual word representations." 21st Nordic Conference on Computational Linguistics, NoDaLiDa, Gothenburg, Sweden, 22-24 May, 2017. Linköping University Electronic Press, 2017.
  6. Brychcín, Tomáš. "Linear transformations for cross-lingual semantic textual similarity." Knowledge-Based Systems 187 (2020): 104819.
  7. Cer, Daniel, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. "SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation." In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). Association for Computational Linguistics, 2017.Comelles, Elisabet, and Jordi Atserias. "VERTa: a linguistic approach to automatic machine translation evaluation." Language Resources and Evaluation1 (2019): 57-86.
  8. Dagan, Ido, Oren Glickman, and Bernardo Magnini. "The PASCAL recognising textual entailment challenge." Machine Learning Challenges Workshop. Springer, Berlin, Heidelberg, 2005.
  9. Das, Arijit, and Diganta Saha. "Deep learning based Bengali question answering system using semantic textual similarity." Multimedia Tools and Applications (2022): 1-25.
  10. Han, Lushan, et al. "UMBC_EBIQUITY-CORE: Semantic textual similarity systems." Second Joint Conference on Lexical and Computational Semantics (* SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity. 2013.
  11. Hochreiter, Sepp, and Jürgen Schmidhuber. "Long short-term memory." Neural computation8 (1997): 1735-1780.
  12. Islam, Aminul, and Diana Inkpen. "Semantic text similarity using corpus-based word similarity and string similarity." ACM Transactions on Knowledge Discovery from Data (TKDD) 2.2 (2008): 1-25.
  13. Lubani, Mohamed, and Shahrul Azman Mohd Noah. "Text Relation Extraction Using Sentence-Relation Semantic Similarity." In Multi-disciplinary Trends in Artificial Intelligence: 13th International Conference, MIWAI 2019, Kuala Lumpur, Malaysia, November 17–19, 2019, Proceedings 13, pp. 3-14. Springer International Publishing, 2019.
  14. Mihalcea, Rada, Courtney Corley, and Carlo Strapparava. "Corpus-based and knowledge-based measures of text semantic similarity." Aaai. Vol. 6. No. 2006. 2006.
  15. Mikolov, Tomas, et al. "Distributed representations of words and phrases and their compositionality." Advances in neural information processing systems. 2013.
  16. Mueller, J., & Thyagarajan, A. (2016, March). Siamese recurrent architectures for learning sentence similarity. In thirtieth AAAI conference on artificial intelligence.
  17. Roul, Rajendra Kumar, and Jajati Keshari Sahoo. "Near-duplicate document detection using semantic-based similarity measure: a novel approach." In Computational Intelligence in Data Mining: Proceedings of the International Conference on ICCIDM 2018, pp. 543-558. Springer Singapore, 2020.
  18. Rychalska B., Pakulska K., Chodorowska K., Walczak W., and Andruszkiewicz P. Samsung Poland NLP Team at SemEval-2016 Task 1: Necessity for diversity; combining recursive autoencoders, wordnet and ensemble methods to measure semantic similarity. Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval 2016), San Diego, CA, USA, 2016.
  19. Schwab, Didier. "Semantic similarity of arabic sentences with word embeddings." 2017.
  20. Shahmirzadi, Omid, Adam Lugowski, and Kenneth Younge. "Text similarity in vector space models: a comparative study." 2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA). IEEE, 2019.
  21. Sultan M.A., Bethard S., and Sumner T. Back to basics for monolingual alignment: Exploiting word similarity and contextual evidence. Transactions of the Association for Computational Linguistics, 2:219–230, 2014a.
  22. Sultan M.A., Bethard S., and Sumner T. DLS@CU: Sentence similarity from word alignment. Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), 241–246, Dublin, Ireland, August 2014b. Association for Computational Ling
  23. Soliman, Abu Bakr, Kareem Eissa, and Samhaa R. El-Beltagy. "Aravec: A set of arabic word embedding models for use in arabic nlp." Procedia Computer Science 117 (2017): 256-265.
  24. Suleiman, Dima, Arafat Awajan, and Nailah Al-Madi. "Deep Learning Based Technique for Plagiarism Detection in Arabic Texts." 2017 International Conference on New Trends in Computing Sciences (ICTCS). IEEE, 2017.
  25. Tian, Junfeng, et al. "Ecnu at semeval-2017 task 1: Leverage kernel-based traditional nlp features and neural networks to build a universal model for multilingual and cross-lingual semantic textual similarity." Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). 2017.
  26. Wu, Hao, et al. "BIT at SemEval-2017 Task 1: Using semantic information space to evaluate semantic textual similarity." Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). 2017.

Keywords