شباهتیابی معنایی متون (STS) از زیرشاخههای مهم پردازش زبان طبیعی است که توجهات زیادی را در چند سال اخیر به خود معطوف کرده است. در شباهتیابی معنایی هدف محاسبه میزان شباهت معنایی بین دو سند متنی، پاراگراف یا جمله است که به دو صورت تک زبانه و بین زبانی مطرح است. در این مقاله محاسبه میزان شباهت معنایی بین دو جمله در زبان عربی و بین زبانی عربی- انگلیسی است که با توجه به عربی زبان بودن بسیاری از متون اسلامی، این پژوهش کاربردهای زیادی دارد. میزان شباهت معنایی بین دو جمله با استفاده از بردارهای معنایی دو جمله امکان پذیر است. در این تحقیق با استفاده از بردارهای از پیش آموزش داده شده بر روی متون عربی موجود در توییتر و ویکیپدیا با استفاده از دو روش CBOW و Skip-Gram که از معروفترین روشهای آموزش تعبیه کلمات میباشند بردارهای کلمات استخراج میگردد. همچنین از مدلهای مبتنی بر مبدلهای نظیر paraphrase-xlm-roberta نیز برای محاسبه شباهت معنایی بین زبانی عربی -انگلیسی مورد استفاده قرار گرفته است. برای ارزیابی مدل و آموزش آن با استفاده از دادههای موجود در کنفرانس شباهتیابی معنایی سال 2017 که بهصورت جفت جمله عربی و جفت جمله عربی -انگلیسی بودند اقدام به آموزش مدل شبکه عصبی عمیق با نام شبکه سیامی با استفاده از لایه LSTM نمودیم. استفاده از LSTM توانائی یادگیری وابستگیهای بلندمدت در شبکه را امکانپذیر میسازد. شبکههای سیامی در عین سادگی نتایج قابل قبولی را از خود نشان میدهند و مدلهای مبتنی بر مبدلها نیز قابلیت یادگیری بین زبانی را دارند. در لایه آخر شبکه، با استفاده از شباهت کسینوسی بین بردارهای متعلق به دو جمله ورودی، میزان شباهت بین آنها، به دست میآید. نتایج بیانگر آن است که با استفاده از روش پیشنهادی میزان همبستگی پیرسون 83.4 درصد برای جفت جمله عربی -عربی و میزان همبستگی پیرسون 82 درصد برای جفت جمله عربی-انگلیسی به دست میآید که از سایر روشهای موجود عملکرد بهتری را از خود نشان میدهد.
Agirre, Eneko, et al. "Semeval-2012 task 6: A pilot on semantic textual similarity." * SEM 2012: The First Joint Conference on Lexical and Computational Semantics–Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012). 2012.
Agirre, Eneko, et al. "Semeval-2014 task 10: Multilingual semantic textual similarity." Proceedings of the 8th international workshop on semantic evaluation (SemEval 2014). 2014.
Agirre, Eneko, et al. "* SEM 2013 shared task: Semantic textual similarity." Second Joint Conference on Lexical and Computational Semantics (* SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity. 2013.
Bar D., Biemann C., Gurevych I., and Zesch T. Ukp: Computing semantic textual similarity by combining multiple content similarity measures. Proceedings of the 6th International Workshop on Semantic Evaluation, in conjunction with the 1st Joint Conference on Lexical and Computational Semantics, 2012.
Bjerva, Johannes, and Robert Östling. "Cross-lingual learning of semantic textual similarity with multilingual word representations." 21st Nordic Conference on Computational Linguistics, NoDaLiDa, Gothenburg, Sweden, 22-24 May, 2017. Linköping University Electronic Press, 2017.
Brychcín, Tomáš. "Linear transformations for cross-lingual semantic textual similarity." Knowledge-Based Systems 187 (2020): 104819.
Cer, Daniel, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. "SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation." In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). Association for Computational Linguistics, 2017.Comelles, Elisabet, and Jordi Atserias. "VERTa: a linguistic approach to automatic machine translation evaluation." Language Resources and Evaluation1 (2019): 57-86.
Das, Arijit, and Diganta Saha. "Deep learning based Bengali question answering system using semantic textual similarity."Multimedia Tools and Applications (2022): 1-25.
Han, Lushan, et al. "UMBC_EBIQUITY-CORE: Semantic textual similarity systems." Second Joint Conference on Lexical and Computational Semantics (* SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity. 2013.
Hochreiter, Sepp, and Jürgen Schmidhuber. "Long short-term memory." Neural computation8 (1997): 1735-1780.
Islam, Aminul, and Diana Inkpen. "Semantic text similarity using corpus-based word similarity and string similarity." ACM Transactions on Knowledge Discovery from Data (TKDD) 2.2 (2008): 1-25.
Lubani, Mohamed, and Shahrul Azman Mohd Noah. "Text Relation Extraction Using Sentence-Relation Semantic Similarity." In Multi-disciplinary Trends in Artificial Intelligence: 13th International Conference, MIWAI 2019, Kuala Lumpur, Malaysia, November 17–19, 2019, Proceedings 13, pp. 3-14. Springer International Publishing, 2019.
Mihalcea, Rada, Courtney Corley, and Carlo Strapparava. "Corpus-based and knowledge-based measures of text semantic similarity." Aaai. Vol. 6. No. 2006. 2006.
Mikolov, Tomas, et al. "Distributed representations of words and phrases and their compositionality." Advances in neural information processing systems. 2013.
Mueller, J., & Thyagarajan, A. (2016, March). Siamese recurrent architectures for learning sentence similarity. In thirtieth AAAI conference on artificial intelligence.
Roul, Rajendra Kumar, and Jajati Keshari Sahoo. "Near-duplicate document detection using semantic-based similarity measure: a novel approach." In Computational Intelligence in Data Mining: Proceedings of the International Conference on ICCIDM 2018, pp. 543-558. Springer Singapore, 2020.
Rychalska B., Pakulska K., Chodorowska K., Walczak W., and Andruszkiewicz P. Samsung Poland NLP Team at SemEval-2016 Task 1: Necessity for diversity; combining recursive autoencoders, wordnet and ensemble methods to measure semantic similarity. Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval 2016), San Diego, CA, USA, 2016.
Schwab, Didier. "Semantic similarity of arabic sentences with word embeddings." 2017.
Shahmirzadi, Omid, Adam Lugowski, and Kenneth Younge. "Text similarity in vector space models: a comparative study." 2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA). IEEE, 2019.
Sultan M.A., Bethard S., and Sumner T. Back to basics for monolingual alignment: Exploiting word similarity and contextual evidence. Transactions of the Association for Computational Linguistics, 2:219–230, 2014a.
Sultan M.A., Bethard S., and Sumner T. DLS@CU: Sentence similarity from word alignment. Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), 241–246, Dublin, Ireland, August 2014b. Association for Computational Ling
Soliman, Abu Bakr, Kareem Eissa, and Samhaa R. El-Beltagy. "Aravec: A set of arabic word embedding models for use in arabic nlp." Procedia Computer Science 117 (2017): 256-265.
Suleiman, Dima, Arafat Awajan, and Nailah Al-Madi. "Deep Learning Based Technique for Plagiarism Detection in Arabic Texts." 2017 International Conference on New Trends in Computing Sciences (ICTCS). IEEE, 2017.
Tian, Junfeng, et al. "Ecnu at semeval-2017 task 1: Leverage kernel-based traditional nlp features and neural networks to build a universal model for multilingual and cross-lingual semantic textual similarity." Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). 2017.
Wu, Hao, et al. "BIT at SemEval-2017 Task 1: Using semantic information space to evaluate semantic textual similarity." Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). 2017.
عبدوس, محمد, و مینایی, بهروز. (1404). شباهتیابی معنایی تک زبانه و بین زبانی متون عربی با استفاده از یادگیری عمیق. علوم اسلامی و انسانی دیجیتال, 1(1), 103-122. https://doi.org/10.22034/disah.2024.716150
MLA
عبدوس, محمد, و مینایی, بهروز. "شباهتیابی معنایی تک زبانه و بین زبانی متون عربی با استفاده از یادگیری عمیق", علوم اسلامی و انسانی دیجیتال, 1, 1, 1404, 103-122. doi: 10.22034/disah.2024.716150
HARVARD
عبدوس محمد, مینایی بهروز. (1404). 'شباهتیابی معنایی تک زبانه و بین زبانی متون عربی با استفاده از یادگیری عمیق', علوم اسلامی و انسانی دیجیتال, 1(1), pp. 103-122. doi: 10.22034/disah.2024.716150
CHICAGO
محمد عبدوس و بهروز مینایی, "شباهتیابی معنایی تک زبانه و بین زبانی متون عربی با استفاده از یادگیری عمیق," علوم اسلامی و انسانی دیجیتال, 1 1 (1404): 103-122, doi: 10.22034/disah.2024.716150
VANCOUVER
عبدوس محمد, مینایی بهروز. شباهتیابی معنایی تک زبانه و بین زبانی متون عربی با استفاده از یادگیری عمیق. علوم اسلامی و انسانی دیجیتال. 1404;1(1):103-122. doi: 10.22034/disah.2024.716150