علوم اسلامی و انسانی دیجیتال

علوم اسلامی و انسانی دیجیتال

موتور پیراسته ساز هوشمند کلمات عربی نور

نوع مقاله : مقاله پژوهشی

نویسنده
مهندس برنامه‌نویس گروه پردازش هوشمند مرکز تحقیقات کامپیوتری علوم اسلامی
چکیده
استمر به الگوریتمی گفته می‌شود که در پردازش زبان طبیعی (NLP) در تجزیه و تحلیل ریخت شناسی، برای استخراج و نشان دادن شکل اصلی یا میانوند کلمات استفاده می­شود. به عبارت دیگر، استمر با حذف پیشوندها، پسوندها و مصوت‌ها از افعال صرف شده و اسماء مشتق و سایر اشکال کلمه حاصل می‌شود. هدف یک استمر عربی این است که با حفظ هویت معنایی و عملکرد نحوی، هر کلمه را به شکل پایه آن تقلیل دهد، بنابراین فهرست‌بندی، جستجو و دسته‌بندی کارآمد مجموعه‌های بزرگ اسناد عربی را تسهیل می‌کند. این ابزار نقش حیاتی در بازیابی اطلاعات، طبقه بندی اسناد، ترجمه ماشینی، خلاصه سازی و سیستم های پاسخگویی به سوالات دارد. ویژگی منحصر به فرد استمر هوشمند کلمات عربی بهره‌گیری همزمان از متد قاعده‌ای (Rule-Based)، متد آماری (Data-driven) و متد یادگیری (Learning-Based) می‌باشد. بر این اساس، کارکرد این موتور در سطح قابل توجهی از سایر موتورهای استمر پیشی گرفته است. در این موتور از داده های عربی متون کهن و جدید به همراه تحلیل های صرفی هوشمند و همچنین تکنیک های پیراسته‌سازی مبتنی بر ساختارِ قاعده مند زبان عربی استفاده شده است و تلفیق این سه متد در جهت ابهام زدایی از کلماتی که مبتنی بر قواعد زبان عربی شکل سالمی ندارند، در کارآمدی آن بسیار موثر واقع شده است و همین امر سبب شده است برخی از چالش‌هایی را که سایر موتورهای استمر با آن مواجه شوند، این موتور پشت سر بگذارد. ضمن اینکه این موتور با استفاده از استم های انسانی که توسط متخصصان ارائه شده است، مقایسه و ارزیابی شده است که دقت آن در مقایسه با سایر استمرها از سطح کیفی بسیار بالایی برخودار است. از دیگر ویژگی های این موتور، امکان ارائه استم‌ها در سطوح مختلف می‌باشد که امکان بسیار ارزشمند و جدیدی است که بر اساس آن می‌توان نیازمندهای بسیاری از کاربران را پوشش داد.
کلیدواژه‌ها

 1.          الهی­ منش، محمدحسین (1394)، ابهام‌زدایی هوشمند صرفی نور، ره­آورد نور، شماره 53، ص13-18.
2.         بدیل یعقوب، إمیل (1425)، المعجم المفصل فی الجموع، دارالکتب العلمیة، بیروت: لبنان.
3.         دانش، سید محمد (1393)، تحلیل‌گر هوشمند صرفی نور، شماره 49، ص15– 23.
4.         سریانی حبیب (1395)، شبکه واژگانی زبان عربی با استفاده از فرآیند نیمه خودکار در داده‌های علوم اسلامی، ره­آورد نور، شماره 57، ص47-56.
5.         سریانی حبیب، مینایی بهروز (1390). سیستم هوشمند برچسب‌گذار ادات سخن زبان عربی؛ لایه صرف، ره­آورد نور، شماره 34، ص18 - 28.
6.         مصطفی إبراهیم، الزیات أحمد حسن، عبدالقادر حامد، النجار محمد علی (1429). المعجم الوسیط، تهران: موسسة الصادق للطباعة و النشر.
7.          Ababneh Mohamad, Al-Shalabi Riyad, Kanaan Ghassan, and Al-Nobani Alaa (2012). Building an Effective Rule-Based Light Stemmer for Arabic Language to Improve Search Effectiveness. The International Arab Journal of Information Technology, Vol. 9, No. 4, July 2012: Pp 368-372.
8.          Abu Ata, B, Al-Omari, A. (2014). A Rule-Based Stemmer for Arabic Gulf Dialect. Journal of King Saud University, Science 50(2), Computer and Information Sciences (JKSU). (Submitted). DOI:10.1016/j.jksuci.2014.04.003.
9.          Al-Fedaghi S. and F. Al-Anzi. (1989). “A new algorithm to generate Arabic root-pattern forms”. In proceedings of the 11th national Computer Conference and Exhibition. PP 391-400.
10.      Aljlayl, M, Frieder, O. (2002). On Arabic search: improving the retrieval effectiveness via a light stemming approach. In: Proceedings of the Eleventh International Conference on Information and Knowledge Management, McLean, VA
11.      Al-Kabi M.N, Al-Radaideh Q.A, Akkawi K.W. (2011). Benchmarking and assessing the performance of Arabic stemmers J. Inform. Sci, Vol 37 (2), pp. 111-119
12.      Al-Kabi Mohammed N, Kazakzeh Saif A, Abu Ata Belal M, Al-Rababah Saif A, Alsmad Izzat M. i (2015). A novel root based Arabic stemmer. Journal of King Saud University – Computer and Information Sciences (2015) 27: Pp 94–103.
13.      Al-Serhan, H, Ayesh, A. (2006). A triliteral word roots extraction using neural network for Arabic, In: The 2006 International Conference on Computer Engineering and Systems: Pp. 436–440.
14.      Al-Shalabi, R, Kanaan, G, Ghwanmeh, S, Nour, F. M. (2007). Stemmer algorithm for Arabic words based on excessive letter locations. In: 4th International Conference on Innovations in Information Technology (IIT ‘07): Pp. 456–460.
15.      Boubas, A, Lulu, L, Belkhouche, B, Harous, S. (2011). GENESTEM: A novel approach for an Arabic stemmer using genetic algorithms. In: International Conference on Innovations in Information Technology (IIT 2011): Pp. 77–82.
16.      Boudchiche Mohamed, Mazroui Azzeddine (2018). A hybrid approach for Arabic lemmatization. International Journal of Speech Technology, https://doi.org/10.1007/s10772-018-9528-3.
17.      Boudlal, A, Lakhouaja, A, Mazroui, A, Meziane, A, Ould Abdallahi Ould Bebah, M,Shoul, M. (2010). Alkhalil Morpho SYS1: a morphosyntactic analysis system for arabic texts. In: International Arab Conference on Information Technology. Benghazi, Libya: Pp 1–6.
18.      Buckwalter T. (2007). Issues in Morphological Analysis, in Arabic Computational Morphology. Eds. A. Soudi, A. van den Bosch, and G. Neumann. Springer, 2007: Pp. 23–41
19.      Chen, A, Gey, F.C. (2002). Building an Arabic stemmer for information retrieval. In: Proceedings of the 11th Text Retrieval Conference (TREC).
20.      Darwish, K. (2003). Probabilistic methods for searching OCR-degraded Arabic text. Unpublished Ph.D. Thesis. University of Maryland, USA.
21.      El-Sadany, T.A, Hashish, M.A. (1989). An Arabic Morphological System, IBM System Journal. Vol.28 (4). Pp 600-612.
22.      Flores F.N, Moreira V.P. (2016). Assessing the impact of stemming accuracy on information retrieval Inf. Process. Manage, Vol 52 (5): Pp 840-854.
23.       Frakes William B, Fox Christopher J. (2003). Strength and Similarity of Affix Removal Stemming Algorithms. ACM SIGIR Forum, Volume 37, No. 1: Pp 26-30.
24.      Galvez C, F, Anegón de Moya, Solana V.H. (2005). Term conflation methods in information retrieval: non-linguistic and linguistic approaches Journal of Document, Vol 61 (4): Pp 520-547
25.      Goweder, A, Poesio, M, De Roeck, A. and Reynolds, J. (2005) Identifying Broken Plurals in Unvowelised Arabic Text. Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing, Vancouver, 6-8 October 2005: Pp 246-253
26.      Hadni Meryeme, El Ouatik Saîd Alaoui, Lachkar Abdelmonaime (2013). Effective Arabic Stemmer Based Hybrid Approach for Arabic Text Categorization. International Journal of Data Mining & Knowledge Management Process (IJDKP) Vol.3, No.4: Pp 1-14.
27.      Hamdy Mubarak (2017). Build Fast and Accurate Lemmatization for Arabic, Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC-2018), Miyazaki, Japan, LREC.
28.      J. Yaghi and S. M. Yagi (2004). Systematic Verb Stem Generation for Arabic, in Proc. of the Workshop on Computational Approaches to Arabic Script-Based Languages (Geneva, Switzerland, August 28 - 28, 2004). ACL Workshops. Association for Computational Linguistics, Morristown, NJ: Pp. 23–30.
29.      Kazem T, Rania E, and Je.rey C. (2005). Arabic Stemming Without A Root Dictionary, International Conference on Information Technology: Coding and Computing (ITCC'05) - Volume II, USA. DOI: 10.1109/ITCC.2005.90
30.      Khoja, S. and Garside, R. (1999). Stemming Arabic Text. Technical report, Computing Department, Lancaster University,UK Lancaster.
31.      Larkey L, Ballesteros L. Connell M.E. (2007). Light stemming for Arabic information retrieval. Arabic Computational Morphology: Knowledge-based and Empirical Methods, Springer, Netherlands (2007): Pp 221-243
32.      Larkey, L, Ballesteros, L, Connell, M.E. (2002). Improving stemming for Arabic information retrieval: light stemming and co-occurrence analysis. In: SIGIR’02, Tampere, Finland: Pp. 275–282.
33.              Lovins, J.B. (1968). Development of a Stemming Algorithm. Mechanical Translation and Computational Linguistics, vol.11, nos.1 and 2, March and June 1968: Pp 22-31.
34.      Mustafa Mohammad, Salah Eldeen Afag, Bani-Ahmad Sulieman, Osman Elfaki Abdelrahman (2017). A Comparative Survey on Arabic Stemming: Approaches and Challenges. Intelligent Information Management, Vol 9: Pp 39-67.
35.      Mustafa Suleiman H. (2012), Word Stemming for Arabic Information Retrieval: The Case for Simple Light Stemming. ABHATH AL-YARMOUK: "Basic Sci. & Eng.", by Yarmouk University, Irbid, Jordan, Vol. 21, No. 1, 2012: Pp. 123-144.
36.      Namly Driss, Tajmout Rachida, Bouzoubaa Karim and Abouenour Lahsen (2016). NAFIS: A Gold Standard Corpus for Arabic Stemmers Evaluation, In Proc of the 28th International Business Information Management Association Conference IBIMA, Seville, Spain. ISBN: 978-0-9860419-8-3.
37.      Pasha Arfath, Al-Badrashiny Mohamed, Diab Mona, El Kholy Ahmed, Eskander Ramy, Habash Nizar, Pooleery Manoj, Rambow Owen, and Roth Ryan. (2014). MADAMIRA: A Fast, Comprehensive Tool for Morphological Analysis and Disambiguation of Arabic. In LREC, vol 14: Pp. 1094-1101.
38.      Rogati, M, McCarley, S, Yang, Y. (2003). Unsupervised learning of Arabic stemming using a parallel corpus. In: Proceedings of the 41st Annual Meeting on Association for Computational Linguistics-Volume 1, Stroudsburg, USA.
39.      Saad, M.K, Ashour, W. (2010). Arabic morphological tools for text mining. In: 6th International Conference on Electrical and Computer Systems (EECS’10), Lefke, North Cyprus, 2010.
40.      Sawalha Majdi and Eric Steven Atwell. (2008). Comparative evaluation of arabic language morphological analysers and stemmers. Proceedings of COLING 2008 22nd International Conference on Computational Linguistics. Companion volume Posters and Demonstrations, Manchester: Pp 107–110.
41.      Sembok, T. M. T, & AbuAta, B. (2013). Arabic word stemming algorithms and retrieval effectiveness. In Lecture Notes in Engineering and Computer Science.Vol. 3 LNECS, London: Pp 1577-1582.
42.      W.B. Frakes, C.J. Fox (2003). Strength and similarity of affix removal stemming algorithms, ACM SIGIR Forum, Vol 37 (1): Pp 26-30. DOI: 10.1145/945546.945548.
43.      Younes Jaafar, Driss Namly, Karim Bouzoubaa, Abdellah Yousfi (2017). Enhancing Arabic stemming process using resources and benchmarking tools. Journal of King Saud University - Computer and Information Sciences. Volume 29, Issue 2, April 2017: Pp 164-170.
44.      Yusof RJR, Zainuddin R, Mohd Sapiyan Baba, Zulkifli Mohd Yusoff (2010). QUR'ANIC WORDS STEMMING, the Arabian Journal for Science and Engineering, Volume 35, Number 2C, December 2010: Pp 37-49.
45.      Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, Nizar Habash (2020), An Open Source Python Toolkit for Arabic Natural Language Processing, European Language Resources Association (ELRA), Pp 7022-7032.