Journal of Studies in Applied Language
زبان کاوی کاربردی
JSAL
Literature & Humanities
http://jsal.ierf.ir
1
admin
2980-9304
2980-9304
10.61186/jsal
fa
jalali
1402
1
1
gregorian
2023
4
1
6
2
online
1
fulltext
fa
اثربخشی ترجمه ماشینی در ÙØ±Ø¢ÛŒÙ†Ø¯ پردازش زبان؛ بهره‌گیری از قرینه سیاق در معنایابی واژگان قرآن
Efficiency of Machine Translation in the Language Processing Process; Using Context Clues in Finding the [Exact] Meaning of Quranic Words [In Persian]
زبان شناسی اجتماعی
Sociolinguistics
پژوهشي
Research
<div dir="rtl" style="text-align: justify;"><span style="line-height:2;"><span style="font-family:nasimYW;"><span style="font-size:14px;"><span style="direction:rtl"><span style="unicode-bidi:embed"><span lang="FA"><span style="background:white">به ÙØ±Ø¢ÛŒÙ†Ø¯ برگرداندن مطلبی از زبان مبدأ به زبان مقصد Ú©Ù‡ با ÛŒØ§ÙØªÙ† هم ارزهای معناشناختی میان دو زبان صورت Ù…ÛŒ­Ú¯ÛŒØ±Ø¯ØŒ ترجمه Ù…ÛŒ­Ú¯ÙˆÛŒÙ†Ø¯. مهم­ØªØ±ÛŒÙ† مشکلات ترجمه، ابهاماتی است Ú©Ù‡ در واژگان Ùˆ ساختار جملات وجود دارند. در یک تقسیم­Ø¨Ù†Ø¯ÛŒØŒ پنج</span></span> <span lang="FA"><span style="background:white">نوع</span></span> <span lang="FA"><span style="background:white">مهم</span></span> <span lang="FA"><span style="background:white">ابهام</span></span> <span lang="FA"><span style="background:white">واژگانی (ابهام­Ù‡Ø§ÛŒ مقوله­Ø§ÛŒØŒ واژه­Ù‡Ø§ÛŒ</span></span> <span lang="FA"><span style="background:white">هم­Ø¢ÙˆØ§ØŒ</span></span> <span lang="FA"><span style="background:white">واژه­Ù‡Ø§ÛŒ</span></span> <span lang="FA"><span style="background:white">هم­Ù†ÙˆÛŒØ³Ù‡ØŒ</span></span> <span lang="FA"><span style="background:white">چند</span></span> <span lang="FA"><span style="background:white">معنایی</span></span> <span lang="FA"><span style="background:white">Ùˆ</span></span> <span lang="FA"><span style="background:white">ابهام</span></span> <span lang="FA"><span style="background:white">انتقالی) Ùˆ دو</span></span> <span lang="FA"><span style="background:white">نوع</span></span> <span lang="FA"><span style="background:white">مهم</span></span> <span lang="FA"><span style="background:white">ابهام ساختاری (ابهام­Ù‡Ø§ÛŒ</span></span> <span lang="FA"><span style="background:white">ساختاری</span></span> <span lang="FA"><span style="background:white">واقعی</span></span> <span lang="FA"><span style="background:white">Ùˆ ابهام­Ù‡Ø§ÛŒ سیستمی) وجود دارد. </span></span><span lang="FA"><span style="background:white">ترجمه</span></span> <span lang="FA"><span style="background:white">ماشینی</span></span> <span lang="FA"><span style="background:white">(</span></span><span dir="LTR">Machine translation: </span><span dir="LTR">MT</span><span lang="FA"><span style="background:white">)</span></span> <span lang="FA"><span style="background:white">Ú©Ù‡ بخشی از ØÙˆØ²Ù‡ پردازش زبان طبیعی(</span></span><span dir="LTR">Natural Language Processing: NLP</span><span lang="FA"><span style="background:white">) مبتنی بر کامپیوتر در زبان­Ø´Ù†Ø§Ø³ÛŒ رایانه­Ø§ÛŒ Ùˆ هوش مصنوعی بوده به عنوان</span></span><span lang="AR-SA"><span style="background:white"> یکی از تکنیکهای خودکاری است Ú©Ù‡ متن بدون ساختار را به دادههای ساختاری تبدیل میکند با تبدیل متن به اطلاعات، توانسته است تØÙ„یلهای بیشتری را به دادهها اعمال کرده تا اطلاعات Ù…Ùیدی استخراج شود</span></span><span lang="AR-SA"><span style="background:white">. در این نوشتار Ú©Ù‡ به روش کتابخانه­Ø§ÛŒ تدوین شده، جهت Ø±ÙØ¹ مسائل پیرامون معنای واژگان در ترجمه ماشینی قرآن، طرØÛŒ به صورت نظری پیشنهاد شده Ú©Ù‡ هد٠آن Ú©Ù…Ú© به Ùهم بهتر معنای واژگان قرآن، با بهرمندی از قرینه سیاق Ùˆ Ø¨Ø§ÙØª عبارت است. </span></span><span lang="FA"><span style="background:white">در روش پیشنهادی با بهره­Ù…ندی از قاعده سیاق Ùˆ تکنیک­Ù‡Ø§ÛŒ متن کاوی، Ùˆ با استناد به آن، واژه معادل مناسب­ØªØ±ÛŒ در زبان مقصد برگزیند. در این Ø·Ø±ØØŒ سیاق را در مقیاس کلمات دانسته Ú©Ù‡ Ù…ÛŒ­ØªÙˆØ§Ù† آن را به شرط Ø§ØØ±Ø§Ø² شرایط به انواع دیگر توسعه داد. به طور خلاصه این Ø·Ø±Ø Ø¯Ùˆ مرØÙ„Ù‡ دارد: اولویت­Ø¨Ù†Ø¯ÛŒ (وزن­Ø¯Ù‡ÛŒ) واژگان هم­Ø¬ÙˆØ§Ø±Ù هم ورودی (هر واژه در Ù…ØØ¯ÙˆØ¯Ù‡ آیاتی Ú©Ù‡ در مورد نزول یک­Ø¨Ø§Ø±Ù‡ آن­Ù‡Ø§ Ø§ØªÙØ§Ù‚ نظر وجود دارد) Ùˆ سپس مقایسه با کلماتی Ú©Ù‡ اشتراک Ù„ÙØ¸ÛŒ (چندمعنا) دارند Ùˆ نیز مقایسه همنظیران یک واژه با همنظیران سایر واژگان (متراد٭یابی). Ù…ÛŒ باشد. برای دقیق­ØªØ± شدن نتایج Ù…ÛŒ­ØªÙˆØ§Ù† مشخصات بیشتری از کلمات را به صورت دستی تهیه نمود، جداولی شامل مواردی چون Ù…Ú©Ù‘ÛŒ یا مدنی بود آیات، ترتیب نزول سوره، Ù…ÙØ§Ù‡ÛŒÙ… Ùˆ تعابیری Ú©Ù‡ در معنای کلمات قرآن در ÙØ±Ù‡Ù†Ú¯ لغاتی چون لسان العرب ابن منظور Ùˆ ÙØ±Ù‡Ù†Ú¯ لغت راغب اصÙهانی آمده است Ùˆ غیره. برای بدست آوردن داده­Ù‡Ø§ÛŒ ورودی از تکنیک­Ù‡Ø§ÛŒ نمایه­Ø³Ø§Ø²ÛŒ Ø§Ø³ØªÙØ§Ø¯Ù‡ Ù…ÛŒ­Ø´ÙˆØ¯</span></span><span lang="FA"><span style="background:white">. </span></span><span lang="FA"><span style="background:white">در مرØÙ„Ù‡ پیش پردازش باید داده­Ù‡Ø§ÛŒÛŒ Ú©Ù‡ دارای اهمیت کمتری است(</span></span><span dir="LTR">Stop Words</span><span lang="FA"><span style="background:white">) (مانند</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white">الذی</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white">ØŒ "</span></span><span lang="FA"><span style="background:white">التی</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white">ØŒ </span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white">لم</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white">ØŒ</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white">کان</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white">ØŒ</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white"> کانما</span></span><span lang="FA"><span style="background:white">"</span></span><span lang="FA"><span style="background:white"> Ùˆ غیره) ØØ°Ù شود تا خروجی بهتری بدست آید. برای تغییر Ø´Ú©Ù„ داده Ù…ÛŒ­ØªÙˆØ§Ù† اعراب را ØØ°Ù کرد تا کدنویسی Ø±Ø§ØØª­ØªØ± انجام شود، برای کاهش نمونه نیز Ù…ÛŒ­ØªÙˆØ§Ù† از ریشه میانوندی کلمات Ø§Ø³ØªÙØ§Ø¯Ù‡ نمود. برای اینکه با استناد به قاعده سیاق، برای یکایک کلماتی Ú©Ù‡ به عنوان ورودی مورد پردازش قرار Ù…ÛŒ­Ú¯ÛŒØ±Ù†Ø¯ØŒ رکوردی از مشخصات تهیه نمود، لازم است ابتدا ÙˆØ§ØØ¯ سازی(</span></span><span dir="LTR">Tokenizer</span><span lang="FA"><span style="background:white">) صورت گیرد، در داده­Ù‡Ø§ÛŒ اولیه تهیه شده، در Ú©Ù„ مجموعه آیات ورودی، بر اساس دو معیار قرابت مکانی Ùˆ ÙØ±Ø§ÙˆØ§Ù†ÛŒ تکرار، به هر کلمه وزنی اختصاص یابد. هر Ú†Ù‡ کلمات به کلمه مورد نظر نزدیک­ØªØ± Ùˆ یا بیشترتکرار شده باشد، وزن بیشتری به آن اختصاص داده Ù…ÛŒ شود Ú©Ù‡ معر٠ارتباط معنایی قوی­ØªØ± آنان است Ùˆ برعکس. طبیعتاَ کلماتی Ú©Ù‡ در یک آیه قرار دارند (شماره آیه یکسانی دارند) نسب به کلماتی Ú©Ù‡ در آیات دیگر Ùˆ ÙØ§ØµÙ„Ù‡ دورتر قرار دارند از ظریب تأثیر بیشتری برخوردار هستند. در سنجش معیار ÙØ±Ø§ÙˆØ§Ù†ÛŒ: برای نشان دادن اهمیت کلمه در سوره از ÙØ±Ø§ÙˆØ§Ù†ÛŒ وزنی (</span></span><span dir="LTR"><span style="background:white">TF/IDF Weigh</span></span><span dir="LTR">t</span><span lang="FA"><span style="background:white">) Ø§Ø³ØªÙØ§Ø¯Ù‡ Ù…ÛŒ­Ø´ÙˆØ¯ØŒ</span></span> <span lang="FA"><span style="background:white">مقدار</span></span> <span dir="LTR"><span style="background:white">TF/IDF</span></span> <span style="background:white"> به تناسب تعداد تکرار کلمه در هر سوره یا مجموعه آیات ورودی، Ø§ÙØ²Ø§ÛŒØ´ مییابد Ùˆ توسط تعداد آیاتی Ú©Ù‡ در سوره هستند Ùˆ شامل کلمه نیز میباشند متعادل میشود</span><span lang="AR-SA">.</span> <span lang="FA"><span style="background:white">در نهایت این نتیجه ØØ§ØµÙ„ آمد Ú©Ù‡</span></span><span lang="AR-SA"><span style="background:white"> از هم­Ø¬ÙˆØ§Ø±ÛŒ کلمات Ùˆ روابط معنایی بین آن­Ù‡Ø§ Ùˆ با Ú©Ù…Ú© تکنیک­Ù‡Ø§ÛŒ متن کاوی، Ùهم بیشتری از واژگان ØØ§ØµÙ„ شده Ú©Ù‡ این مهم گزینش مناسب­ØªØ± واژه معادل در زبان مقصد را منجر Ù…ÛŒ شود. </span></span></span></span></span></span></span></div>
<div style="text-align: justify;"><span style="font-size:12px;"><span style="font-family:Times New Roman;"><span style="line-height:2;"><span style="unicode-bidi:embed"><span style="color:black">Translation is the transfer of the content of a text from the source language in to the target language, which is done by finding semantic equivalents between the two languages. The most important problems facing translation are the ambiguities in vocabulary and sentence structure. In a division, there are five important types of lexical ambiguity (categorical ambiguities, homophones, homographs, polysemy and transitive ambiguity), and two important types of structural ambiguity (real structural ambiguities and systemic ambiguities). Machine translation (MT), which is a part of the computer-based field of natural language processing (NLP) in computational linguistics and artificial intelligence, is considered as one of the automatic techniques that that convert unstructured text into structured data, and by converting text into information, it has been able to apply further analysis to the data to extract useful information. In this article, which was compiled in a library method, a theoretical plan has been proposed to resolve the issues surrounding the meaning of words in the machine translation of the Quran, the purpose of which is to help better understand the meaning of the words of the Quran, by taking advantage of the context clues and styles of the expressions. In the proposed method, a more suitable equivalent word is chosen in the target language by taking advantage of the context rule and text mining techniques, and referring to it. In this plan, the context is considered in the scale of words, which can be developed to other types if the conditions are met. In short, this plan has two steps: prioritizing (weighting) the adjacent words next to each other (any word within the range of verses where there is a consensus about their simultaneous descent) and then, comparing with the homonyms words (polysemous), and also comparing the equivalents of a word with the equivalents of other words (synonymization). In order to make the results more accurate, more specifications of the words can be prepared manually, tables that include things such as whether the verses are Meccan or Medinan, the order of revelation of the Surahs, the concepts and interpretations that are mentioned in the meaning of the words of the Qur'an in dictionaries such as Lisan al-Arab by Ibn Manzur and The Book of Vocabulary in the Strange Qur'an by Al-Ragheb Al-Isfahani and so on. Indexing techniques are used to obtain input data. In the pre-processing stage, the data that is less important (Stop Words) (such as “al-lazi (which)”, “al-lati (that is)”, “lam (not)”, “k'ana (was)”, “kaannama (as if)”, etc.) should be removed to get a better output. To change the shape of the data, the diacritic can be removed to make coding easier, and to reduce the sample size, the infix of the words can be used. In order to prepare a record of specifications for each word that is processed as input, based on the rule of context clues, at first, it is necessary to create a tokenizer, to prepare it in the primary data, and in the entire collection of input verses, a weight should be assigned to each word based on the two criteria of spatial proximity and frequency of repetition. The closer the words are to the desired word or the more it is repeated, the more weight is assigned to it, which represents their stronger semantic connection, and vice versa. Naturally, the words that are in the same verse (have the same number of the verse) have a greater influence than the words that are in other verses and at a further distance. In measuring the frequency criterion, weighted frequency (TF/IDF Weight) is used to show the importance of the word in the surah, the value (TF/IDF value) increases proportionally to the number of times a word appears in each surah or set of input verses, and is balanced by the number of verses that are in the Surah and contain the word. Finally, it was concluded that by using the contiguity of words and the semantic relations between them, and with the help of text mining techniques, a greater understanding of the vocabulary was obtained, which leads to a more appropriate selection of the equivalent word in the target language.</span></span></span></span></span></div>
زبان شناسی رایانه ای, زبان شناسی اجتماعی, ترجمه ماشینی, قرآن, قرینه سیاق, معادل یابی واژگان
Computational Linguistics, Sociolinguistics, Machine Translation, Qur'an, Context Correlation, Finding Equivalents for the Words
101
130
http://jsal.ierf.ir/browse.php?a_code=A-10-2-11&slc_lang=fa&sid=1
Zaynab
Shams
زینب
شمس
Z.Shams@qom.ac.ir
1003194753284600509
1003194753284600509
Yes
PhD student of Qur'anic and Hadith Sciences, Faculty of Theology, Kashan University, Iran
دانشجوی دکتری علوم قرآن Ùˆ ØØ¯ÛŒØ«ØŒ دانشکده الهیات، دانشگاه کاشان، ایران
Sepideh
Chehreh
سپیده
چهره
1003194753284600510
1003194753284600510
No
Master of Artificial Intelligence, Islamic Azad University Science and Research Branch, Iran
کارشناس ارشد هوش مصنوعی دانشگاه علوم تØÙ‚یقات تهران