Hybrid Translation with Classification: Revisiting Rule-Based and Neural Machine Translation

Jin-Xia Huang; Kyung-Soon Lee; Young-Kil Kim

doi:10.3390/electronics9020201

Hybrid Translation with Classification: Revisiting Rule-Based and Neural Machine Translation

Electronics ◽

10.3390/electronics9020201 ◽

2020 ◽

Vol 9 (2) ◽

pp. 201

Author(s):

Jin-Xia Huang ◽

Kyung-Soon Lee ◽

Young-Kil Kim

Keyword(s):

Machine Translation ◽

Classification Accuracy ◽

Training Data ◽

Translation System ◽

Rule Based ◽

Neural Machine Translation ◽

Machine Translation System ◽

Text Classifiers ◽

Hybrid Machine Translation ◽

Translation Accuracy

This paper proposes a hybrid machine-translation system that combines neural machine translation with well-developed rule-based machine translation to utilize the stability of the latter to compensate for the inadequacy of neural machine translation in rare-resource domains. A classifier is introduced to predict which translation from the two systems is more reliable. We explore a set of features that reflect the reliability of translation and its process, and training data is automatically expanded with a small, human-labeled dataset to solve the insufficient-data problem. A series of experiments shows that the hybrid system’s translation accuracy is improved, especially in out-of-domain translations, and classification accuracy is greatly improved when using the proposed features and the automatically constructed training set. A comparison between feature- and text-based classification is also performed, and the results show that the feature-based model achieves better classification accuracy, even when compared to neural network text classifiers.

Download Full-text

English to Hindi Machine Translation System in the Context of Homoeopathy Literature

International Journal of Artificial Life Research ◽

10.4018/ijalr.2016010103 ◽

2016 ◽

Vol 6 (1) ◽

pp. 46-62

Author(s):

Pramod P. Sukhadeve

Keyword(s):

Machine Translation ◽

Translation System ◽

Simple Complex ◽

Rule Based ◽

Machine Translation System ◽

Translation Accuracy

Over the years, researches in machine translation (MT) systems have gain momentum due to their widespread applicability. A number of systems have come up doing the task successfully for different language pairs. However, to the best of the author's knowledge, no significant work has been done in clinical and medical related domain especially in Homoeopathy. This paper describes a rule based English-Hindi MT system for Homoeopathic sentences. It has been designed to translate a variety of sentences from Homoeopathic literature. To achieve the task, the author developed English and Hindi Homoeopathic corpuses presently having the size 21096 and 23145 sentences respectively. For translation, the input sentences (in English) have been categorised in four different type's i.e. simple, complex, interrogative and ambiguous sentences. The authors tested the translation accuracy using BLEU score. At present, the overall Bleu score of the system is 0.7808 and the accuracy percentage is 82.25%.

Download Full-text

Neural machine translation system for the Kazakh language based on synthetic corpora

MATEC Web of Conferences ◽

10.1051/matecconf/201925203006 ◽

2019 ◽

Vol 252 ◽

pp. 03006

Author(s):

Ualsher Tukeyev ◽

Aidana Karibayeva ◽

Balzhan Abduali

Keyword(s):

Machine Translation ◽

Training Data ◽

Translation System ◽

Natural Languages ◽

Neural Machine Translation ◽

Translation Quality ◽

Parallel Data ◽

Machine Translation System ◽

Turkic Languages

The lack of big parallel data is present for the Kazakh language. This problem seriously impairs the quality of machine translation from and into Kazakh. This article considers the neural machine translation of the Kazakh language on the basis of synthetic corpora. The Kazakh language belongs to the Turkic languages, which are characterised by rich morphology. Neural machine translation of natural languages requires large training data. The article will show the model for the creation of synthetic corpora, namely the generation of sentences based on complete suffixes for the Kazakh language. The novelty of this approach of the synthetic corpora generation for the Kazakh language is the generation of sentences on the basis of the complete system of suffixes of the Kazakh language. By using generated synthetic corpora we are improving the translation quality in neural machine translation of Kazakh-English and Kazakh-Russian pairs.

Download Full-text

Otedama: Fast Rule-Based Pre-Ordering for Machine Translation

Prague Bulletin of Mathematical Linguistics ◽

10.1515/pralin-2016-0015 ◽

2016 ◽

Vol 106 (1) ◽

pp. 159-168 ◽

Cited By ~ 1

Author(s):

Julian Hitschler ◽

Laura Jehl ◽

Sariya Karimova ◽

Mayumi Ohta ◽

Benjamin Körner ◽

...

Keyword(s):

Open Source ◽

Machine Translation ◽

State Of The Art ◽

Statistical Machine Translation ◽

Training Data ◽

Translation System ◽

Rule Based ◽

Machine Translation System ◽

Target Languages ◽

Established Technique

Abstract We present Otedama, a fast, open-source tool for rule-based syntactic pre-ordering, a well established technique in statistical machine translation. Otedama implements both a learner for pre-ordering rules, as well as a component for applying these rules to parsed sentences. Our system is compatible with several external parsers and capable of accommodating many source and all target languages in any machine translation paradigm which uses parallel training data. We demonstrate improvements on a patent translation task over a state-of-the-art English-Japanese hierarchical phrase-based machine translation system. We compare Otedama with an existing syntax-based pre-ordering system, showing comparable translation performance at a runtime speedup of a factor of 4.5-10.

Download Full-text

English to Sanskrit machine translation system: a rule-based approach

International Journal of Advanced Intelligence Paradigms ◽

10.1504/ijaip.2012.048144 ◽

2012 ◽

Vol 4 (2) ◽

pp. 168 ◽

Cited By ~ 2

Author(s):

Vimal Mishra ◽

R.B. Mishra

Keyword(s):

Machine Translation ◽

Translation System ◽

Rule Based ◽

System A ◽

Machine Translation System ◽

Rule Based Approach

Download Full-text

Handling Unknown Words in Neural Machine Translation System

2020 International Conference on Decision Aid Sciences and Application (DASA) ◽

10.1109/dasa51403.2020.9317169 ◽

2020 ◽

Author(s):

Kamal Deep Garg ◽

Jatin Gupta ◽

Vandana Saini

Keyword(s):

Machine Translation ◽

Translation System ◽

Neural Machine Translation ◽

Machine Translation System ◽

Unknown Words

Download Full-text

SisHiTra : A Hybrid Machine Translation System from Spanish to Catalan

Advances in Natural Language Processing - Lecture Notes in Computer Science ◽

10.1007/978-3-540-30228-5_31 ◽

2004 ◽

pp. 349-359 ◽

Cited By ~ 1

Author(s):

José R. Navarro ◽

Jorge González ◽

David Picó ◽

Francisco Casacuberta ◽

Joan M. de Val ◽

...

Keyword(s):

Machine Translation ◽

Translation System ◽

Machine Translation System ◽

Hybrid Machine ◽

Hybrid Machine Translation

Download Full-text

Translation of Medical Texts using Neural Networks

International Journal of Reliable and Quality E-Healthcare ◽

10.4018/ijrqeh.2016100104 ◽

2016 ◽

Vol 5 (4) ◽

pp. 51-66 ◽

Cited By ~ 5

Author(s):

Krzysztof Wolk ◽

Krzysztof P. Marasek

Keyword(s):

Machine Translation ◽

Statistical Machine Translation ◽

European Medicines Agency ◽

Translation System ◽

Training Methods ◽

Neural Machine Translation ◽

Machine Translation System ◽

Source Sentence ◽

Parallel Text ◽

Translation Systems

The quality of machine translation is rapidly evolving. Today one can find several machine translation systems on the web that provide reasonable translations, although the systems are not perfect. In some specific domains, the quality may decrease. A recently proposed approach to this domain is neural machine translation. It aims at building a jointly-tuned single neural network that maximizes translation performance, a very different approach from traditional statistical machine translation. Recently proposed neural machine translation models often belong to the encoder-decoder family in which a source sentence is encoded into a fixed length vector that is, in turn, decoded to generate a translation. The present research examines the effects of different training methods on a Polish-English Machine Translation system used for medical data. The European Medicines Agency parallel text corpus was used as the basis for training of neural and statistical network-based translation systems. A comparison and implementation of a medical translator is the main focus of our experiments.

Download Full-text

Neural Machine Translation System of Indic Languages - An Attention based Approach

2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP) ◽

10.1109/icaccp.2019.8882969 ◽

2019 ◽

Cited By ~ 1

Author(s):

Parth Shah ◽

Vishvajit Bakrola

Keyword(s):

Machine Translation ◽

Translation System ◽

Neural Machine Translation ◽

Machine Translation System

Download Full-text

Estimating Machine Translation Quality of Any Input Sentence

International Journal of Asian Language Processing ◽

10.1142/s2717554520500022 ◽

2020 ◽

Vol 30 (01) ◽

pp. 2050002

Author(s):

Taichi Aida ◽

Kazuhide Yamamoto

Keyword(s):

Machine Translation ◽

Evaluation Model ◽

Joint Probability ◽

Translation System ◽

Neural Machine Translation ◽

Translation Quality ◽

Machine Translation System ◽

Input Sentence ◽

The One

Current methods of neural machine translation may generate sentences with different levels of quality. Methods for automatically evaluating translation output from machine translation can be broadly classified into two types: a method that uses human post-edited translations for training an evaluation model, and a method that uses a reference translation that is the correct answer during evaluation. On the one hand, it is difficult to prepare post-edited translations because it is necessary to tag each word in comparison with the original translated sentences. On the other hand, users who actually employ the machine translation system do not have a correct reference translation. Therefore, we propose a method that trains the evaluation model without using human post-edited sentences and in the test set, estimates the quality of output sentences without using reference translations. We define some indices and predict the quality of translations with a regression model. For the quality of the translated sentences, we employ the BLEU score calculated from the number of word [Formula: see text]-gram matches between the translated sentence and the reference translation. After that, we compute the correlation between quality scores predicted by our method and BLEU actually computed from references. According to the experimental results, the correlation with BLEU is the highest when XGBoost uses all the indices. Moreover, looking at each index, we find that the sentence log-likelihood and the model uncertainty, which are based on the joint probability of generating the translated sentence, are important in BLEU estimation.

Download Full-text

Attention-Based Syllable Level Neural Machine Translation System for Myanmar to English Language Pair

International Journal on Natural Language Computing ◽

10.5121/ijnlc.2019.8201 ◽

2019 ◽

Vol 8 (2) ◽

pp. 01-11

Author(s):

Yi Mon Shwe Sin ◽

Khin Mar Soe

Keyword(s):

Machine Translation ◽

English Language ◽

Translation System ◽

Neural Machine Translation ◽

Machine Translation System ◽

Language Pair

Download Full-text