Main Article Content

Abstract

Modern Neural Machine Translation (NMT) systems have achieved state-of-the-art, performance, largely due to the availability of large-scale parallel corpora. However, the translation quality of NMT for Low-Resource Languages ​​(LRL) remains limited due to data sparsity. Numerous studies have proposed different strategies to address this challenge. Among the most widely adopted strategies are Transfer Learning (TL) and Data Augmentation (DA) strategies. This research aims to present a systematic review of how these techniques, including Back-Translation (BT), Hybrid Transfer Learning (HTL), and the utilization of self-supervised objectives such as Masked Language Modeling (MLM), Causal Language Modeling (CLM), and Denoising Autoencoder (DAE), affect the quality improvement of NMT for LRL. The findings show that a hybrid combination of TL and DA with a self-supervised objective is the most effective solution for extremely low-resource scenarios, capable of producing the highest translation quality (highest BLEU score) and outperforming baseline models and traditional methods such as Statistical Machine Translation (SMT).

Keywords

transfer learning data augmentation neural machine translation self-supervised learning

Article Details

How to Cite
Khuluq, N. F., Muzhaffar, M. N., & ’Uyun, S. (2026). A Systematic Review of Transfer Learning and Data Augmentation in Neural Machine Translation of Low-Resource Languages. Jurnal Sains, Nalar, Dan Aplikasi Teknologi Informasi, 5(2), 138–145. https://doi.org/10.20885/snati.v5.i2.47528

References

  1. D. Ataman, A. Birch, N. Habash, M. Federico, P. Koehn, and K. Cho, “Machine Translation in the Era of Large Language Models:A Survey of Historical and Emerging Problems,” Information 2025, Vol. 16, Page 723, vol. 16, no. 9, p. 723, Aug. 2025, doi: 10.3390/INFO16090723. DOI: https://doi.org/10.3390/info16090723
  2. Z. Tan et al., “Neural machine translation: A review of methods, resources, and tools,” AI Open, vol. 1, pp. 5–21, Jan. 2020, doi: 10.1016/J.AIOPEN.2020.11.001. DOI: https://doi.org/10.1016/j.aiopen.2020.11.001
  3. A. Ramesh, V. B. Parthasarathy, R. Haque, and A. Way, “Comparing Statistical and Neural Machine Translation Performance on Hindi-To-Tamil and English-To-Tamil,” Digital 2021, Vol. 1, Pages 86-102, vol. 1, no. 2, pp. 86–102, Apr. 2021, doi: 10.3390/DIGITAL1020007. DOI: https://doi.org/10.3390/digital1020007
  4. A. Solyman et al., “Optimizing the impact of data augmentation for low-resource grammatical error correction,” Journal of King Saud University - Computer and Information Sciences, vol. 35, no. 6, Jun. 2023, doi: 10.1016/j.jksuci.2023.101572. DOI: https://doi.org/10.1016/j.jksuci.2023.101572
  5. S. Sen, M. Hasanuzzaman, A. Ekbal, P. Bhattacharyya, and A. Way, “Neural machine translation of low-resource languages using SMT phrase pair injection,” Nat. Lang. Eng., vol. 27, no. 3, pp. 271–292, May 2021, doi: 10.1017/S1351324920000303. DOI: https://doi.org/10.1017/S1351324920000303
  6. J. Dong, “Transfer Learning-Based Neural Machine Translation for Low-Resource Languages,” ACM Transactions on Asian and Low-Resource Language Information Processing, Sep. 2023, doi: 10.1145/3618111. DOI: https://doi.org/10.1145/3618111
  7. R. Kimera, D. N. Heo, D. N. Rim, and H. Choi, “Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda,” in NLPIR 2024 - 2024 8th International Conference on Natural Language Processing and Information Retrieval, Association for Computing Machinery, Inc, Apr. 2025, pp. 142–148. doi: 10.1145/3711542.3711594. DOI: https://doi.org/10.1145/3711542.3711594
  8. J. Zhang et al., “Neural Machine Translation for Low-Resource Languages from a Chinese-centric Perspective: A Survey,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 6, Jun. 2024, doi: 10.1145/3665244. DOI: https://doi.org/10.1145/3665244
  9. T. O. Tafa et al., “Machine Translation Performance for Low-Resource Languages: A Systematic Literature Review,” IEEE Access, vol. 13, pp. 72486–72505, 2025, doi: 10.1109/ACCESS.2025.3562918. DOI: https://doi.org/10.1109/ACCESS.2025.3562918
  10. W.-H. Her and U. Kruschwitz, “Investigating Neural Machine Translation for Low-Resource Languages: Using Bavarian as a Case Study,” Apr. 2024, Accessed: Nov. 25, 2025. [Online]. Available: http://arxiv.org/abs/2404.08259
  11. N. R. Haddaway, M. J. Page, C. C. Pritchard, and L. A. McGuinness, “PRISMA2020: An R package and Shiny app for producing PRISMA 2020-compliant flow diagrams, with interactivity for optimised digital transparency and Open Synthesis,” Campbell Systematic Reviews, vol. 18, no. 2, p. e1230, Jun. 2022, doi: 10.1002/CL2.1230. DOI: https://doi.org/10.1002/cl2.1230
  12. A. Javed et al., “Transformer-Based Re-Ranking Model for Enhancing Contextual and Syntactic Translation in Low-Resource Neural Machine Translation,” Electronics (Switzerland), vol. 14, no. 2, Jan. 2025, doi: 10.3390/electronics14020243. DOI: https://doi.org/10.3390/electronics14020243
  13. J. Pang et al., “Rethinking the Exploitation of Monolingual Data for Low-Resource Neural Machine Translation under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) license,” Computational Linguistics, vol. 50, no. 1, 2024, doi: 10.1162/coli. DOI: https://doi.org/10.1162/coli_a_00496
  14. T. N. Quoc, H. Le Thanh, and H. P. Van, “Khmer-Vietnamese Neural Machine Translation Improvement Using Data Augmentation Strategies,” Informatica (Slovenia), vol. 47, no. 3, pp. 349–360, Sep. 2023, doi: 10.31449/inf.v47i3.4761. DOI: https://doi.org/10.31449/inf.v47i3.4761
  15. H. Vu and N. D. Bui, “On the scalability of data augmentation techniques for low-resource machine translation between Chinese and Vietnamese,” Journal of Information and Telecommunication, vol. 7, no. 2, pp. 241–253, 2023, doi: 10.1080/24751839.2023.2186625. DOI: https://doi.org/10.1080/24751839.2023.2186625
  16. M. Maimaiti, Y. Liu, H. Luan, and M. Sun, “Enriching the Transfer Learning with Pre-Trained Lexicon Embedding for Low-Resource Neural Machine Translation,” 2022. DOI: https://doi.org/10.26599/TST.2020.9010029
  17. C. Samuel and I. T. Ali, “Batak Toba language-Indonesian machine translation with transfer learning using no language left behind,” International Journal of Advances in Applied Sciences, vol. 13, no. 4, pp. 830–839, Dec. 2024, doi: 10.11591/ijaas.v13.i4.pp830-839. DOI: https://doi.org/10.11591/ijaas.v13.i4.pp830-839
  18. L. Li, W. Hu, and M. Luo, “PNMT: Zero-Resource Machine Translation with Pivot-Based Feature Converter,” Computers, Materials and Continua, vol. 84, no. 3, pp. 5915–5935, 2025, doi: 10.32604/cmc.2025.064349. DOI: https://doi.org/10.32604/cmc.2025.064349
  19. X. Wu and R. Deng, “Research on the application of cross-language transfer learning model in English translation for low-resource scenarios,” IEEE Access, 2025, doi: 10.1109/ACCESS.2025.3628643. DOI: https://doi.org/10.1109/ACCESS.2025.3628643
  20. J. Pang et al., “Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models”, doi: 10.1162/tacl.
  21. B. K. Yazar and E. Kilic, “Improving Low-Resource Kazakh-English and Turkish-English Neural Machine Translation Using Transfer Learning and Part of Speech Tags,” IEEE Access, vol. 13, pp. 32341–32356, 2025, doi: 10.1109/ACCESS.2025.3542491. DOI: https://doi.org/10.1109/ACCESS.2025.3542491
  22. K. Ismail, S. Abdou, M. Farouk, and A. Salem, “Transformers to the rescue: alleviating data scarcity in arabic grammatical error correction with pre-trained models,” Neural Comput. Appl., vol. 37, no. 18, pp. 13011–13038, Jun. 2025, doi: 10.1007/s00521-025-11145-1. DOI: https://doi.org/10.1007/s00521-025-11145-1
  23. V. Lara-Ortiz, R. Q. Fuentes-Aguilar, and I. Chairez, “Spanish to Mexican Sign Language glosses corpus for natural language processing tasks,” Scientific Data , vol. 12, no. 1, Dec. 2025, doi: 10.1038/s41597-025-04871-7. DOI: https://doi.org/10.1038/s41597-025-04871-7
  24. X. Ma, M. Lan, W. Hu, and Y. Lu, “Dongba Machine Translation with Transfer Learning: Leveraging Pre-trained Ancient Chinese Models,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 24, no. 5, Apr. 2025, doi: 10.1145/3721980. DOI: https://doi.org/10.1145/3721980
  25. M. R. Costa-jussà et al., “Scaling neural machine translation to 200 languages,” Nature, vol. 630, no. 8018, pp. 841–846, Jun. 2024, doi: 10.1038/s41586-024-07335-x. DOI: https://doi.org/10.1038/s41586-024-07335-x
  26. W. Zhang, X. Li, Y. Yang, R. Dong, and V. Basile, “Pre-Training on Mixed Data for Low-Resource Neural Machine Translation,” 2021, doi: 10.3390/info1203. DOI: https://doi.org/10.3390/info12030133
  27. M. Tars, A. Tättar, and M. Fishel, “Cross-lingual Transfer from Large Multilingual Translation Models to Unseen Under-resourced Languages,” Baltic Journal of Modern Computing, vol. 10, no. 3, pp. 435–446, 2022, doi: 10.22364/bjmc.2022.10.3.16. DOI: https://doi.org/10.22364/bjmc.2022.10.3.16
  28. J. Yan, T. Lin, and S. Zhao, “Migration Learning and Multi-View Training for Low-Resource Machine Translation Migration Learning and Multi-View Training,” 2024. [Online]. Available: www.ijacsa.thesai.org DOI: https://doi.org/10.14569/IJACSA.2024.0150572
  29. Y. Wen, J. Guo, Z. Yu, and Z. Yu, “Chinese–Vietnamese Pseudo-Parallel Sentences Extraction Based on Image Information Fusion,” Information (Switzerland), vol. 14, no. 5, May 2023, doi: 10.3390/info14050298. DOI: https://doi.org/10.3390/info14050298
  30. A. Slim, A. Melouah, U. Faghihi, and K. Sahib, “Improving Neural Machine Translation for Low Resource Algerian Dialect by Transductive Transfer Learning Strategy,” Arab. J. Sci. Eng., vol. 47, no. 8, pp. 10411–10418, Aug. 2022, doi: 10.1007/s13369-022-06588-w. DOI: https://doi.org/10.1007/s13369-022-06588-w
  31. T. V. Ngo, P. T. Nguyen, V. V. Nguyen, T. Le Ha, and L. M. Nguyen, “An Efficient Method for Generating Synthetic Data for Low-Resource Machine Translation: An empirical study of Chinese, Japanese to Vietnamese Neural Machine Translation,” Applied Artificial Intelligence, vol. 36, no. 1, 2022, doi: 10.1080/08839514.2022.2101755. DOI: https://doi.org/10.1080/08839514.2022.2101755
  32. Z. Xu, S. Zhan, W. Yang, and Q. Xie, “Based on Gated Dynamic Encoding Optimization, the LGE-Transformer Method for Low-Resource Neural Machine Translation,” IEEE Access, vol. 12, pp. 162861–162869, 2024, doi: 10.1109/ACCESS.2024.3488186. DOI: https://doi.org/10.1109/ACCESS.2024.3488186
  33. M. Sun, H. Wang, M. Pasquine, and I. A. Hameed, “Machine translation in low-resource languages by an adversarial neural network,” Applied Sciences (Switzerland), vol. 11, no. 22, Nov. 2021, doi: 10.3390/app112210860. DOI: https://doi.org/10.3390/app112210860
  34. Y. Sang, Y. Chen, and J. Zhang, “Neural Machine Translation Research on Syntactic Information Fusion Based on the Field of Electrical Engineering,” Applied Sciences (Switzerland), vol. 13, no. 23, Dec. 2023, doi: 10.3390/app132312905. DOI: https://doi.org/10.3390/app132312905
  35. Z. Mao, C. Chu, and S. Kurohashi, “Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine Translation,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 21, no. 4, Jul. 2022, doi: 10.1145/3491065. DOI: https://doi.org/10.1145/3491065
  36. V. Karyukin, D. Rakhimova, A. Karibayeva, A. Turganbayeva, and A. Turarbek, “The neural machine translation models for the low-resource Kazakh–English language pair,” PeerJ Comput. Sci., vol. 9, 2023, doi: 10.7717/peerj-cs.1224. DOI: https://doi.org/10.7717/peerj-cs.1224