Enhancing Low-Resource Lampung Speech Recognition through Cross-Lingual XLSR-Wav2Vec 2.0 Pretraining
Abstract
This study investigates the application of Wav2Vec 2.0 (W2V2) and Cross-Lingual Speech Representation (XLSR) models to Lampung language speech recognition. LampungNyow v1.0 is introduced, a speech corpus designed to provide a baseline for training and evaluating Automatic Speech Recognition (ASR) for this low-resource regional language of Indonesia. The dataset enables supervised fine-tuning and standardized evaluation, addressing the lack of publicly available linguistic resources for Lampung. Several pre-trained W2V2 models on Lampung speech recognition using Word Error Rate (WER) as the evaluation metric. The evaluated models include W2V2-Base, W2V2-Large, W2V2-Large-XLSR-Indonesian, W2V2-Large-XLSR-Sundanese, W2V2-Large-XLSR-53, and the multilingual W2V2-Large-XLSR-Indonesia-Javanese-Sundanese model. Monolingual models have higher WER values, according to experimental results: W2V2-Base achieved 36,23%, while W2V2-Large achieved 36,30%. XLSR models, such as XLSR-53 (33,88%), Sundanese (33,99%), and Indonesian (33,70%), demonstrated modest improvements. The W2V2-Large-XLSR-Indonesian-Javanese-Sundanese model, which was the foundation for the Lampung automatic speech recognition system in this study, achieved lower WER of 17,39%. These findings suggest that, in contrast to more comprehensive multilingual or monolingual pretraining models, multilingual pretraining utilizing a number of Indonesian regional languages can produce acoustic and contextual speech representations that are better suited for the resource-constrained Lampung automatic speech recognition task. When compared to the baseline W2V2-Large model, the obtained WER of 17,39% indicates a relative improvement of more than 50%.
Keywords
Full Text:
PDFReferences
D. Yu and L. Deng, Automatic Speech Recognition: A Deep Learning Approach. in Signals and Communication Technology. London: Springer London, 2015. doi: 10.1007/978-1-4471-5779-3.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. in Adaptive Computation and Machine Learning series. Cambridge, MA, USA: MIT Press, 2016. Accessed: Aug. 17, 2025. [Online]. Available: https://mitpress.mit.edu/9780262035613/deep-learning/
S. Karpagavalli and E. Chandra, “A Review on Automatic Speech Recognition Architecture and Approaches,” IJSIP, vol. 9, no. 4, pp. 393–404, Apr. 2016, doi: 10.14257/ijsip.2016.9.4.34.
Y. He et al., “Streaming End-to-end Speech Recognition for Mobile Devices,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2019, pp. 6381–6385. doi: 10.1109/ICASSP.2019.8682336.
M. Schuster, “Speech Recognition for Mobile Devices at Google,” in PRICAI 2010: Trends in Artificial Intelligence, B.-T. Zhang and M. A. Orgun, Eds., Berlin, Heidelberg: Springer, 2010, pp. 8–10. doi: 10.1007/978-3-642-15246-7_3.
H. H. O. Nasereddin and A. A. R. Omari, “Classification techniques for automatic speech recognition (ASR) algorithms used with real time speech translation,” in 2017 Computing Conference, Jul. 2017, pp. 200–207. doi: 10.1109/SAI.2017.8252104.
X. Huang, A. Acero, H.-W. Hon, and R. Reddy, Spoken Language Processing: A Guide to Theory, Algorithm, and System Development, 1st ed. USA: Prentice Hall PTR, 2001.
D. O’Shaughnessy, “Trends and developments in automatic speech recognition research,” Computer Speech & Language, vol. 83, p. 101538, Jan. 2024, doi: 10.1016/j.csl.2023.101538.
J. Wang and Y. Chen, Introduction to Transfer Learning: Algorithms and Practice. in Machine Learning: Foundations, Methodologies, and Applications. Singapore: Springer Nature Singapore, 2023. doi: 10.1007/978-981-19-7584-4.
X. Wei, C. Dong, and Y. Xu, “The Mechanical and System Design of Finger Training Rehabilitation Device Based on Speech Recognition,” Journal of Applied Data Sciences, vol. 3, no. 2, pp. 60–65, May 2022.
H. I. Pratiwi, I. H. Kartowisastro, B. Soewito, and W. Budiharto, “Dispute on Security Framework Model of MFCC Mixed Methods in Speech Recognition System,” Journal of Applied Data Sciences, vol. 6, no. 3, pp. 1542–1550, Jun. 2025.
Badan Pusat Statistik, “Jumlah Penduduk Pertengahan Tahun - Tabel Statistik.” Accessed: Feb. 14, 2026. [Online]. Available: https://www.bps.go.id/id/statistics-table/2/MTk3NSMy/jumlah-penduduk-pertengahan-tahun--ribu-jiwa-.html
A. F. Aji et al., “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio, Eds., Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 7226–7249. doi: 10.18653/v1/2022.acl-long.500.
Wagiati, N. Darmayanti, Y. Yohanarisagarniwa, and D. Zein, “Mapping the Dimensions of Linguistic Distance: A Study on Quantitative and Qualitative Geolinguistics of Banjar Sundanese Dialect,” European Journal of Language and Culture Studies, vol. 2, no. 4, pp. 8–17, Jul. 2023, doi: 10.24018/ejlang.2023.2.4.87.
O. Kjartansson, S. Sarin, K. Pipatsrisawat, M. Jansche, and L. Ha, “Crowd-Sourced Speech Corpora for Javanese, Sundanese, Sinhala, Nepali, and Bangladeshi Bengali,” in 6th Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU 2018), ISCA, Aug. 2018, pp. 52–55. doi: 10.21437/SLTU.2018-11.
H. Nasution, R. Rahayu, E. Wibowo, D. M. Harum, and F. Moses, Persebaran Bahasa-Bahasa di Provinsi Lampung. Bandar Lampung: Kantor Bahasa Provinsi Lampung, 2008.
A. E. Sanusi, S. Zamzanah, M. Widodo, I. Sunarti, and S. Samhati, Kata Tugas Bahasa Lampung Dialek Abung. Jakarta: Pusat Pembinaan dan Pengembangan Bahasa, Departemen Pendidikan dan Kebudayaan, 1997.
N. W. Putri, “Pergeseran Bahasa Daerah Lampung Pada Masyarakat Kota Bandar Lampung,” PRASASTI: Journal of Linguistics, vol. 3, no. 1, Art. no. 1, 2018, doi: 10.20961/prasasti.v3i1.16550.
R. Dewi, F. Ariyani, and N. E. Rusminto, “Pemertahanan Bahasa Lampung dalam Ranah Pendidikan,” Tiyuh Lampung, vol. 3, no. 1, pp. 48–57, 2018.
A. N. Agency, “Kantor Bahasa sebut jumlah penutur Bahasa Lampung 6.250 orang,” Antara News Lampung. Accessed: Feb. 14, 2026. [Online]. Available: https://lampung.antaranews.com/berita/706452/kantor-bahasa-sebut-jumlah-penutur-bahasa-lampung
C. Moseley, Atlas of the World’s Languages in Danger. UNESCO Digital Library, 2012.
A. Raharjo and A. Zahra, “Javanese and Sundanese speech recognition using Whisper,” Computer Science and Information Technologies, vol. 6, no. 3, pp. 253–261, Nov. 2025, doi: 10.11591/csit.v6i3.p253-261.
A. Ramadhan, A. Junaidi, Aristoteles, F. R. Lumbanraja, and A. Faisol, “Implementation of MFCC Features Extraction of Recurrent Neural Network and Bidirectional LSTM for Speech to Text Transcription: A Case Study on the Lampung Language Dialect Api,” in Proceedings of the 5th International Conference on Applied Sciences, Mathematics, and Informatics (ICASMI), Lampung: Springer, 2024. doi: https//doi.org/10.2991/978-94-6463-730-4_10.
M. Zainudin, A. Junaidi, F. R. Lumbanraja, Aristoteles, and Tristiyanto, “Perceptual Linear Prediction Features Extraction For Lampung Voice to Text Transcription (Case Study: Lampung Pepadun Dialect A),” in Proceedings of the 5th International Conference on Applied Sciences, Mathematics, and Informatics (ICASMI), Lampung: Springer, 2024. doi: https//doi.org/10.2991/978-94-6463-730-4_10.
B. Gold, N. Morgan, and D. Ellis, Speech and Audio Signal Processing: Processing and Perception of Speech and Music, 1st ed. Wiley, 2011. doi: 10.1002/9781118142882.
M. Johnson and T. Taniguchi, “Online and offline computational reduction techniques using backward filtering in CELP speech coders,” IEEE Transactions on Signal Processing, vol. 40, no. 8, pp. 2090–2093, Aug. 1992, doi: 10.1109/78.149977.
H. Fletcher, “The Nature of Speech and its Interpretations,” Bell System Technical Journal, vol. 1, pp. 129–144, 1922.
H. Dudley, R. R. Riesz, and S. S. A. Watkins, “A synthetic speaker,” Journal of the Franklin Institute, vol. 227, no. 6, pp. 739–764, Jun. 1939, doi: 10.1016/S0016-0032(39)90816-1.
K. H. Davis, R. Biddulph, and S. Balashek, “Automatic Recognition of Spoken Digits,” J. Acoust. Soc. Am., vol. 24, no. 6, pp. 637–642, Nov. 1952, doi: 10.1121/1.1906946.
H. F. Olson and H. Belar, “Phonetic Typewriter,” J. Acoust. Soc. Am., vol. 28, no. 6, pp. 1072–1081, Nov. 1956, doi: 10.1121/1.1908561.
J. W. Forgie and C. D. Forgie, “Results Obtained from a Vowel Recognition Computer Program,” J. Acoust. Soc. Am., vol. 31, no. 6_Supplement, p. 844, Jun. 1959, doi: 10.1121/1.1936151.
J. Suzuki and K. Nakata, “Recognition of Japanese Vowels—Preliminary to the Recognition of Speech,” J. Radio Res. Lab, vol. 37, no. 8, pp. 193–212, 1961.
T. Sakai and S. Doshita, “Phonetic Typewriter,” The Journal of the Acoustical Society of America, vol. 33, no. 11_Supplement, pp. 1664–1664, Nov. 1961, doi: 10.1121/1.1936652.
K. Nagata, Y. Kato, and S. Chiba, “Spoken Digit Recognizer for The Japanese Language,” NEC Res. Develop, no. 6, 1963.
P. Denes, “The design and operation of the mechanical speech recognizer at University College London,” Journal of the British Institution of Radio Engineers, vol. 19, no. 4, pp. 219–229, Apr. 1959, doi: 10.1049/jbire.1959.0027.
B. S. Atal and S. L. Hanauer, “Speech Analysis and Synthesis by Linear Prediction of the Speech Wave,” J. Acoust. Soc. Am, vol. 50, no. 2, pp. 637–655, 1971.
F. Itakura and S. Saito, “A statistical method for estimation of speech spectral density and formant frequencies,” Electronics and Communications in Japan, vol. 53A, pp. 36–43, 1970.
F. Itakura, “Minimum prediction residual principle applied to speech recognition,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 23, no. 1, pp. 67–72, Feb. 1975, doi: 10.1109/TASSP.1975.1162641.
L. Rabiner, S. Levinson, A. Rosenberg, and J. Wilpon, “Speaker-independent recognition of isolated words using clustering techniques,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 4, pp. 336–349, Aug. 1979, doi: 10.1109/TASSP.1979.1163259.
M. Khudhair and A. Talib, “Improving Low Resources Arabic Speech Recognition using Data Augmentation,” in 2022 Fifth College of Science International Conference of Recent Trends in Information Technology (CSCTIT), 2022, pp. 60–65. doi: 10.1109/CSCTIT56299.2022.10145613.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, no. 1, pp. 30–42, Jan. 2012, doi: 10.1109/TASL.2011.2134090.
R. Sobti, K. Guleria, and V. Kadyan, “Automatic Speech Recognition System for Low Resource Punjabi Language using Deep Neural Network-Hidden Markov Model (DNN-HMM),” International Journal of Intelligent Systems and Applications in Engineering, vol. 12, no. 19s, Art. no. 19s, Mar. 2024.
T. G. Fantaye, J. Yu, and T. T. Hailu, “Advanced Convolutional Neural Network-Based Hybrid Acoustic Models for Low-Resource Speech Recognition,” Computers, vol. 9, no. 2, p. 36, May 2020, doi: 10.3390/computers9020036.
T. N. Sainath, A. Mohamed, B. Kingsbury, and B. Ramabhadran, “Deep convolutional neural networks for LVCSR,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, May 2013, pp. 8614–8618. doi: 10.1109/ICASSP.2013.6639347.
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, May 2013, pp. 6645–6649. doi: 10.1109/ICASSP.2013.6638947.
D. Amodei et al., “Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin,” in Proceedings of The 33rd International Conference on Machine Learning, PMLR, Jun. 2016, pp. 173–182. Accessed: Aug. 25, 2025. [Online]. Available: https://proceedings.mlr.press/v48/amodei16.html
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 4945–4949. doi: 10.1109/ICASSP.2016.7472618.
E. Battenberg et al., “Exploring neural transducers for end-to-end speech recognition,” in 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Dec. 2017, pp. 206–213. doi: 10.1109/ASRU.2017.8268937.
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A Comparison of Sequence-to-Sequence Models for Speech Recognition,” in Interspeech 2017, ISCA, Aug. 2017, pp. 939–943. doi: 10.21437/Interspeech.2017-233.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2, in NIPS’14, vol. 2. Cambridge, MA, USA: MIT Press, Dec. 2014, pp. 3104–3112.
E. Trentin and M. Gori, “A survey of hybrid ANN/HMM models for automatic speech recognition,” Neurocomputing, vol. 37, no. 1, pp. 91–126, Apr. 2001, doi: 10.1016/S0925-2312(00)00308-8.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2016, pp. 4960–4964. doi: 10.1109/ICASSP.2016.7472621.
A. Gulati et al., “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Interspeech 2020, ISCA, Oct. 2020, pp. 5036–5040. doi: 10.21437/Interspeech.2020-3015.
A. Hannun et al., “Deep Speech: Scaling up end-to-end speech recognition,” Dec. 19, 2014, arXiv: arXiv:1412.5567. doi: 10.48550/arXiv.1412.5567.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” in Proceedings of the 34th International Conference on Neural Information Processing Systems, Red Hook, NY, USA: Curran Associates Inc., Dec. 2020, pp. 12449–12460.
W.-H. Tsai, P. L. Thi, T.-C. Tai, C.-L. Huang, and J.-C. Wang, “Low-Resource Speech Recognition Based on Transfer Learning,” in 2022 RIVF International Conference on Computing and Communication Technologies (RIVF), Ho Chi Minh City, Vietnam: IEEE, Dec. 2022, pp. 145–149. doi: 10.1109/RIVF55975.2022.10013881.
J. Lehečka, J. V. Psutka, and J. Psutka, “Transfer Learning of Transformer-Based Speech Recognition Models from Czech to Slovak,” in Text, Speech, and Dialogue, K. Ekštein, F. Pártl, and M. Konopík, Eds., Cham: Springer Nature Switzerland, 2023, pp. 328–338. doi: 10.1007/978-3-031-40498-6_29.
A. R. Bharadwaj, K. Laxmi Srina, K. Snuhith, R. Bhavani, and M. A. Jabbar, “Telugu Language Low-resource ASR Fine-Tuning with Wav2Vec2-XLS-R-300 and Interface Development,” in 2025 Global Conference in Emerging Technology (GINOTECH), 2025, pp. 1–6. doi: 10.1109/GINOTECH63460.2025.11077079.
A. Cryssiover and A. Zahra, “Speech recognition model design for Sundanese language using WAV2VEC 2.0,” Int J Speech Technol, vol. 27, no. 1, pp. 171–177, Mar. 2024, doi: 10.1007/s10772-023-10066-5.
P. Arisaputra, A. T. Handoyo, and A. Zahra, “XLS-R Deep Learning Model for Multilingual ASR on Low-Resource Languages: Indonesian, Javanese, and Sundanese,” ICIC Express Letters, Part B: Applications, vol. 15, no. 6, pp. 551–559, 2024, doi: 10.24507/icicelb.15.06.551.
P. Arisaputra and A. Zahra, “Indonesian Automatic Speech Recognition with XLSR-53,” ISI, vol. 27, no. 6, pp. 973–982, Dec. 2022, doi: 10.18280/isi.270614.
A. Bawitlung, S. K. Dash, and R. M. Pattanayak, “Mizo Automatic Speech Recognition: Leveraging Wav2vec 2.0 and XLS-R for Enhanced Accuracy in Low-Resource Language Processing,” ACM Trans. Asian Low-Resour. Lang. Inf. Process., vol. 24, no. 7, p. 72:1-72:15, Jul. 2025, doi: 10.1145/3746063.
L. R. Stefanel Gris, E. Casanova, F. S. De Oliveira, A. Da Silva Soares, and A. Candido Junior, “Brazilian Portuguese Speech Recognition Using Wav2vec 2.0,” in Computational Processing of the Portuguese Language, vol. 13208, V. Pinheiro, P. Gamallo, R. Amaro, C. Scarton, F. Batista, D. Silva, C. Magro, and H. Pinto, Eds., in Lecture Notes in Computer Science, vol. 13208. , Cham: Springer International Publishing, 2022, pp. 333–343. doi: 10.1007/978-3-030-98305-5_31.
J. C. Vásquez-Correa and A. Álvarez Muniain, “Novel Speech Recognition Systems Applied to Forensics within Child Exploitation: Wav2vec2.0 vs. Whisper,” Sensors, vol. 23, no. 4, p. 1843, Jan. 2023, doi: 10.3390/s23041843.
A. Barcovschi, R. Jain, and P. Corcoran, “A comparative analysis between Conformer-Transducer, Whisper, and wav2vec2 for improving the child speech recognition,” in 2023 International Conference on Speech Technology and Human-Computer Dialogue (SpeD), Oct. 2023, pp. 42–47. doi: 10.1109/SpeD59241.2023.10314867.
K. Bauyrzhan, M. Madina, and O. Assel, “Fine-Tuning the Wav2vec2 Model for Kazakh Speech: A Study on a Limited Corpus,” in 2023 IEEE International Conference on Smart Information Systems and Technologies (SIST), May 2023, pp. 124–128. doi: 10.1109/SIST58284.2023.10223504.
U. Kimanuka, C. wa Maina, and O. Büyük, “Speech recognition datasets for low-resource Congolese languages,” Data in Brief, vol. 52, p. 109796, Feb. 2024, doi: 10.1016/j.dib.2023.109796.
A. Vaswani et al., “Attention is All you Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA: Curran Associates Inc., Dec. 2017, pp. 6000–6010.
L. Chen, M. Asgari, and H. H. Dodge, “Optimize Wav2vec2s Architecture for Small Training Set Through Analyzing its Pre-Trained Models Attention Pattern,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2022, pp. 7112–7116. doi: 10.1109/ICASSP43922.2022.9747831.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning, in ICML ’06. New York, NY, USA: Association for Computing Machinery, Jun. 2006, pp. 369–376. doi: 10.1145/1143844.1143891.
P. Von Platen, “Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with Transformers.” Accessed: Feb. 17, 2026. [Online]. Available: https://huggingface.co/blog/fine-tune-xlsr-wav2vec2
“facebook/wav2vec2-base · Hugging Face.” Accessed: Feb. 18, 2026. [Online]. Available: https://huggingface.co/facebook/wav2vec2-base
“facebook/wav2vec2-large · Hugging Face.” Accessed: Feb. 18, 2026. [Online]. Available: https://huggingface.co/facebook/wav2vec2-large
“facebook/wav2vec2-large-xlsr-53 · Hugging Face.” Accessed: Feb. 18, 2026. [Online]. Available: https://huggingface.co/facebook/wav2vec2-large-xlsr-53
“indonesian-nlp/wav2vec2-large-xlsr-indonesian · Hugging Face.” Accessed: Feb. 18, 2026. [Online]. Available: https://huggingface.co/indonesian-nlp/wav2vec2-large-xlsr-indonesian
“cahya/wav2vec2-large-xlsr-sundanese · Hugging Face.” Accessed: Feb. 18, 2026. [Online]. Available: https://huggingface.co/cahya/wav2vec2-large-xlsr-sundanese
“indonesian-nlp/wav2vec2-indonesian-javanese-sundanese · Hugging Face.” Accessed: Feb. 18, 2026. [Online]. Available: https://huggingface.co/indonesian-nlp/wav2vec2-indonesian-javanese-sundanese
H. Aldarmaki, A. Ullah, S. Ram, and N. Zaki, “Unsupervised Automatic Speech Recognition: A review,” Speech Communication, vol. 139, pp. 76–91, Apr. 2022, doi: 10.1016/j.specom.2022.02.005.
S. Jothilakshmi and V. N. Gudivada, “Chapter 10 - Large Scale Data Enabled Evolution of Spoken Language Research and Applications,” in Handbook of Statistics, vol. 35, V. N. Gudivada, V. V. Raghavan, V. Govindaraju, and C. R. Rao, Eds., in Cognitive Computing: Theory and Applications, vol. 35. , Elsevier, 2016, pp. 301–340. doi: 10.1016/bs.host.2016.07.005.
DOI: https://doi.org/10.47738/jads.v7i3.1388
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)