Skip to main navigation Skip to search Skip to main content

Construction of Large Language Models for Taigi and Hakka Using Transfer Learning

  • Yen Chun Lai
  • , Yi Jun Zheng
  • , Wen Han Hsu
  • , Yan Ming Lin
  • , Cheng Hsiu Cho
  • , Chlh Chung Kuo
  • , Chao Shih Huang
  • , Yuan Fu Liao

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Taigi (also known as Taiwanese Hokkien) and Hakka were once the predominant native languages in Taiwan, but they are now on the verge of extinction. To address this crisis, this study explores how to construct large language models (LLMs) for Taigi and Hakka to aid in their revitalization despite the scarcity of linguistic resources (the low-resource language issue). Our approach leverages pre-trained models from LLMs of resource-rich languages, following three key steps: (1) introducing specific characters for Taigi and Hakka into the LLM's tokenizer, enabling the model to recognize the orthography of both languages; (2) conducting continual pre-training with a small amount of Taigi and Hakka text data to familiarize the LLM with the lexical choices of these languages; and (3) generating task-specific training data via machine translation for instruction fine-tuning, allowing the LLM to perform tasks in Taigi and Hakka. Finally, we evaluated the performance of the constructed LLMs using a reading comprehension question bank (consisting of 1,000 multiple-choice questions) developed by Taiwanese and Hakka language experts. The experimental results indicate that based on LLaMA2-70B, the LLMs achieved accuracy rates of 85.6% for Taigi and 87.7% for Hakka, respectively, suggesting that the performance is approaching human-level proficiency.

Original languageEnglish
Title of host publication2024 27th Conference on the Oriental COCOSDA International Committee for the Co-Ordination and Standardisation of Speech Databases and Assessment Techniques, O-COCOSDA 2024 - Proceedings
EditorsMing-Hsiang Su, Jui-Feng Yeh, Yuan-Fu Liao, Chi-Chun Lee, Yu Taso
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331506032
DOIs
StatePublished - 2024
Event27th Conference on the Oriental COCOSDA International Committee for the Co-Ordination and Standardisation of Speech Databases and Assessment Techniques, O-COCOSDA 2024 - Hsinchu, Taiwan
Duration: 17 Oct 202419 Oct 2024

Publication series

Name2024 27th Conference on the Oriental COCOSDA International Committee for the Co-Ordination and Standardisation of Speech Databases and Assessment Techniques, O-COCOSDA 2024 - Proceedings

Conference

Conference27th Conference on the Oriental COCOSDA International Committee for the Co-Ordination and Standardisation of Speech Databases and Assessment Techniques, O-COCOSDA 2024
Country/TerritoryTaiwan
CityHsinchu
Period17/10/2419/10/24

Keywords

  • Hakka
  • Taigi
  • computer reading comprehension test
  • large language model
  • machine translation

Fingerprint

Dive into the research topics of 'Construction of Large Language Models for Taigi and Hakka Using Transfer Learning'. Together they form a unique fingerprint.

Cite this