Instructions to use dsfsi/zabantu-xlm-roberta with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dsfsi/zabantu-xlm-roberta with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="dsfsi/zabantu-xlm-roberta")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("dsfsi/zabantu-xlm-roberta") model = AutoModelForMaskedLM.from_pretrained("dsfsi/zabantu-xlm-roberta", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Zabantu - Exploring Multilingual Language Model training for South African Bantu Languages
Zabantu-XLM-RoBERTa is a multilingual masked language model developed to investigate whether cross-lingual learning, via pretraining on related Bantu languages, can improve NLP coverage for Tshivenda, a language with very little text of its own. It uses the XLM-RoBERTa architecture and was pre-trained from scratch, inspired by AfriBERTa, developed by Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. Their paper, Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages (2021), demonstrated that competitive multilingual language models could be trained from scratch on limited amounts of text from African languages. This work applies cross-lingual learning to advance NLP applications in Tshivenda, and it also serves as a benchmark for future work on other Bantu languages.
The word Zabantu is derived from: "Za" for South Africa, and "bantu" for Bantu languages
This repository hosts the Zabantu large variant, trained on nine South African Bantu languages: Tshivenda, Sepedi, Sesotho, Setswana, Xitsonga, isiNdebele, isiXhosa, isiZulu, and siSwati. The wider Zabantu project explores different combinations of Bantu languages through monolingual, bilingual, and multilingual experiments.
Model details
| Property | This checkpoint |
|---|---|
| Hub ID | dsfsi/zabantu-xlm-roberta |
| Model class | XLMRobertaForMaskedLM |
| Training objective | Masked language modelling (MLM) |
| Transformer layers | 8 |
| Hidden size | 768 |
| Attention heads per layer | 6 |
| Feed-forward intermediate size | 3,072 |
| Vocabulary size | 250,002, including special tokens |
| Maximum sequence length | 512 tokens |
| Tokenizer | SentencePiece BPE |
| Model size | ~250M parameters |
| Licence | CC BY 4.0 |
See config.json for the architecture configuration.
Usage example(s)
from transformers import pipeline
unmasker = pipeline('fill-mask', model='dsfsi/zabantu-xlm-roberta')
sample_sentences = {
'zulu': "Le ndoda ithi izo____ ukudla.", # Masked word for Zulu
'tshivenda': "Mufana uyo____ vhukuma.", # Masked word for Tshivenda
'sepedi': "Mosadi o ____ pheka.", # Masked word for Sepedi
'tswana': "Monna o ____ tsamaya.", # Masked word for Tswana
'tsonga': "N'wana wa xisati u ____ ku tsaka." # Masked word for Tsonga
}
for language, sentence in sample_sentences.items():
masked_sentence = sentence.replace('____', unmasker.tokenizer.mask_token)
# Get the model predictions
results = unmasker(masked_sentence)
print(f"Original sentence ({language}): {sentence}")
print(f"Top prediction for the masked token: {results[0]['sequence']}\n")
Model Variants
- Zabantu-VEN: A monolingual language model trained on 73k raw sentences in Tshivenda
- Zabantu-NSO: A monolingual language model trained on 179k raw sentences in Sepedi
- Zabantu-NSO+VEN: A bilingual language model trained on 179k raw sentences in Sepedi and 73k sentences in Tshivenda
- Zabantu-SOT+VEN: A multilingual language model trained on 479k raw sentences from Sesotho, Sepedi, Setswana, and Tshivenda
- Zabantu-BANTU: A multilingual language model trained on 1.4M raw sentences from 9 South African Bantu languages
Intended Use
Like any Masked Language Model (MLM), Zabantu models can be adapted to a variety of semantic tasks such as:
- Text Classification/Categorization: Assigning categories or labels to a whole document, or sections of a document, based on its content.
- Sentiment Analysis: Determining the sentiment of a text, such as whether the opinion is positive, negative, or neutral.
- Named Entity Recognition (NER): Identifying and classifying key information (entities) in text into predefined categories such as the names of people, organizations, locations, expressions of times, quantities, monetary values, percentages, etc.
- Part-of-Speech Tagging (POS): Assigning word types to each word (like noun, verb, adjective, etc.), based on both its definition and its context.
- Semantic Text Similarity: Measuring how similar two pieces of texts are, which is useful in various applications such as information retrieval, document clustering, and duplicate detection.
- etc.
Performance and Limitations
- Performance: The Zabantu models demonstrate promising performance on various NLP tasks, including news topic classification with competitive results compared to similar pre-trained cross-lingual models such as AfriBERTa and AfroXLMR.
Monolingual test F1 scores on News Topic Classification
| Weighted F1 [%] | Afriberta-large | Afroxlmr | zabantu-nsoven | zabantu-sotven | zabantu-bantu |
|---|---|---|---|---|---|
| nso | 71.4 | 71.6 | 74.3 | 69 | 70.6 |
| ven | 74.3 | 74.1 | 77 | 76 | 75.6 |
Few-shot(50 shots) test F1 scores on News Topic Classification
| Weighted F1 [%] | Afriberta | Afroxlmr | zabantu-nsoven | zabantu-sotven | zabantu-bantu |
|---|---|---|---|---|---|
| ven | 60 | 62 | 66 | 69 | 55 |
Limitations:
Although efforts have been made to include a wide range of South African languages, the model's coverage may still be limited for certain dialects. We note that the training set was largely dominated by Setwana and IsiXhosa.
We also acknowledge the potential to further improve the model by training it on more data, including additional domains and topics.
As with any language model, the generated output should be carefully reviewed and post-processed to ensure accuracy and cultural sensitivity.
Training Data
The models have been trained on a large corpus of text data collected from various sources, including SADiLaR, Leipnets, Flores, CC-100, Opus and various South African government websites. The training data covers a wide range of topics and domains, notably religion, politics, academics and health (mostly Covid-19).
Citation
If you use Zabantu in research, please cite:
@mastersthesis{nemakhavhani2023tshivenda,
author = {Nemakhavhani, Ndamulelo},
title = {Exploring cross-lingual learning techniques for advancing Tshivenda NLP coverage},
school = {University of Pretoria},
type = {Mini-dissertation, MIT (Big Data Science)},
year = {2023},
note = {Supervised by Vukosi Marivate and Jocelyn Mazarura},
url = {https://repository.up.ac.za/handle/2263/98198}
}
Background
Zabantu was developed in Ndamulelo Nemakhavhani's 2023 mini-dissertation at the University of Pretoria, supervised by Prof. Vukosi Marivate and Dr Jocelyn Mazarura.
- Downloads last month
- 71