dialect-router-v0.1 / README.md
yrrhall's picture oddadmix's picture
Duplicate from oddadmix/dialect-router-v0.1
a00f776
|
Raw History Blame Contribute Delete
4.78 kB
---
language:
- ar
license: mit
tags:
- arabic
- dialect-identification
- text-classification
- arabic-nlp
- arabic-dialect
- nlp
- sequence-classification
- lahgtna
pipeline_tag: text-classification
library_name: transformers
datasets:
- custom
metrics:
- accuracy
- f1
model-index:
- name: dialect-router-v0.1
results:
- task:
type: text-classification
name: Arabic Dialect Identification
metrics:
- type: accuracy
value: null
name: Accuracy
- type: f1
value: null
name: Macro F1
base_model:
- asafaya/bert-mini-arabic
---
# dialect-router-v0.1
A lightweight Arabic dialect identification model that classifies input text into one of **11 Arabic dialect / language codes**. It is used as the routing backbone in the [Lahgtna](https://github.com/Oddadmix/lahgtna-chatterbox) pipeline to automatically select the correct voice reference and Chatterbox language token for speech synthesis.
---
## Model Details
| Property | Value |
|---|---|
| **Architecture** | Transformer encoder (sequence classification head) |
| **Task** | Multi-class text classification |
| **Input** | Raw Arabic text (up to 512 tokens) |
| **Output** | One of 11 dialect codes |
| **Language** | Arabic (`ar`) |
| **License** | MIT |
### Dialect Labels
| Label | Dialect | Region |
|---|---|---|
| `eg` | Egyptian | Egypt |
| `sa` | Saudi | Saudi Arabia |
| `mo` | Moroccan (Darija) | Morocco |
| `iq` | Iraqi | Iraq |
| `sd` | Sudanese | Sudan |
| `tn` | Tunisian | Tunisia |
| `lb` | Lebanese | Lebanon |
| `sy` | Syrian | Syria |
| `ly` | Libyan | Libya |
| `ps` | Palestinian | Palestine |
| `ar` | Modern Standard Arabic (MSA) | — |
---
## Intended Use
### Primary use
Dialect-aware TTS routing — given an Arabic utterance, predict the dialect so the correct speaker reference audio and Chatterbox language code can be selected automatically.
### Secondary use
Standalone Arabic dialect identification for NLP pipelines, content filtering, dataset analysis, or any application that needs to distinguish Arabic dialects programmatically.
### Out-of-scope use
- Non-Arabic languages
- Code-switched text (Arabic + English mixed)
- Dialect intensity scoring or fine-grained subdialect classification
- High-stakes decisions without human review
---
## How to Use
### Direct inference
```python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "oddadmix/dialect-router-v0.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
text = "اه ياراسي الواحد دماغه وجعاه"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
pred_id = torch.argmax(logits, dim=-1).item()
dialect = model.config.id2label[pred_id]
print(dialect) # e.g. "eg"
```
### With the Transformers pipeline
```python
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="oddadmix/dialect-router-v0.1",
)
result = classifier("اه ياراسي الواحد دماغه وجعاه")
print(result)
# [{'label': 'eg', 'score': 0.94}]
```
### Inside Lahgtna TTS
```python
from inference import run_pipeline
# Dialect is detected automatically
run_pipeline(
text="اه ياراسي الواحد دماغه وجعاه",
output_path="output.wav",
)
```
---
## Limitations & Biases
- **Short texts** (< 5 tokens) may produce unreliable predictions — the model benefits from sentence-length input.
- **Code-switched text** (e.g. Arabic + French in Moroccan Darija, or Arabic + English) may confuse the classifier.
- **Dialect continuum** — dialects from geographically adjacent regions (e.g. `sy` / `lb`, `eg` / `ly`) may be confused by the model.
- **Corpus bias** — label distribution in training data may not reflect real-world dialect prevalence; some dialects (e.g. `sd`, `ly`) may have lower recall.
- This model should **not** be used for identity classification of individuals.
---
## Citation
If you use this model in your research or product, please cite:
```bibtex
@misc{lahgtna-dialect-router-2025,
title = {dialect-router-v0.1: Arabic Dialect Identification for TTS Routing},
author = {Oddadmix},
year = {2025},
url = {https://huggingface.co/oddadmix/dialect-router-v0.1}
}
```
---
## Related Resources
- 🔊 **Lahgtna TTS checkpoint** → [`oddadmix/lahgtna-chatterbox-v1`](https://huggingface.co/oddadmix/lahgtna-chatterbox-v1)
- 💻 **Inference code** → [lahgtna-tts on GitHub](https://github.com/Oddadmix/lahgtna-chatterbox)
- 🗣️ **Chatterbox backbone** → [resemble-ai/chatterbox](https://github.com/resemble-ai/chatterbox)