LIDstral Arabic

LIDstral Arabic is a fast language and dialect classifier for text written in Arabic script. It identifies Modern Standard Arabic, all Arabic dialects, and non-Arabic languages that use the same script.

We include these non-Arabic languages in training to help the model distinguish languages that share a script and avoid classifying text as Arabic based on its script alone.

Method

The model combines a fastText one-vs-all (OVA) classifier, an MLP stacker, and per-class isotonic calibration:

  1. fastText OVA: produces independent scores for each class, preserving evidence for competing languages and dialects.
  2. MLP stacker: A 64-unit LayerNorm MLP maps these scores to probabilities over the classes. It uses 144 features derived from raw scores, prior corrections, confidence margins, entropy, and regional summaries. The checkpoint stores feature normalization and label ordering.
  3. Per-class isotonic regressors: fitted on development predictions calibrate the MLP probabilities. The pipeline then renormalizes them across classes. Calibration can change both confidence and the predicted label.

The full pipeline requires three files:

  • lid_ova_full.bin
  • mlp_stacker_ln.pt
  • isotonic_calibration.pkl

classify.py runs the pipeline on CPU, including preprocessing and feature extraction. Each input that remains nonempty after preprocessing receives a predicted label and calibrated probabilities over all 51 classes.

Evaluation

We evaluate on 84,870 Arabic examples drawn from various benchmarks including NADI, MADAR, QADI, Casablanca, FLEURS, SMOL, Omnilingual ASR, and Hassaniya. On Moroccan Darija, LIDstral Arabic achieves 88.67% F1, compared with 72.65% for LahjatBERT ALDi CL, the strongest tested baseline for this dialect, and 71.28% for GlotLID v3.

Per-label F1 on the Arabic benchmark

Usage

Download the private repository with a Hugging Face account that has access:

pip install huggingface_hub
hf auth login
hf download mistralai/LIDstral-Arabic --local-dir LIDstral-Arabic
cd LIDstral-Arabic
pip install -r requirements.txt
python classify.py --text "ياك نتا بخير؟ آش خبارك مع الخدمة؟ نتمنى تكون الأمور كلها مزيانة من جيهتك."

Illustrative output, formatted for readability. Only selected probabilities are shown here; the full output includes all classes.

{
  "label": "Morocco",
  "confidence": 0.9523,
  "probabilities": {
    "Morocco": 0.9523,
    "Algeria": 0.0125,
    "Mauritania": 0.0115,
    "MSA": 0.0062,
    "Libya": 0.0048,
    "..."
  }
}

Keep all three model files in the same directory. The script loads them from its own directory by default. Use --model-dir to select another local bundle.

For JSONL input, each line must contain a string text field:

python classify.py --input input.jsonl > predictions.jsonl

Use --input - to read from standard input. The script returns one JSON object per input, with label, confidence, and probabilities. Empty inputs and text removed entirely by preprocessing raise an error.

To load the model once and classify multiple texts, use the Python API:

from classify import ArabicLID

model = ArabicLID(".").load()
predictions = model.predict([
    "ياك نتا بخير؟ آش خبارك مع الخدمة؟ نتمنى تكون الأمور كلها مزيانة من جيهتك.",
    "هذا نص باللغة العربية الفصحى.",
])

Illustrative batch output, with selected probabilities shown for each input:

[
  {                                                                                                                                                                        
    "label": "Morocco",                                                                                                                                                    
    "confidence": 0.9466,
    "probabilities": {                                                                                                                                                     
      "Morocco": 0.9466,                                    
      "MSA": 0.0185,                                                                                                                                                       
      "Algeria": 0.0125,                                    
      "..."
    }
  },
  {
    "label": "MSA",
    "confidence": 0.7295,
    "probabilities": {
      "MSA": 0.7295,
      "Iraq": 0.1203,
      "Egypt": 0.0266,
      "..."
    }
  }
]

Preprocessing removes web and ASR artifacts, diacritics, decorative elongation, and repeated punctuation. It preserves alef variants.

Limitations

The model always assigns one of its 51 supported classes. It may therefore misclassify unsupported languages or mixed-language text.

License

This model is licensed under the Apache 2.0 License.

You must not use this model in a manner that infringes, misappropriates, or otherwise violates any third party’s rights, including intellectual property rights.

Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mistralai/LIDstral-Arabic