← voidwest    research notes

tokenization in Arabic embedding models

On the Importance of Tokenization in Arabic Embedding Models
alkaoud & syed · WANLP 2020
Arabic NLP tokenization morphology embeddings

the result

The study reports gains from morphology-aware tokens with Word2Vec and BERT. It reports a 60% smaller vocabulary and evaluates NER, sentiment, and POS tagging. The gains belong to those experimental settings.

results

context

this worked because Arabic's morphological system has a finite (and relatively small) set of roots, patterns, and affixes.

a morphology-aware tokenizer maps surface forms to these primitives, reducing the effective vocabulary from ~1M surface forms to ~20K morphemes, and since the embedding layer accounts for a large fraction of parameter count in smaller models, shrinking the embedding table while keeping semantic compositionality intact produces gains.

Hypothesis

comparison with the 2026 results

The 2020 embedding study and the 2026 generation study ask different questions. Their results can coexist: an explicit morphology intervention can help particular embedding tasks while tokenizer alignment alone fails to predict generation.

Scale is one hypothesis to test, alongside task, architecture, and training differences; neither comparison identifies a transition point.