The study reports gains from morphology-aware tokens with Word2Vec and BERT. It reports a 60% smaller vocabulary and evaluates NER, sentiment, and POS tagging. The gains belong to those experimental settings.
this worked because Arabic's morphological system has a finite (and relatively small) set of roots, patterns, and affixes.
a morphology-aware tokenizer maps surface forms to these primitives, reducing the effective vocabulary from ~1M surface forms to ~20K morphemes, and since the embedding layer accounts for a large fraction of parameter count in smaller models, shrinking the embedding table while keeping semantic compositionality intact produces gains.
comparison with the 2026 results
The 2020 embedding study and the 2026 generation study ask different questions. Their results can coexist: an explicit morphology intervention can help particular embedding tasks while tokenizer alignment alone fails to predict generation.
Scale is one hypothesis to test, alongside task, architecture, and training differences; neither comparison identifies a transition point.