← voidwest    research notes

the tokenizer isn't the problem

what i learned reading arabic nlp papers for a week
mohammed al-thobaiti · 2026-05-16

Tokenizer alignment failed to predict generation performance in the reported Arabic tasks. GPT-4o reached 97% on nonce roots with 17% boundary precision. The result leaves the underlying computation unresolved.

why Arabic is a real stress test for tokenizers

Arabic uses a root-and-pattern (non-concatenative) morphological system. a 3-consonant root like k-t-b (write) combines with templatic vowel patterns to produce: kataba (he wrote), kitaab (book), maktab (office), maktūb (written), yaktubu (he writes).

the root never appears as a standalone word, it's always embedded inside a surface form.

a single Arabic word can encode what would be 4 separate tokens in English. fasayaktubūnahā (فسيكتبونها) = "and they will write it", one word with proclitics, stem, and enclitic fused together.

BPE merges contiguous character sequences by frequency; Arabic root-pattern structure can be non-contiguous. Token boundaries therefore need not follow morphological structure. Optional diacritics, spacing variation, and dialect differences add further complications.

the obvious hypothesis: fix the tokenizer

the field tried this. a lot:

reasonable assumption: if you segment morphemes correctly before the model sees them, it should learn morphological structure better.

then a 2026 paper broke the assumption

Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
alakeel, qwaider, aldarmaki, alqahtani · LREC 2026, arXiv:2603.15773
from SDAIA, MBZUAI, and PNU

they evaluated 7 Arabic LLMs (ALLAM, FANAR, GPT-4, GPT-4o, LLaMA-3, Qwen-3, Cohere) on two dimensions: (1) how well their tokenizers align with gold-standard morpheme boundaries, and (2) how well the models can productively generate Arabic root-pattern forms, including nonce (invented) roots, which tests real generalization, not memorization.

Key Finding

alignment doesn't predict competence

ALLAM had the best reported morphological alignment (MCR 83-86%) but reached only 20% accuracy on nonce words. Better alignment did not ensure generalization on this task; that alone does not prove a memorization mechanism.

GPT-4 had the worst tokenizer alignment (fertility 4× higher than ideal, boundary precision 17%) but scored 92% on nonce root-pattern generation. second best overall.

GPT-4o scored 97% on nonce words despite similarly bad tokenizer alignment.

no correlation between tokenizer alignment metrics and morphological generation performance. morpheme F1 and MCR have zero or weak negative correlation with generation accuracy.

FANAR’s MorphBPE tokenizer accompanied steady, middling performance. The contribution of instruction-following remains unresolved.

English prompts performed better for most models in the reported comparison. Instruction-tuning language is a plausible contributor, but the experiment does not isolate training-language composition.

The paper recommends measuring productive generalization directly. Compositional reasoning and instruction-following remain proposed explanations for the scores.

what this means

The two studies differ in tasks, models, and interventions. Testing scale as an explanation would require a controlled comparison.

the field has been asking "which tokenizer wins on classification tasks" for years. multiple papers (2023, 2024) kept landing on "it depends on the task and dataset." the 2026 paper reframes the question entirely: stop asking about surface segmentation and start asking about productive generalization.

the real gap might not be at the tokenizer or even the embedding layer, it might be in the internal representations the model builds. GPT-4 and GPT-4o are doing something ALLAM isn't, despite ALLAM having more Arabic training data and a better tokenizer.

what is it?

what's next

i'm going to look at the internal activation angle, what do the representations actually look like inside models that succeed vs fail on nonce Arabic words?

Small open-weight models make a local activation study feasible. Local logit measurement still requires a working model or equivalent inference endpoint; choosing logits as the readout does not remove the cost of inference. Ember provides access to layer states for the supported local models.

other threads i'm tracking:

this started as a curiosity from a linkedin post. now i have actual research questions and a direction that doesn't seem fully explored. more updates as i go.


papers referenced

  1. alakeel, qwaider, aldarmaki, alqahtani, "morphemes without borders", LREC 2026, arXiv:2603.15773
  2. alkaoud & syed, "on the importance of tokenization in Arabic embedding models", WANLP 2020, ACL Anthology
  3. attia, "Arabic tokenization system", ~2007 (rule-based, finite-state)
  4. alrefaie et al., "exploring tokenization strategies and vocabulary sizes for enhanced Arabic language models", arXiv:2403.11130, 2024