Closer tokenizer–morpheme alignment failed to predict better generation in the reported tasks. GPT-4o reached 97% on nonce roots with 17% boundary precision. The comparison tests this association; it leaves the mechanism behind the performance unresolved.
Arabic uses a root-and-pattern (non-concatenative) system. consonantal roots combine with templatic vowel patterns to form words.
example: root ktb (write) + pattern mafūl → maktūb (written). the root and pattern interleave; they aren't concatenated like prefix+stem+suffix in English, which makes Arabic a stress test for subword tokenizers (BPE, Unigram, WordPiece) built for concatenative morphology.
measured how well each tokenizer's segments match gold-standard morpheme boundaries from CAMEL and Farasa analyzers, on MSA (ATB3) and dialectal (BOLT) Arabic. metrics:
three probing tasks using real roots and nonce (invented) roots:
ALLAM, FANAR, GPT-4, GPT-4o, LLaMA-3, Qwen-3, Cohere. FANAR uses MorphBPE (morphologically-informed tokenization); the rest use standard BPE/Unigram/WordPiece. zero-shot and one-shot prompts, tested in both English and Arabic.
finding 1
no correlation between tokenizer alignment and generation performance. GPT-4o scored highest across all tasks (97% nonce accuracy) with one of the worst alignment scores (17% boundary precision). ALLAM had the best MCR (83-86%) but fell to 20% on nonce words.
finding 2
ALLAM and FANAR drop sharply on the reported nonce-root tasks. The result exposes a generalization gap in these settings; it does not prove that their internal computation is only lexical memorization. FANAR’s morphological tokenizer did not ensure strong nonce performance.
finding 3
Most evaluated models performed better with English instructions than Arabic instructions. Training-language composition is one possible explanation, but this comparison does not isolate it.
finding 4
One-shot prompting improved LLaMA-3, Qwen-3, and Cohere; GPT-4 and GPT-4o stayed flat. The weaker models benefited from an example of the transformation.
finding 5
five error modes. pattern misapplication (root right, template wrong), root deformation (consonants changed), real-word substitution (outputting a valid word instead of applying the pattern), incorrect affix ordering, and partial truncation.
Morphology-aware tokenizers introduce engineering choices that need task-specific evaluation. GPT-4o’s reported fertility > 3 and boundary precision of 17%, alongside 97% nonce accuracy, show that poor alignment did not prevent success here. They do not measure the value of changing only its tokenizer.
The authors propose compositional reasoning and instruction-following as an explanation. The behavioral results motivate that account; they do not isolate an internal mechanism or show that explicit morphological parsing is unnecessary in every setting.
ALLAM and FANAR underperformed GPT-4 and GPT-4o on the reported tasks. Architecture, training data, scale, and tuning differ together, so these results do not isolate which factor explains the gap.
ALLAM scored 67% on real roots and 20% on nonce roots, a 47-point generalization gap that real-word-only evaluation would miss.