← voidwest    research notes

morphemes without borders

Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
Alakeel, Qwaider, Aldarmaki, Alqahtani · LREC 2026
Arabic NLP morphology tokenization LLM evaluation arXiv:2603.15773

the core claim

Closer tokenizer–morpheme alignment failed to predict better generation in the reported tasks. GPT-4o reached 97% on nonce roots with 17% boundary precision. The comparison tests this association; it leaves the mechanism behind the performance unresolved.

why Arabic morphology is a good test

Arabic uses a root-and-pattern (non-concatenative) system. consonantal roots combine with templatic vowel patterns to form words.

example: root ktb (write) + pattern mafūl → maktūb (written). the root and pattern interleave; they aren't concatenated like prefix+stem+suffix in English, which makes Arabic a stress test for subword tokenizers (BPE, Unigram, WordPiece) built for concatenative morphology.

experimental design

part 1: tokenizer morphological alignment

measured how well each tokenizer's segments match gold-standard morpheme boundaries from CAMEL and Farasa analyzers, on MSA (ATB3) and dialectal (BOLT) Arabic. metrics:

part 2: morphological generation

three probing tasks using real roots and nonce (invented) roots:

models evaluated

ALLAM, FANAR, GPT-4, GPT-4o, LLaMA-3, Qwen-3, Cohere. FANAR uses MorphBPE (morphologically-informed tokenization); the rest use standard BPE/Unigram/WordPiece. zero-shot and one-shot prompts, tested in both English and Arabic.


key findings

Key Finding

finding 1

no correlation between tokenizer alignment and generation performance. GPT-4o scored highest across all tasks (97% nonce accuracy) with one of the worst alignment scores (17% boundary precision). ALLAM had the best MCR (83-86%) but fell to 20% on nonce words.

Key Finding

finding 2

ALLAM and FANAR drop sharply on the reported nonce-root tasks. The result exposes a generalization gap in these settings; it does not prove that their internal computation is only lexical memorization. FANAR’s morphological tokenizer did not ensure strong nonce performance.

Key Finding

finding 3

Most evaluated models performed better with English instructions than Arabic instructions. Training-language composition is one possible explanation, but this comparison does not isolate it.

Key Finding

finding 4

One-shot prompting improved LLaMA-3, Qwen-3, and Cohere; GPT-4 and GPT-4o stayed flat. The weaker models benefited from an example of the transformation.

Observation

finding 5

five error modes. pattern misapplication (root right, template wrong), root deformation (consonants changed), real-word substitution (outputting a valid word instead of applying the pattern), incorrect affix ordering, and partial truncation.


what this means for Arabic NLP research

1. evaluate tokenizer changes on the target task

Morphology-aware tokenizers introduce engineering choices that need task-specific evaluation. GPT-4o’s reported fertility > 3 and boundary precision of 17%, alongside 97% nonce accuracy, show that poor alignment did not prevent success here. They do not measure the value of changing only its tokenizer.

2. instruction-following is an explanation to test

The authors propose compositional reasoning and instruction-following as an explanation. The behavioral results motivate that account; they do not isolate an internal mechanism or show that explicit morphological parsing is unnecessary in every setting.

3. the model comparison does not isolate training choices

ALLAM and FANAR underperformed GPT-4 and GPT-4o on the reported tasks. Architecture, training data, scale, and tuning differ together, so these results do not isolate which factor explains the gap.

4. nonce conditions expose a generalization gap

ALLAM scored 67% on real roots and 20% on nonce roots, a 47-point generalization gap that real-word-only evaluation would miss.

5. questions for a controlled follow-up


limitations