The setup, estimates, and next steps below belong to the date of this note. The recorded results are retained; this editorial update is not a rerun or a claim that the historical implementation is still current.
This note proposes testing whether probes can recover Arabic roots and patterns from model states. The planned stimulus grid combines about 200 nonce roots with 10 patterns. It is an experiment plan; no probe result is reported here.
The morphemes without borders paper reports 97% for GPT-4o on nonce Arabic root-pattern generation and 20% for ALLAM despite better tokenizer alignment in ALLAM. This motivates a representation question; it does not establish where morphology is computed, or that one model lacks morphological representations.
the 2026 paper leaves an open question: morphological competence in LLMs is defined by productive generalization, not tokenizer alignment, but we don't know how the model achieves it. the authors suggest "compositional reasoning + instruction-following" as the mechanism, but that's a behavioral description, not a mechanistic one.
three facts make this tractable right now:
GPT-2, LLaMA, and ALLAM differ in architecture, training, and size. Within-family comparisons reduce some differences, but do not automatically isolate scale. Every comparison needs its exact checkpoints, prompts, and measurement interface recorded.
build a stimulus set of ~200 nonce triliteral roots (consonant triplets that don't exist in Arabic, filtered against a lexicon) crossed with ~10 common patterns (fa3ala, maf3ūl, yaf3alu, fā3il, etc.). each stimulus is: "apply pattern X to root Y" → expected surface form. example:
the gold-standard dataset from Alakeel et al. is public (github). start there, extend with more patterns if needed.
for each stimulus:
Q1: where does root identity live?
train linear probes on the hidden states at each layer to classify the root (which of the 200 nonce roots produced this activation?). a high-accuracy probe at a given layer means root identity is linearly decodable there. compare across models.
Q2: where does pattern identity live?
same probe, different target: classify which pattern was applied. does pattern information appear at the same layers as root information? does it appear earlier or later?
Q3: are root and pattern disentangled?
For an open-weight model with accessible activations, compare root-probe and pattern-probe weight subspaces at each layer. Low overlap describes these fitted readouts under the chosen metric; it does not establish causal or symbolic disentanglement inside the model.
Q4: where does the model "figure it out"?
Compare representations associated with correct and incorrect outputs within the same model. A layerwise difference can motivate a follow-up intervention, but it cannot by itself identify where the computation succeeds or fails.
Q5: the scaling question
if possible, run the same probes on GPT-2 small → medium → large (or LLaMA 3 1B → 3B → 8B). at what parameter count does the model transition from memorization (ALLAM-like, real roots only) to productive generalization (GPT-4o-like, nonce roots)?
this answers the inflection-point question from the tokenizer writeup.
ember already has everything needed for the base inference. the probing pipeline needs two additions:
| model | params | why |
|---|---|---|
| GPT-2 small | 124M | baseline; english-centric, bad Arabic tokenizer. tested in 2026 paper indirectly |
| LLaMA 3 1B / 3B | 1B / 3B | middle of the pack in 2026 paper. likely has some Arabic in pretraining |
| AraBERT / CAMeLBERT | ~110M | Arabic-specific BERT. good tokenizer alignment. should behave like ALLAM, good on real, bad on nonce? |
| LLaMA 3 8B | 8B | largest feasible on consumer CPU. where does the inflection happen? |
These were planning estimates, not measured runtime results: ~50ms per GPT-2 forward pass, 200 stimuli × 12 layers × 2 snapshots = 4,800 vectors, and roughly 5 GB for an 8B Q4_K model. Actual latency and memory depend on the loader, context, and captured states; they need separate measurement.
The plan requires open-weight models with accessible hidden states. GPT-4o’s reported 97% provides a behavioral reference; its internal organization remains inaccessible to this experiment.
Chance-level results would show that this probe, extraction point, and dataset did not recover the target. Nonlinear encoding, insufficient data, and a weak measurement setup would remain possible explanations.
Overlapping probe subspaces would describe the fitted readouts. They would not distinguish a lookup table from a compositional rule without additional behavioral and intervention controls.
A difference between 1B and 3B would motivate a controlled scaling study. It would not establish a parameter threshold below which morphological injection helps, or above which models learn the structure unaided.