← voidwest    research notes
historical context

The setup, estimates, and next steps below belong to the date of this note. The recorded results are retained; this editorial update is not a rerun or a claim that the historical implementation is still current.

probing Arabic morphology inside LLMs

a research plan · 2026-05-16
Arabic NLP probing ember morphology research plan
method clarification · 2026-08 audit

This note proposes testing whether probes can recover Arabic roots and patterns from model states. The planned stimulus grid combines about 200 nonce roots with 10 patterns. It is an experiment plan; no probe result is reported here.

Open Question

The morphemes without borders paper reports 97% for GPT-4o on nonce Arabic root-pattern generation and 20% for ALLAM despite better tokenizer alignment in ALLAM. This motivates a representation question; it does not establish where morphology is computed, or that one model lacks morphological representations.

motivation

the 2026 paper leaves an open question: morphological competence in LLMs is defined by productive generalization, not tokenizer alignment, but we don't know how the model achieves it. the authors suggest "compositional reasoning + instruction-following" as the mechanism, but that's a behavioral description, not a mechanistic one.

three facts make this tractable right now:

  1. ember gives direct access to hidden states. after every transformer block, i can call backend.data(&x) and read the activations. no hooks, no CUDA synchronization, no framework overhead.
  2. behavioral probing is cheap. i don't need to train probes on 100K examples. the nonce root-pattern task from the 2026 paper gives a clean ground-truth signal: feed a root + pattern, check if the output is correct. probe at every layer.
  3. the comparison writes itself. run the same probe on GPT-2 (terrible tokenizer alignment, unknown Arabic performance), LLaMA 3 (tested in the 2026 paper, middle of the pack), and a dedicated Arabic model (ALLAM if accessible, otherwise AraBERT). layer-by-layer, same stimuli.

GPT-2, LLaMA, and ALLAM differ in architecture, training, and size. Within-family comparisons reduce some differences, but do not automatically isolate scale. Every comparison needs its exact checkpoints, prompts, and measurement interface recorded.

experimental design

stimuli: root-pattern nonce pairs

build a stimulus set of ~200 nonce triliteral roots (consonant triplets that don't exist in Arabic, filtered against a lexicon) crossed with ~10 common patterns (fa3ala, maf3ūl, yaf3alu, fā3il, etc.). each stimulus is: "apply pattern X to root Y" → expected surface form. example:

the gold-standard dataset from Alakeel et al. is public (github). start there, extend with more patterns if needed.

probing setup

for each stimulus:

  1. run forward pass through the model with caching disabled (we want all hidden states, not just the final logit)
  2. extract hidden states after each transformer block, specifically after the attention residual add and after the MLP residual add (two snapshots per layer)
  3. extract the output embedding (logits) for the final token position
  4. record correctness: does the argmax token match the expected surface form?

analysis questions

Open Question

Q1: where does root identity live?

train linear probes on the hidden states at each layer to classify the root (which of the 200 nonce roots produced this activation?). a high-accuracy probe at a given layer means root identity is linearly decodable there. compare across models.

Open Question

Q2: where does pattern identity live?

same probe, different target: classify which pattern was applied. does pattern information appear at the same layers as root information? does it appear earlier or later?

Open Question

Q3: are root and pattern disentangled?

For an open-weight model with accessible activations, compare root-probe and pattern-probe weight subspaces at each layer. Low overlap describes these fitted readouts under the chosen metric; it does not establish causal or symbolic disentanglement inside the model.

Open Question

Q4: where does the model "figure it out"?

Compare representations associated with correct and incorrect outputs within the same model. A layerwise difference can motivate a follow-up intervention, but it cannot by itself identify where the computation succeeds or fails.

Hypothesis

Q5: the scaling question

if possible, run the same probes on GPT-2 small → medium → large (or LLaMA 3 1B → 3B → 8B). at what parameter count does the model transition from memorization (ALLAM-like, real roots only) to productive generalization (GPT-4o-like, nonce roots)?

this answers the inflection-point question from the tokenizer writeup.

implementation in ember

what needs to change

ember already has everything needed for the base inference. the probing pipeline needs two additions:

  1. activation capture mode, a flag or separate function that runs the forward pass but saves the hidden state after each block instead of discarding it. currently Gpt2::forward_with_cache overwrites x at each layer. a variant that pushes x.clone() to a Vec<B::Tensor> before the next block is ~5 lines.
  2. probe training harness, a small Rust module or Python script that takes the saved activations, fits logistic regression probes (via linfa or numpy), and reports accuracy per layer. this doesn't need to be fast, we're running ~200 stimuli, not 200K.

models to test

modelparamswhy
GPT-2 small124Mbaseline; english-centric, bad Arabic tokenizer. tested in 2026 paper indirectly
LLaMA 3 1B / 3B1B / 3Bmiddle of the pack in 2026 paper. likely has some Arabic in pretraining
AraBERT / CAMeLBERT~110MArabic-specific BERT. good tokenizer alignment. should behave like ALLAM, good on real, bad on nonce?
LLaMA 3 8B8Blargest feasible on consumer CPU. where does the inflection happen?

compute requirements

These were planning estimates, not measured runtime results: ~50ms per GPT-2 forward pass, 200 stimuli × 12 layers × 2 snapshots = 4,800 vectors, and roughly 5 GB for an 8B Q4_K model. Actual latency and memory depend on the loader, context, and captured states; they need separate measurement.

expected results & interpretation

if GPT-4 succeeds but we can't probe it

The plan requires open-weight models with accessible hidden states. GPT-4o’s reported 97% provides a behavioral reference; its internal organization remains inaccessible to this experiment.

if probes are near chance everywhere

Chance-level results would show that this probe, extraction point, and dataset did not recover the target. Nonlinear encoding, insufficient data, and a weak measurement setup would remain possible explanations.

if root and pattern are in the same subspace

Overlapping probe subspaces would describe the fitted readouts. They would not distinguish a lookup table from a compositional rule without additional behavioral and intervention controls.

if the inflection point is between 1B and 3B

A difference between 1B and 3B would motivate a controlled scaling study. It would not establish a parameter threshold below which morphological injection helps, or above which models learn the structure unaided.

related work to cite

next steps

  1. add activation capture to ember, ~30 lines in model.rs. a new method forward_with_activations that returns hidden states alongside logits.
  2. build stimulus set, start with the public dataset from Alakeel et al., filter to nonce roots only, expand to ~200 stimuli with 10+ patterns.
  3. run GPT-2 probes, baseline. GPT-2's Arabic is probably poor, but the probe structure works regardless.
  4. run LLaMA 3 1B/3B probes, the interesting comparison. where's the inflection?
  5. write up, if the scaling curve is clear, this is a short paper (findings or workshop). 6 pages, clean story: "here's what the 2026 paper showed at the behavioral level; here's what's happening inside the model to explain it."