POS recovery survived the stricter lexical splits. Under lemma-heldout evaluation, its lift over character n-gram baselines was +19.2pp for Qwen3 and +15.2pp for Llama. This supports linear recovery in the corrected setup; causal use of morphology remains untested.
Ember could load GGUF models, run small CPU-first inference, extract hidden states, and save enough activation data to train probes. That was already useful.
It meant I could stop treating model internals as something locked behind a Python stack and start building the measurement path myself in Rust.
This week, the center of the work moved. The hard question is no longer only whether Ember can extract hidden states. It is whether I can trust what the probe is measuring.
The revised setup supports fewer claims and makes their limits explicit.
The draft’s new title, Leakage-Aware Probing of Arabic Morphology in Small Language Models, reflects the work: split policy, token position, prompt content, baselines, and runtime validity determine what the probe scores mean.
The paper is still a preprint draft, but the core claim has finally become narrow enough that I can defend it.
A probe result is only useful if the split tests the kind of generalization it claims to test.
The original Ember probing loop was straightforward:
The pipeline remains the same. Hidden-state extraction is the measurement instrument, so its validity shapes every interpretation.
The current study uses 4,701 Arabic morphology stimuli derived from PADT / CAMeL-style annotations. The labels include root, lemma, part of speech, abstract pattern, concrete pattern, gender, and number.
These are not all equally easy to evaluate. Some are small closed sets.
Some are high-cardinality lexical or quasi-lexical labels. Some are directly visible in the prompt unless the prompt is ablated.
This week also narrowed the model set to two valid runs:
I removed the Qwen2.5 results because my runtime and tokenizer-loading path produced invalid hidden states. Those activations tell us nothing reliable about Qwen2.5’s Arabic morphology.
That is the theme of the week. The table got less exciting. The methodology got stronger.
The first version used random cross-validation in places where random splits were too forgiving.
For many classification tasks, random CV is fine. For Arabic morphology, it can be misleading.
Related surface forms can share a lemma or root. If related lexical items appear in both train and test, a probe can look strong because it has learned lexical families rather than a more general morphological representation.
Arabic forms built around a shared root can appear in both training and test data. Success on a related form may reflect lexical overlap, leaving transfer to unseen groups unresolved.
So the stricter evaluation uses grouped splits:
Lower scores under grouped splits expose how much the random split benefited from lexical overlap.
A smaller claim can be a stronger result.
The strongest result right now is part of speech.
POS survives lemma/root-heldout evaluation in both Qwen3-0.6B and Llama-3.2-1B. It also shows positive lift over character n-gram surface baselines: +19.2pp for Qwen3 under lemma-heldout evaluation, and +15.2pp for Llama under the same split.
That does not prove the models have a complete theory of Arabic morphology. It does show that, under stricter splits, the hidden states contain linearly recoverable syntactic/morphological category information beyond what the simple surface baseline captures.
Gender and number show more modest positive lift. Their lower-cardinality labels suit heldout evaluation, though prompt and pattern effects still complicate interpretation.
The important part is that some signal remains after the easy leakage path is made harder.
The stricter split exposed a second issue: root, lemma, and pattern are not like POS, gender, and number.
POS, gender, and number are low-cardinality labels. Their classes are mostly present in both train and test. A closed-set classifier can reasonably be asked to predict them under heldout lexical groups.
Root, lemma, abstract pattern, and concrete pattern have many more classes. Under lemma-heldout or root-heldout splits, the test set can contain labels that never appeared during training. A standard closed-set classifier cannot predict a class it has never seen.
A strict heldout split can ask a closed-set classifier to predict unseen labels. Failure under that setup leaves the presence of root information unresolved.
So root, lemma, and pattern need a different evaluation framework. Possibilities include retrieval-style evaluation, representation geometry, nearest-neighbor structure, contrastive tests, or controlled nonce stimuli where the label space is designed around generalization.
The current paper now treats high-cardinality labels more carefully instead of pretending the same classifier setup works for every feature.
I removed labels that this method cannot validly evaluate under the revised splits.
That is a better paper.
Another correction is token position.
The study extracts hidden states at the prompt’s final period token. Earlier descriptions left that position implicit.
The current framing is prompt-final representation probing.
This position is intentional. The final period token is tokenizer-stable across models and sits after the full prompt.
It can aggregate information from the surface word and the morphological fields included in the prompt. That makes cross-model comparison cleaner than choosing a model-specific Arabic subword position.
But it also changes the interpretation. This is not direct word-token probing.
I am not claiming that the Arabic word's own final subword contains the measured information. I am probing the representation at a stable prompt-final position after the model has read the whole formatted stimulus.
That is a narrower claim, and it needs to be said plainly.
The original prompt included:
Lemma, root, and pattern fields supply analysis to the model. A prompt-final probe can recover that supplied analysis, so the result needs a separate surface-only control.
So I added an ablated prompt that removes Lemma, Root, and Pattern.
The ablation is informative:
This is not a clean "one model is better" result. It is a prompt-dependence result. The same probing setup can mean different things across architectures, even for the same label.
The lesson is simple: task-informative prompts need ablation checks. If the prompt contains the answer, or contains features close to the answer, the probe may still be measuring something real, but it is measuring the representation of the whole prompt context.
This week’s corrections changed which results the paper could support.
I removed the invalid Qwen2.5 results. I clarified that the probed position is the final period token.
I fixed layer indexing: layer 0 is the first transformer block output, not the embedding layer. I pinned the character n-gram baseline so comparisons are traceable.
I fixed number drift across runs and made the table values easier to audit. I also fixed Table 5 formatting so it no longer hides important distinctions.
None of this is glamorous, but it is the work that makes the remaining claims usable.
The systems code is part of the measurement instrument.
That sentence has become more true as the project has matured. If the tokenizer path is wrong, the hidden states are wrong.
If the layer index is mislabeled, the interpretation is wrong. If the baseline drifts, the lift is not traceable.
If the split leaks lexical families, the result can look stronger than it is.
The Rust code, dataset builder, probe scripts, metadata, and paper tables are not separate pieces anymore. They are one measurement pipeline.
Ember started as a CPU-first Rust inference engine for GGUF models. That is still the base. But this week made it clearer that Ember is becoming something more specific: a reproducible probing pipeline.
The engineering matters because hidden-state extraction is the measurement instrument. The probe is only as meaningful as the activations, metadata, tokenizer handling, prompt construction, split policy, and baseline around it.
That changes how I think about the project. Speed and model support still matter, but correctness and traceability matter more.
A probing engine should make it hard to confuse invalid activations with results. It should record enough metadata that a table can be traced back to the exact model, prompt, token position, split policy, and baseline.
That is less exciting than a big accuracy number, but it is the foundation for results I can stand behind.
The immediate next steps are straightforward: add more valid models, strengthen confidence intervals, and design better evaluation for root, pattern, and lemma. Direct-token probing should be added as complementary evidence, especially to separate word-local representations from prompt-final aggregated representations.
Representation-geometry methods also look more appropriate for high-cardinality morphology than plain closed-set classifiers under heldout lexical splits.
The claim is narrower now. POS survives stricter evaluation in two small models, with positive lift over surface baselines.
Gender and number show more modest signal. Root, lemma, and pattern remain central to the larger research direction, but need a better evaluation setup before I treat them as established.
The remaining claims depend on the corrected runs and their documented controls.