Ember runs GGUF on CPU in a way you can inspect and replay. Capture states, patch them, and pack the run into a bundle you can verify later. It is not a llama.cpp replacement: llama.cpp remains the external performance and correctness reference.
Inspectable forward pass → artifacts with checksums → replayable probes and interventions.
Golden path: Llama-3.2-1B-Instruct Q8_0. Start with the five-minute CLI workflow: the pinned model and tokenizer produce a bundle with captures across all 16 layers, ready for verification and intervention comparison.
A probe can read a feature that the model's answer path does not use; more fitting can even make transfer worse. That distinction is why this instrument exists. Read the probe can read it. can the model use it? for the diagnostic evidence and its limits.
The 1.0 priority is capture, intervention, bundles, and verification on a short, explicitly validated model list. Consoles, agent traces, and EmberSEC are related tools and research tracks; follow their separate documentation.
| smoke | structural execution only: the command loaded artifacts and produced output. |
| golden logits | output-logit comparison against a trusted reference for the same prompt, tokenizer, model, and quantization path. |
| activation checks | hidden-state comparison by prompt, tokenizer, model, layer, and token position. |
| probes | linear or MLP decodability/recoverability, not causal model use. |
| interventions | only supports behavioral claims when downstream logits or continuations change. |
| area | status | read |
|---|---|---|
| CPU runtime | works locally across small/medium GGUF paths | engineering artifact, not production parity |
| Qwen3 0.6B | generation/probe paths run | needs trusted golden-logit reference |
| LLaMA 1B/3B/8B | local smoke/probe artifacts exist | research conclusions remain preliminary |
| Gemma 4 E2B | dense text-only path runs local smoke/benchmark | experimental until golden checks cover architecture details |
| encoder benchmarks | mBERT PADT smoke completed; suite manifest exists | full XLM-R/AraBERTv2 suite still pending |
the newest engineering pass added a thread-count benchmark section under engineering. local results show that larger dense Q8_0 models benefit from threaded runtime paths on this machine, while the small Qwen3 0.6B run does not. the page keeps that claim deliberately local; it is not a cloud-speed forecast.