← voidwest    engineering    internals
Inspectable CPU inference for reproducible LLM systems research.
rust 1.92 ● mit ● capture · intervene · verify

Ember runs GGUF on CPU in a way you can inspect and replay. Capture states, patch them, and pack the run into a bundle you can verify later. It is not a llama.cpp replacement: llama.cpp remains the external performance and correctness reference.

Inspectable forward pass → artifacts with checksums → replayable probes and interventions.

Golden path: Llama-3.2-1B-Instruct Q8_0. Start with the five-minute CLI workflow: the pinned model and tokenizer produce a bundle with captures across all 16 layers, ready for verification and intervention comparison.

A probe can read a feature that the model's answer path does not use; more fitting can even make transfer worse. That distinction is why this instrument exists. Read the probe can read it. can the model use it? for the diagnostic evidence and its limits.

The 1.0 priority is capture, intervention, bundles, and verification on a short, explicitly validated model list. Consoles, agent traces, and EmberSEC are related tools and research tracks; follow their separate documentation.

start here

validation ladder

smokestructural execution only: the command loaded artifacts and produced output.
golden logitsoutput-logit comparison against a trusted reference for the same prompt, tokenizer, model, and quantization path.
activation checkshidden-state comparison by prompt, tokenizer, model, layer, and token position.
probeslinear or MLP decodability/recoverability, not causal model use.
interventionsonly supports behavioral claims when downstream logits or continuations change.

current status

areastatusread
CPU runtimeworks locally across small/medium GGUF pathsengineering artifact, not production parity
Qwen3 0.6Bgeneration/probe paths runneeds trusted golden-logit reference
LLaMA 1B/3B/8Blocal smoke/probe artifacts existresearch conclusions remain preliminary
Gemma 4 E2Bdense text-only path runs local smoke/benchmarkexperimental until golden checks cover architecture details
encoder benchmarksmBERT PADT smoke completed; suite manifest existsfull XLM-R/AraBERTv2 suite still pending

latest update

the newest engineering pass added a thread-count benchmark section under engineering. local results show that larger dense Q8_0 models benefit from threaded runtime paths on this machine, while the small Qwen3 0.6B run does not. the page keeps that claim deliberately local; it is not a cloud-speed forecast.

deeper pages

a developing thread on inference-engine attack surfaces, parser trust boundaries, memory safety, and reproducible security work.
lifecycle ordering, context fields, tensor ownership, family-specific semantics, and failure propagation.
architecture, design decisions, math primitives, attention, KV cache, bugs, and the first coherent output.
the current systems work, benchmark plot, and engineering subpages.
Arabic NLP notes, morphology probing, tokenizer papers, and running research direction.