← voidwest ember engineering

tool calling without changing the inference core

2026-08-25 · Ember v0.6.7 · structured execution and research tracing

once a model can call tools, generation attracts responsibilities it should not own. tool syntax starts leaking into token loops, retries become decode branches, and external side effects get confused with KV state. soon the forward pass is carrying execution policy that has nothing to do with producing the next token.

Ember v0.6.7 keeps that boundary hard. tool calling was added as an explicit state machine above the existing inference and session stack. attention, KV storage, tokenization, model forward passes, and tensor code did not change. the runtime lives in the additive src/agent/ module tree; a thin CLI exposes it. the model still does exactly one job: generate tokens.

src/agent/session.rs       explicit execution state machine + commit ledger
        │
        ├── protocol.rs    model-family render/parse codecs
        ├── schema.rs      raw JSON → ValidatedArguments
        ├── tool.rs        bounded execution, cancellation, timeout
        ├── trace.rs       authoritative ordered event record
        └── artifact.rs    atomic, hashed run artifacts
        │
        ▼
ChatModelEngine           narrow seam over the existing model session
        │
        ▼
inference core            tokenizer · forward pass · attention · KV · tensors
                          unchanged in v0.6.7

that layout is the design. the agent layer can decide that generated text describes an action, but it cannot redefine inference to make the action happen.

the loop is a state machine

  1. model turn: generate one assistant turn against the currently committed session.
  2. parse action: classify the turn as final text, a tool call, or an explicitly malformed call.
  3. validate: resolve the registered tool and convert raw JSON into typed, validated arguments.
  4. execute: run the tool under cancellation, timeout, panic containment, and hard-call limits.
  5. reinject: render the result through the same family protocol and commit it to the same session.
  6. continue: start the next model turn, or commit final text and stop.

there is no recursive call back into an "agent prompt." one loop owns the state transitions. it checks maximum model steps, tool calls, wall time, per-tool timeout, per-turn output tokens, and reinjected result size at explicit boundaries. cancellation is a first-class outcome, not another string placed in context.

model syntax ends at the protocol boundary

Qwen2.5 and Llama 3.x do not describe tool calls the same way. Qwen uses a ChatML assistant message containing <tool_call>...</tool_call>. Llama uses <|python_tag|> followed by a JSON custom-function call, then expects the result under an ipython role. stop tokens and result rendering differ too.

family |assistant action |result form
Qwen2.5<tool_call>{"name":…,"arguments":…}</tool_call>ChatML user message with <tool_response>
Llama 3.x<|python_tag|>{"name":…,"parameters":…}ipython role message

the generic loop knows neither syntax. it receives one of three protocol-level actions: FinalText, ToolCall, or MalformedToolCall. codecs render the system tool schemas, user messages, assistant framing, and tool results, then parse model output back into that common action type. their byte-exact renders are pinned by tests.

this is where model-family special cases belong. adding another syntax means implementing another codec, not branching inside session logic or teaching the KV cache what a tool response is.

raw model output is never executable input

parsing a JSON-looking span is only syntax recognition. it is not permission to execute. the runtime resolves the named tool, parses its argument object, validates it against the registered schema, and only then constructs ValidatedArguments. the Tool trait accepts that type; it has no entry point that accepts raw model text.

model bytes
  → protocol parser
  → raw tool call
  → schema validation
  → ValidatedArguments
  → Tool::execute(...)

validation is strict and recursive. it checks strings, numbers, integers, booleans, arrays, enums, and nested objects. missing required fields, unknown fields, wrong types, enum violations, malformed JSON, and excessive nesting fail closed. schema validation collects every field error in the call instead of stopping at the first one.

{
  "ok": false,
  "kind": "invalid_arguments",
  "errors": [
    {"path":"operation", "reason":"missing required argument"},
    {"path":"precision", "reason":"unknown field"},
    {"path":"a", "reason":"expected number, got string"}
  ]
}

malformed calls and validation failures are traced as rejections and reinjected through the protocol as ok:false tool results. the model can repair the call on its next turn. recovery remains bounded by the same step and wall-time limits; a bad call does not quietly become prose, and it never reaches the tool.

generation can roll back; an executed side effect cannot

the difficult boundary is not parsing. it is commit semantics. model generation is speculative until the turn completes. an external tool effect becomes real when the tool executes. those are different kinds of state, and the runtime refuses to pretend otherwise.

cancel during generation
  → truncate speculative KV state
  → commit no assistant turn

tool executes
  → external side effect now exists

cancel before tool-result commit
  → keep the external effect visible
  → do not inject the result into the session
  → emit tool_result_uncommitted

the committed ledger records the actual conversation transition: system → user → assistant_tool_call → tool_result → … → final. cancellation during generation rolls back to the previous KV boundary. cancellation after execution leaves a valid committed prefix and no fictional tool-result message, while the trace states that the external effect already happened.

this distinction matters most for writes. deleting a generated suffix cannot unwrite a file, reverse an API mutation, or retract a message. rollback is a property of speculative model state, not a universal undo mechanism.

the trace is the execution record

tracing in v0.6.7 is not diagnostic logging added around the real result. the JSONL trace is the authoritative execution history. every line carries the ember.agent.trace.v1 schema, a run id, and a monotonic sequence number. wall-clock timestamps help with latency; ordering comes from seq.

{"schema":"ember.agent.trace.v1","run_id":"run-…","seq":29,
 "event_type":"tool_execution_finished","step":"tool-3","phase":"execute",
 "data":{"tool":"write_artifact","ok":true,"duration_ms":0.3,
         "artifact_ids":["0000-94474f5d7426"]}}

the provenance event pins the runtime, model and tokenizer identity, protocol and tool-schema snapshot, and run configuration. subsequent events record model turns, parse decisions, validation, tool execution, session mutations, result commits, terminal states, and artifact hashes.

each JSONL event is flushed when appended, so a crash leaves a readable prefix; the parser reports a torn trailing line instead of rejecting the history. privacy controls can replace prompts or generated text with lengths and hashes, summarize or hash tool payloads, and keep per-token events off. auditable does not have to mean copying every payload into the trace.

the model said it wrote the file

the most useful failure in the real-model runs was not a crash. in one attempt, the model narrated that it had written an artifact. it had not called write_artifact. the trace contained no tool-execution event, and the artifact store was empty.

model narration ≠ execution history

without the execution boundary, the final sentence is tempting to treat as success. with it, the claim is easy to falsify. an artifact exists only when the tool ran, the store published the bytes, and the trace recorded the resulting identity. generated text can describe that history; it cannot manufacture it.

one successful run

a Qwen2.5-1.5B Q8_0 run exercised the same boundary successfully. the model read experiment summaries containing a baseline score of 0.612 and a tuned score of 0.744, called the calculator, wrote a Markdown artifact, and returned DONE 21.57%.

source scores     baseline_q4 = 0.612, tuned_q8 = 0.744
calculation       (0.744 - 0.612) / 0.612 × 100 = 21.5686…%
artifact          0000-94474f5d7426-best-config.md
sha256            94474f5d742605e55400ce8d35ff27d6ef204c13b8820c02acb89fdef0e01d6e
final answer      DONE 21.57%

the interesting result is not that the model can divide. it is that the reads, calculation, write, content hash, and final synthesis are distinct, attributable events. the trace shows which facts came from files, which number came from a tool, and which bytes became an artifact.

orchestration is not the bottleneck

a release-mode scripted benchmark measured the runtime around model inference: parsing, validation, mock tool execution, state transitions, and tracing. these are the reported results from 200 repetitions on the report host, not tool-latency or model-throughput numbers.

scenario |wall time / run
one tool, tracing off0.494 ms
one tool, JSONL tracing0.700 ms
three tools, memory trace1.898 ms

agent orchestration is not the bottleneck; model inference still dominates by orders of magnitude. the real Qwen workflow spent about 63.7 seconds in model turns and 2.8 milliseconds in tools. the boundary is explicit without being expensive.

the claim stays narrow

this is not an autonomous-agent framework, browser agent, shell executor, MCP implementation, multi-agent system, or approval framework. v0.6.7 is a small structured execution layer with explicit state and an auditable history. its built-ins are deterministic local tools; there is no shell, network, browser, or delete tool.

the current protocol surface executes one tool call per assistant step; additional calls are counted and traced rather than fanned out. tool timeouts use a watchdog around synchronous work, so a timed-out worker is detached and its result discarded rather than preempted unsafely. those limits are visible parts of the contract, not hidden behind the word "agent."

tool calling did not need to become part of inference. the model still generates tokens. the runtime above it decides whether those tokens describe a final answer or a valid action, and the trace records what actually happened.