← Back to all fieldnotes
Agent safety

Jev: Typed Decision Gates for Agent Safety

A fast reflex

The Jev source example uses a low-cost decision model, with a reported latency around 420 ms, that returns typed probabilities rather than prose. It classifies command action and blast radius before an agent tool call; a reasoning model handles work that passes the gate. Latency is an example measurement, not a general guarantee.

Three command outcomes

The provisional bands are safe ≥0.85: RUN silently; 0.5–0.85: CONFIRM with a human; and safe <0.5: BLOCK. The related prompt gate either runs a high-quality prompt or proposes a context-grounded rewrite and waits for approval. Decisions and prompt versions are recorded in JSONL.

The miss that matters

A destructive command scored 0.54 in shadow data, landing in CONFIRM rather than BLOCK. This is a reviewable miss, not evidence the gate is calibrated. Logs, human override, and a review of roughly 20 prompts are needed before treating the thresholds as reliable.

How the pieces connect

Illustrative architecture and workflow.

Open diagram full size ↗