Blog

Log every Jev decision: the minimum record

Our system-of-record checklist had one step that does most of the work: send every Jev call through one wrapper, and have that wrapper write a record. This post is that wrapper. It's about 110 lines of standard-library Python, it writes one record per decision call, and by default the record contains nothing sensitive.

The goal is modest. When someone asks about a specific decision ("why was this ticket escalated?", "which version answered last Tuesday's batch?", "which project's usage doubled?"), you should be able to answer from a log in a few minutes, without keeping copies of your customers' data.

What a Jev response gives you to log

Everything in this section comes from TypeSafe's API reference and models page, checked on Oct 9, 2026.

A call is a POST to https://api.typesafe.ai/v1/systemone with a bearer key. The body has three parts: model, state (the content to evaluate), and questions, a map of typed questions whose keys you choose.

The response gives you four things worth recording:

  • model: the versioned ID that actually answered. The models page says the response reports it "so you can log which model produced each result." This matters because the jev-latest alias moves when a new release ships, and "the answers behind it can change without a change on your side."
  • answers: one per question, returned under your own keys. A Choice answer has the chosen option, probabilities and a confidence. A Score answer has a probability-weighted score, a legend, probabilities and a confidence. A Noul answer is a single value from 0 (no) to 1 (yes). The docs give a confidence only for Choice and Score.
  • usage.input_tokens: TypeSafe charges for input tokens only, and output tokens are free.
  • The HTTP status. The documented errors are 401 (bad or missing key), 422 (validation failed), 429 (rate limit) and 529 (overloaded). For 429 and 529, the docs recommend retrying with exponential backoff, and say the official SDKs do this by default.

What the response doesn't give you is everything on your side: who called, which version of your questions they sent, and which thresholds your code applied to the answers. The wrapper has to supply those.

The minimum record

Field Comes from Why it matters
ts Your code When the decision happened
request_id Your code A handle to join with your app logs. We didn't find a request ID documented in TypeSafe's API reference, so the wrapper generates its own.
caller.project, caller.service Your code A shared key can't tell you who called, so attribution has to come from you
model_requested Your code What you pinned or asked for
model_answered Response model The version that actually answered
questions_hash Your code Ties the call to the exact instructions and criteria you sent
answers Response, summarized Each answer's type, value, confidence (where given) and the gate result
thresholds, thresholds_version Your code The rule that turned an answer into an action
input_tokens Response usage Cost attribution by project
status, latency_ms HTTP layer Errors, 429 trends and slow calls

The wrapper

The example is a support-triage call: a Choice that picks a department, and a Noul that asks whether the message is urgent. The questions come straight from the examples in TypeSafe's API reference.

"""Minimal Jev wrapper that writes one metadata-only record per decision call.

Standard library only. Record field names are ours, not TypeSafe's.
"""
import hashlib
import json
import os
import sys
import time
import urllib.error
import urllib.request
import uuid

JEV_URL = os.environ.get("JEV_URL", "https://api.typesafe.ai/v1/systemone")
PINNED_MODEL = "jev-1.13.0"  # pin the versioned ID you tuned thresholds against


def questions_hash(questions):
    """Stable hash of the question definitions (instructions and criteria)."""
    canonical = json.dumps(questions, sort_keys=True, separators=(",", ":"))
    return "sha256:" + hashlib.sha256(canonical.encode()).hexdigest()


def summarize(answer):
    """Keep the decision, not the input: the value, plus confidence where Jev gives one."""
    kind = answer.get("type")
    if kind == "choice":
        return {"type": kind, "value": answer.get("choice"), "confidence": answer.get("confidence")}
    if kind == "score":
        return {"type": kind, "value": answer.get("score"), "confidence": answer.get("confidence")}
    if kind == "noul":  # a probability of "yes"; no separate confidence
        return {"type": kind, "value": answer.get("noul")}
    return {"type": kind}


def gate(summary, threshold):
    """Return accept or escalate. Your code owns the thresholds, so log them."""
    if threshold is None:
        return None
    measure = summary.get("confidence", summary.get("value"))
    return "accept" if measure is not None and measure >= threshold else "escalate"


def stderr_sink(record):
    """Replace with your log pipeline or an insert-only table."""
    print(json.dumps(record), file=sys.stderr)


def ask_jev(state, questions, *, project, service, thresholds=None,
            thresholds_version=None, model=PINNED_MODEL, sink=stderr_sink,
            timeout=10):
    thresholds = thresholds or {}
    record = {
        "ts": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "request_id": str(uuid.uuid4()),  # generated by us, not by TypeSafe
        "caller": {"project": project, "service": service},
        "model_requested": model,
        "questions_hash": questions_hash(questions),
        "thresholds": thresholds,
        "thresholds_version": thresholds_version,
    }
    body = json.dumps({"model": model, "state": state, "questions": questions}).encode()
    req = urllib.request.Request(JEV_URL, data=body, method="POST", headers={
        "Authorization": "Bearer " + os.environ.get("TYPESAFE_API_KEY", ""),
        "Content-Type": "application/json",
    })
    start = time.monotonic()
    try:
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            data = json.load(resp)
            record["status"] = resp.status
    except urllib.error.HTTPError as err:  # 401, 422, 429, 529 ...
        record["status"] = err.code
        record["latency_ms"] = round((time.monotonic() - start) * 1000)
        sink(record)  # failed calls get a record too
        raise
    record["latency_ms"] = round((time.monotonic() - start) * 1000)
    record["model_answered"] = data.get("model")  # the versioned ID that answered
    record["input_tokens"] = data.get("usage", {}).get("input_tokens")
    answers = {}
    for name, answer in data.get("answers", {}).items():
        summary = summarize(answer)
        summary["gate"] = gate(summary, thresholds.get(name))
        answers[name] = summary
    record["answers"] = answers
    sink(record)  # no state, no instruction text, no key
    return data, record


TRIAGE_QUESTIONS = {
    "department": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
            "billing": "Payments, invoicing, refunds",
            "technical": "Bugs, outages, integrations",
            "sales": "Pricing, upgrades, new accounts",
        },
    },
    "is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"},
}

if __name__ == "__main__":
    ask_jev(
        "Help! My payouts have been failing for 3 days.",
        TRIAGE_QUESTIONS,
        project="support",
        service="ticket-triage",
        thresholds={"department": 0.8, "is_urgent": 0.7},
        thresholds_version="triage-thresholds@4",
    )

A few design choices are worth calling out:

  • Hash the questions, don't store them. questions_hash is a SHA-256 of the question definitions, serialized with sorted keys. Keep the definitions in your repo. The hash then points to a reviewed commit, and any instruction edit shows up as a new hash.
  • Summarize the answers. The record keeps each answer's value and confidence but drops the full probability map. Keep it if your reviews need it.
  • Log the gate, not just the answer. For Choice and Score, the gate compares confidence to your threshold. For Noul, it compares the probability itself. The thresholds and their version go in the record, because they're your rule, not Jev's.
  • Failed calls get a record too. A 429 or 529 is written with its status and latency before the error is raised, so rate-limit trends show up per project.
  • No retries in the sketch. If you call the HTTP API directly, add exponential backoff on 429 and 529 as TypeSafe's docs advise, and log each attempt. If you use an official SDK, the docs say it retries by default, so put the wrapper around the SDK call instead.
  • Pin on purpose. jev-1.13.0 is the current versioned ID on TypeSafe's models page. The docs recommend pinning a versioned ID if you've tuned confidence thresholds against it, and moving on your own schedule.

Expected output

Running the script with JEV_URL pointed at a local stand-in that returns the sample Choice and Noul answers from TypeSafe's API reference writes one line like this (pretty-printed here; ts, request_id and latency_ms will differ):

{
  "ts": "2026-10-09T16:06:13Z",
  "request_id": "a4c7d420-14fe-4636-be7c-043265dfc458",
  "caller": {"project": "support", "service": "ticket-triage"},
  "model_requested": "jev-1.13.0",
  "questions_hash": "sha256:4d3a3eb496b797909182b84dff1f02d2f359706ceaa224bbc0bdc8acfd098223",
  "thresholds": {"department": 0.8, "is_urgent": 0.7},
  "thresholds_version": "triage-thresholds@4",
  "status": 200,
  "latency_ms": 10,
  "model_answered": "jev-1.13.0",
  "input_tokens": 318,
  "answers": {
    "department": {"type": "choice", "value": "billing", "confidence": 0.81, "gate": "accept"},
    "is_urgent": {"type": "noul", "value": 0.95, "gate": "accept"}
  }
}

Raise the department threshold to 0.85 and the same answer is recorded with "gate": "escalate", meaning a person picks the queue. A small test suite checks exactly that, along with the 429 path, the stability of the hash, and that the record contains no state, no instruction text and no key. It runs against the local stand-in only. This code has not been tested against the live Jev API, so try it in a non-production project first.

What stays out by default

The record leaves out three things on purpose:

  • Raw state. It's often the most sensitive thing in the request: support messages, records, chat logs. A log full of it becomes a risk of its own.
  • Instruction text. The hash points to the version in your repo. There's no need to copy it into every log line.
  • The key. The Authorization header never touches the record.

If you need request bodies, for debugging a specific disputed decision for example, make that a deliberate choice: turn it on per project, redact personal data first, and keep retention short.

Questions the record answers

With records like these in one table, common questions become simple queries:

  • Which version answered last Tuesday's batch? Filter by time and project, and group by model_answered.
  • How often do we escalate? Count gate = escalate per question per week. If a gate never fires, the threshold may be too loose. If it fires constantly, look at the questions.
  • Did anything shift when the alias moved? If you're on jev-latest, compare confidence distributions before and after a change in model_answered.
  • What does each project cost? Sum input_tokens by project and multiply by the current price. TypeSafe's models page listed $0.042 per million input tokens on Oct 9, 2026, so check it before you rely on it.
  • Are we hitting limits? Count 429s by project and hour. TypeSafe says its rate limits are "adjusting dynamically" and can change without notice.

The other half: a change log

Per-call records tell you what happened. A change log tells you why behavior changed. Keep an append-only log of pin moves, threshold changes, instruction edits and key rotations, each with who, when, from what, to what, and why. With questions_hash and thresholds_version in every call record, you can line each call up against the changes that preceded it. The checklist post covers the change log in more detail.

What's next

The natural next step is comparing versions on your own traffic before moving a pin or shipping an instruction edit: choice flips, threshold flips, confidence shifts. In ModelMesa, replay against a new Jev version or revised instructions is coming soon. It will run only on traffic a project opts in to capture, and it's off by default.

A record like this only covers calls that go through the wrapper. Calls made with a copy of the key that skip it won't show up, so make the wrapper the only way your services call Jev.


We're building ModelMesa to keep this record automatically, along with project-scoped keys, rules for who can call Jev, spend caps, and an audit log of keys, policy changes and decision calls.

Want a system of record for the Jev traffic that goes through ModelMesa's gateway? Join the early-access waitlist.

The ModelMesa team