Blog

Who called Jev, under which rules, and when? A system-of-record checklist

Jev is easy to adopt: one endpoint, a bearer key, typed questions in, typed answers out. That simplicity is also why governance tends to lag behind. A few weeks after the first integration, the key lives in three services and a CI job, someone has edited a question's instructions twice, and nobody is quite sure which Jev version answered last Tuesday's batch.

That's fine until someone asks about a specific decision. A customer disputes an outcome. A reviewer asks why a ticket was flagged. Finance asks which team's usage doubled. At that point you need a system of record: a dependable account of who called what, under which rules, and when.

This post covers what that record should contain, and gives you a checklist you can apply today with your own logging, whether or not you ever use ModelMesa.

Why decision calls need a record

A Jev call isn't a chat transcript. You send state plus a map of named, typed questions (choice, score, noul) and get back typed answers. Per TypeSafe's API reference, choice and score answers include probabilities and a confidence value. Those answers usually drive something downstream: a triage step, a flag, a threshold check.

Three things can change the answer to the same input:

  1. The model version. The jev-latest alias moves when a new release ships. TypeSafe's models page warns that "the answers behind it can change without a change on your side."
  2. Your instructions. Each question carries the instructions and criteria you write. Edit them and the answers can shift.
  3. Your preprocessing. If you redact or mask data before sending, Jev sees different input.

If your logs don't capture all three, you can see that an answer changed but not why.

What a useful record contains

Who. The service, environment and project that made the call. If several services share one key, nothing in the key tells you which one called. Attribution has to come from your side.

What. The model you requested and the versioned model ID that actually answered (TypeSafe's docs say the response's model field reports it "so you can log which model produced each result"). Also the question names, the answers and their confidence, usage.input_tokens, the HTTP status, and latency.

Under which rules. A version or hash of your question definitions, the model pin in effect, the confidence thresholds your code applied, and your data-handling settings (for example, whether personal data was redacted first).

When, and who changed what. A timestamp on every call, plus a separate change log for anything that alters behavior: key creation and rotation, pin changes, instruction edits, threshold changes. Each entry records who, when, from what to what, and why.

An illustrative per-call record (field names are ours, not TypeSafe's):

{
  "ts": "2026-10-12T15:04:05Z",
  "request_id": "req_01J...",
  "caller": { "service": "support-triage", "env": "prod", "project": "support" },
  "model_requested": "jev-1.13.0",
  "model_answered": "jev-1.13.0",
  "questions_hash": "sha256:9f2c...",
  "thresholds_version": "triage-thresholds@4",
  "answers": { "category": { "value": "billing", "confidence": 0.87 } },
  "input_tokens": 1840,
  "status": 200,
  "latency_ms": 412
}

Notice what's missing: the raw state. Keep metadata by default, and make full request bodies opt-in, with redaction and a short retention window. state often contains personal data, and a log full of it becomes a risk of its own.

A checklist you can apply today

  1. Inventory your key. List every service, job and machine that holds your TypeSafe key. If you can't produce that list, that's your first finding.
  2. Send every call through one wrapper. Make a single internal function that requires a service and project name and writes the record above. No untagged calls.
  3. Log the version that answered, not just the one you asked for. Store the response's model field on every call.
  4. Pin deliberately. If you've tuned confidence thresholds against a version, TypeSafe's docs recommend pinning that version's ID (currently jev-1.13.0) instead of the alias and moving on your own schedule. Record every pin change.
  5. Version your questions like code. Keep question definitions in your repo, hash them, and log the hash with each call. An instruction edit should be a reviewed commit, not a quiet config change.
  6. Keep an append-only change log for key rotations, pin moves, and threshold and instruction changes. Even an insert-only database table beats searching chat history.
  7. Estimate cost at the source. TypeSafe prices Jev per input token, and output tokens are free. Its models page listed $0.042 per million input tokens when we checked on Oct 8, 2026, so check the current price. Multiply usage.input_tokens by that price per project and you have estimated cost by team before the invoice arrives.
  8. Log upstream errors separately. Track 401, 422, 429 and 529 as distinct statuses. TypeSafe says its rate limits are "adjusting dynamically" and can change without notice, so you want to see a 429 trend early.
  9. Decide retention on purpose. Write down how long metadata lives, whether bodies are stored at all, and who can view them.
  10. Test the record. Pick a random call from last week and answer: which service, which version, which instructions, which thresholds, and who last changed any of them. If that takes more than a few minutes, close the gaps.

The limit of any record

A system of record only covers what passes through it. Calls that skip your wrapper, or go out on a copy of the key nobody inventoried, won't show up. That's why steps 1 and 2 matter more than any dashboard.


We're building ModelMesa to keep this record for Jev teams, covering the traffic that goes through its gateway: project-scoped keys, rules for who can call Jev, spend caps, and an audit log. Early access is open at modelmesa.com.

The ModelMesa team