What is Jev?

Everyone is talking about it this week. Here’s what it is, told small — and then one real question about using it with four agents, worked all the way through.

Jev is a new AI model from a startup called TypeSafe AI. It launched into early access in September 2026 and everyone is talking about it because it is deliberately not a chatbot.

A normal AI model

You ask a question. It writes you a paragraph. It might be right, it might make something up, and you have to read the paragraph to find out. Takes a second or two and costs real money for every word it writes.

Jev

You ask a question and give it a list of allowed answers. It picks one and tells you how sure it is, as a number. It cannot write words. It cannot go off your list. Takes about a tenth of a second and costs almost nothing.

TypeSafe calls this a “System One model” — the fast, gut-feeling part of thinking, not the slow, talky part. It’s built for other programs to ask it quick questions, thousands of times a second: which team gets this ticket, is this a spam message, is this tool call dangerous, which agent should go next.

Three things to know before you use it: it’s closed-source and API-only (no downloading the model), it’s US-hosted so don’t send it sensitive data, and it can still pick the wrong answer — it just can’t pick an answer you didn’t offer.

An example: my question

Here is the question that started this page:

“I have four agents. Can I use Jev as the hand-off, or as an orchestrator that decides which tool call should be made from which agent?”
Short answer: Jev is not the boss. Jev is the boss’s pointing finger. Your code is still the boss.

Let’s see why, with the four agents drawn as four helpers.

You have four helpers

Imagine four friends in a room. Each one is good at one thing. When a new job comes in, somebody has to decide who gets it.

đŸ—ș
Planner
breaks big asks into steps
🔹
Builder
writes the code
đŸ§Ș
Tester
runs and writes tests
🔍
Reviewer
checks for mistakes

Jev is a helper who can’t talk

Jev never writes sentences. Jev can only do one thing: you show Jev a message and a list of choices, and Jev points at one choice and tells you how sure it is. That’s it. No chatting, no explaining, no making things up — because Jev literally cannot say anything that isn’t on your list.

That sounds like a small thing. It is exactly the thing an orchestrator needs most, and it happens in a fraction of a second for a fraction of a cent.

Try it: who gets this job?

Pick a message. Watch Jev point. (These numbers are pretend — the real Jev returns real ones.)

Pick a message above.

So is Jev the orchestrator?

Half right. There are two different jobs hiding inside the word “orchestrator”:

Deciding — Jev’s job ✅

“Which helper should take this?” “Is this tool call safe to run?” “How risky is this change, 0 to 2?” These are pick-one questions. Jev answers them fast, with a probability you can put a threshold on.

Handing off — your code’s job ❌

Actually calling the Builder, passing the message along, keeping the conversation history, retrying, logging, asking a human when Jev is unsure. Jev can’t do any of this. It only points.

So the picture is: your orchestrator is a small loop of normal code. Every time it has to make a fuzzy decision, it asks Jev instead of writing a giant if-else or calling a slow chat model.

How one turn goes

  1. A message arrives. Could be from a user, could be output from one of your agents saying “I’m done, here’s what I found.”
  2. The boss asks Jev. “Here’s the message. Which of these four should take it next? Also: does it need a human?”
  3. Jev points. Tester, 91% sure. Human needed: 4%.
  4. The boss hands off. Your code calls the Tester agent with the message. If Jev had been under, say, 60% sure, the boss asks a person instead.
  5. Repeat. Tester finishes, its output becomes the next message, back to step 2.

Three kinds of questions Jev can answer

Which one?

choice — pick from your list. Use it for “which agent.”

Yes or no?

noul — a 0 to 1 probability. Use it for “should this tool call run?”

How much?

score — a spot on a scale you define. Use it for “how risky, 0–2?”

The grown-up version

This is what the boss actually sends to Jev for one turn. Notice the four helpers are just the four keys under criteria.

POST https://api.typesafe.ai/v1/systemone

{
  "state": "Tests in PaymentServiceTest are failing after the refactor. Stack trace attached.",
  "model": "jev-latest",
  "questions": {
    "next_agent": {
      "type": "choice",
      "instructions": "Which agent should handle this next?",
      "criteria": {
        "planner":  "Vague or large request that needs breaking down",
        "builder":  "Clear code change to implement",
        "tester":   "Failing tests, missing coverage, verification",
        "reviewer": "Finished work that needs a quality check"
      }
    },
    "needs_human": {
      "type": "noul",
      "instructions": "Is this too ambiguous or risky for an agent to handle alone?"
    }
  }
}

And what comes back — no prose, just a decision your code can branch on:

{
  "answers": {
    "next_agent": { "choice": "tester", "probabilities": { "tester": 0.91, "builder": 0.06, "reviewer": 0.02, "planner": 0.01 }, "confidence": 0.88 },
    "needs_human": { "noul": 0.04 }
  }
}

Then the boss does the handoff:

if (needsHuman > 0.5 || confidence < 0.6) askAPerson(message);
else agents.get(nextAgent).run(message);

“It uses fewer tokens” — is that true?

Half true, and the half that’s true is the half that mattered. Think of tokens as words you pay for. Every AI model charges for the words you send in and the words it sends back — and the words coming back cost about five times more.

Words going in: same as any model

Jev still has to read your message and your list of choices. For the four-agent router above that’s roughly 200–300 tokens per decision. A longer message means a bigger bill, just like everywhere else. Jev doesn’t shrink what you send.

Words coming back: zero

Jev doesn’t write sentences, so there is nothing to charge for. TypeSafe bills output at $0. A chat-model router would write 20–150 tokens of JSON every time — or thousands if it “thinks” first. That’s where the money was going.

The savings people forget to count

  1. No “you are a router” boilerplate. With a chat model you spend 200–500 tokens every call explaining the rules and the JSON schema. With Jev the schema is the request.
  2. No retries. Jev can’t return broken JSON or add commentary, so you never re-send the whole prompt because the model wrapped its answer in prose.
  3. No thinking tokens. A reasoning model can burn 2,000 tokens deciding between four agents. Jev can’t.

What your orchestrator would pay

DecisionsJev (~300 tokens in, 0 out)Small chat model (~500 in, 50 out)
1$0.0000126$0.00075
100,000$1.26$75
1,000,000$12.60$750

Jev at $0.042 per million input tokens, output free. Chat model at a typical $1 in / $5 out per million. Rough numbers, but the gap is real: about 60x cheaper on routing.

The honest caveats

It’s early-access pricing

Treat $0.042 as today’s floor, not a promise. If output gets billed later you’re still fine — output is tiny — but budget for it.

Your agents still cost the same

Jev only replaces the routing decision. Your four helpers still run on a real LLM. Routing was maybe 10–20% of your token bill; that’s the part that goes to near zero.

Speed is the bigger win

70–500 ms per decision means you can afford to ask Jev on every handoff instead of skipping some to save time.

The number is the real prize

“91% sure” is something you can set a threshold on. A paragraph from a chat model is not. That’s worth more than the cost saving.

When Jev is the wrong helper

You need words back

Jev can’t write the plan, the code, or the review. Your four agents still need a real language model for that. Jev only picks who goes next.

Only one possible next step

If Builder always hands to Tester, that’s a plain line of code. Don’t pay for a decision that isn’t one.

Sensitive data

Every question goes to a US-hosted API. Don’t send real customer data or internal logs unless your data policy allows it.

Jev can still be wrong

It can’t answer off-list, but it can pick the wrong item. Always keep the “low confidence → ask a human” branch.