AI & Machine Learning

The Jev principle explained: AI that ticks boxes instead of writing

Jev does not write text; it answers multiple-choice questions with a probability for each answer. The principle in plain language, with PHP examples you can copy, up to Claude and Jev working together.

Erik van de Blaak
Erik van de Blaak
22 min read • 6 views
The Jev principle explained: AI that ticks boxes instead of writing

An online shop gets two hundred emails on a Monday morning. "Where is my coat?" "Can I swap these shoes for a size up?" "You charged me twice!" Someone reads them one by one and drags each into the right folder. Every email takes a few seconds, and not a single word gets written.

That is exactly the kind of work TypeSafe AI built Jev for, an AI model announced on 15 September 2026. Jev writes nothing. It reads a text, gets a few questions with fixed answers, and ticks boxes, with how sure it is next to each one. This article explains that principle with everyday examples and short pieces of PHP you can copy, up to a setup in which Claude and Jev work together. Benchmarks and price comparisons are in my earlier article on Jev.

Checked against TypeSafe's documentation on 1 October 2026. The answers from Jev and Claude in the examples are illustrations: they show what an answer looks like, they are not results of real calls.

What is the Jev principle?

Jev answers multiple-choice questions about a text. You supply the text and the questions, including every allowed answer. For each question Jev returns the most likely answer plus a probability for every possible answer. What happens next is decided by your own code.

The difference with ChatGPT, Claude or Gemini is easiest to see in an exam. A language model gets an open question and writes an answer, word by word. That answer can go anywhere: it can ramble, invent a category that does not exist, or neatly say "delivery". Jev gets a multiple-choice question and can only tick a box that is on the form. That is where TypeSafe's claim comes from that Jev does not hallucinate: there is no room to make anything up. Ticking the wrong box is still possible. I will come back to that.

Because Jev does not have to write anything, it is fast and cheap. A language model computes again for every word; according to TypeSafe, Jev computes all probabilities in one go. TypeSafe states 70 to 500 milliseconds per call, independent measurements came out around 0.4 seconds. Input costs 0.042 dollars per million tokens (a token is roughly a syllable or a short word) and the answer is free.

Left: an LLM loops through the model for every token and produces text that still has to be parsed. Right: Jev does one pass and returns a probability distribution over the supplied options.
Left, a language model writing word by word; right, Jev computing a probability per option in one go. The list of options is everything Jev can answer. (The probabilities in the figure are an illustration.)

Jev's three question types

Jev has three kinds of question. That is all there is.

Noul: yes or no, with a probability

A Noul is a yes/no question. The answer is a single number between 0 and 1: the probability that the answer is "yes". Read it like the chance of rain in a weather forecast. 0.9 is a clear yes, 0.1 a clear no, and around 0.5 Jev does not know.

The number is a probability, not a measure. Ask "Is the customer angry?" and get 0.6, and it does not mean "a bit angry". It means a sixty percent chance that the answer is yes. The documentation warns about this itself. If you want to know how angry, use a Score.

Choice: one box from a list

A Choice picks one option from a list you provide, up to 255 options. You get the chosen option, a probability per option (adding up to 1) and a number for certainty. Jev always has to pick something, so the documentation recommends an "other" option. Without it, an unsolicited job application in your customer service inbox still gets squeezed into "return".

Score: a place on a scale

A Score places the text on a scale of 2 to 10 levels that you describe in words, for example "calm", "irritated" and "furious". The result is a weighted average of the level numbers, which start at 0. That is why you often get a decimal: 1.4 sits between "irritated" (1) and "furious" (2), closer to irritated.

Your first Jev call in PHP

TypeSafe has official libraries for Python and JavaScript, not for PHP. That hardly matters: Jev is a single web address you send a JSON message to. This function does that with curl, which ships with almost every PHP installation. I built it myself from the API documentation; it is not TypeSafe's code.

function askJev(string|array $text, array $questions): array
{
    $ch = curl_init('https://api.typesafe.ai/v1/systemone');
    curl_setopt_array($ch, [
        CURLOPT_POST           => true,
        CURLOPT_RETURNTRANSFER => true,
        CURLOPT_TIMEOUT        => 10,
        CURLOPT_HTTPHEADER     => [
            'Authorization: Bearer ' . getenv('TYPESAFE_API_KEY'),
            'Content-Type: application/json',
        ],
        CURLOPT_POSTFIELDS     => json_encode([
            'model'     => 'jev-latest',
            'state'     => $text,
            'questions' => $questions,
        ], JSON_THROW_ON_ERROR),
    ]);

    $body   = curl_exec($ch);
    $status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);

    if ($body === false || $status !== 200) {
        throw new RuntimeException("Jev returned no usable answer (HTTP $status)");
    }

    return json_decode($body, true, flags: JSON_THROW_ON_ERROR)['answers'];
}

Three things go in. model says which version of Jev you use, state is the text you are asking about, and questions are your questions. What comes back is a list of answers, one per question. You create the API key in TypeSafe's console and put it in an environment variable, so it does not live in your code.

Example 1: an online shop's inbox

$mail = 'Hi, last week I ordered a winter coat (order 48213) and I still have not '
      . 'received anything. I leave for a ski trip on Saturday, so I really need it!';

$topicQuestion = [
    'type'         => 'choice',
    'instructions' => 'What is this customer email about?',
    'criteria'     => [
        'delivery'         => 'Parcel did not arrive, arrived late or arrived damaged',
        'return'           => 'Customer wants to send something back or exchange it',
        'payment'          => 'Invoice, charge or refund',
        'product_question' => 'Question about size, material or stock',
        'other'            => 'Fits none of the other options',
    ],
];

$answers = askJev($mail, [
    'topic'    => $topicQuestion,
    'deadline' => [
        'type'         => 'noul',
        'instructions' => 'Does the customer mention a day or date by which the product must arrive?',
    ],
]);

Each question gets a name you choose yourself, here topic and deadline. That name is where you will find the answer. In criteria you describe in plain language what each option means. With Jev, that is all the "prompting" there is: writing clear descriptions.

The answer looks like this (illustration):

{
  "topic": {
    "type": "choice",
    "choice": "delivery",
    "probabilities": { "delivery": 0.94, "return": 0.01, "payment": 0.01, "product_question": 0.01, "other": 0.03 },
    "confidence": 0.91
  },
  "deadline": { "type": "noul", "noul": 0.97 }
}

From there it is ordinary PHP again. moveToFolder() and markAsUrgent() stand for your own functions:

$topic   = $answers['topic']['choice'];      // 'delivery'
$certain = $answers['topic']['confidence'];  // 0.91
$hurry   = $answers['deadline']['noul'];     // 0.97

if ($certain >= 0.9) {
    moveToFolder($topic);
} elseif ($certain >= 0.5) {
    moveToFolder($topic, needsReview: true);
} else {
    moveToFolder('manual');
}

if ($hurry >= 0.8) {
    markAsUrgent();
}

Look at who does what. Jev says what is in the email. Your code decides what happens. Jev sends nothing, moves nothing and cannot break anything. That is the core of the principle: the AI gives a judgement in a form a computer can read, and the rules stay yours. The thresholds of 0.9 and 0.5 are TypeSafe's starting advice: above 0.9 automatic, in between with a check, below that a human.

According to TypeSafe, English is the language in which Jev is most accurate; other languages work, but less well. If your emails are not in English, test on your own data before you build on it.

What does "91 percent sure" mean?

The probabilities are the most interesting part of Jev, and the easiest to misread.

TypeSafe trains Jev for calibrated probabilities. That works like a good weather forecast: if the forecaster predicts a 70 percent chance of rain on a hundred days, it should actually rain on about seventy of them. Not on all hundred, and not on thirty. TypeSafe means it the same way: of all answers Jev puts 0.8 on, about 80 percent should be right. That says something about a group of answers, never about the single answer in front of you.

The confidence on a Choice or Score is something else. It tells you how concentrated the probabilities are: everything on one option gives nearly 1, an even spread across all options nearly 0. An email that is half about a return and half about a double charge therefore gets a low confidence, even if Jev recognises both topics perfectly well. Confidence tells you how decisive the answer is. How often such an answer is right, it does not tell you.

Whether the probabilities are honest for your texts is something you have to check yourself. Two small independent tests reached opposite conclusions on that (see the earlier article), and the documentation says so too: tune thresholds on your own data and start cautiously.

Thresholds follow from the consequences

What do you do with 0.7? That does not depend on Jev but on what it costs when the answer is wrong. For Nouls the documentation gives a rule of thumb: 0.5 when yes and no are equally easy to undo, higher when a wrong "yes" is expensive, lower when a missed "yes" is expensive.

$answers = askJev($mail, [
    'wants_return' => [
        'type'         => 'noul',
        'instructions' => 'Does the customer want to send a product back?',
    ],
    'hazard' => [
        'type'         => 'noul',
        'instructions' => 'Does the customer report smoke, fire, overheating or an electric shock from a product?',
    ],
]);

// A return form sent by mistake is awkward, but quickly put right
if ($answers['wants_return']['noul'] >= 0.7) {
    sendReturnForm();
}

// A missed safety report is far worse than a colleague checking for nothing
if ($answers['hazard']['noul'] >= 0.2) {
    alertTeamLead();
}

A missed report about a hair dryer that started smoking weighs more than a colleague reading a harmless email. That is why that threshold is low. Refunding money is something you never leave to a probability alone, however high: irreversible and financial actions need a person or a hard rule in your code.

Example 2: ask every question at once

With an ordinary language model you often ask questions one after another: first "what is it about?", and only for a return, "what is the reason?". Every follow-up question is another call and another wait.

For Jev, TypeSafe recommends the opposite and calls it speculative fan-out: ask every question you might need in a single call. Jev answers them in parallel, so an extra question costs hardly any time. Because the answer is free, you only pay for the words of the question itself. Your code ignores whatever turns out to be irrelevant.

$answers = askJev($mail, [
    'topic'         => $topicQuestion,
    'return_reason' => [
        'type'         => 'choice',
        'instructions' => 'Why does the customer want to send the product back?',
        'criteria'     => [
            'size'    => 'Does not fit, too big or too small',
            'defect'  => 'Broken, damaged or not working',
            'wrong'   => 'Received a different product than ordered',
            'taste'   => 'Colour or style is disappointing',
            'unknown' => 'The reason is not in the email',
        ],
    ],
    'anger' => [
        'type'         => 'score',
        'instructions' => 'How angry does the customer sound?',
        'criteria'     => ['Calm and matter-of-fact', 'Irritated but polite', 'Furious, harsh language'],
    ],
]);

if ($answers['topic']['choice'] === 'return') {
    $reason = $answers['return_reason']['choice'];  // only relevant for a return
}

if ($answers['anger']['score'] > 1.2) {            // scale 0 to 2
    givePriority();
}

If the email is about a late delivery, there is still an answer to the return question. You simply do not use it. It cost a few tokens, and it saves you a second call on every email that is about a return.

Example 3: assessing a CV without a black box

Ask an AI to "Rate this candidate from 1 to 10" and you get a 7 without knowing why. Was it the experience? A nicely worded cover letter?

TypeSafe describes a different approach, composite scoring: cut the judgement into small, separate questions, let Jev score each part, and add up the scores in your own code with weights you choose. An example for a PHP developer vacancy:

$cv = file_get_contents('cv-candidate-17.txt');

$levels = [
    'No evidence in the CV',
    'Mentioned, but without concrete projects',
    'One or two concrete projects',
    'Several years of demonstrable experience',
    'Deep experience: trains others or builds the foundations',
];

$answers = askJev($cv, [
    'php'      => ['type' => 'score', 'instructions' => 'Experience with PHP', 'criteria' => $levels],
    'laravel'  => ['type' => 'score', 'instructions' => 'Experience with the Laravel framework', 'criteria' => $levels],
    'teamwork' => ['type' => 'score', 'instructions' => 'Experience working in a development team', 'criteria' => $levels],
]);

// The weights are your choice, not Jev's
$weights = ['php' => 0.5, 'laravel' => 0.3, 'teamwork' => 0.2];

$total = 0;
foreach ($weights as $question => $weight) {
    $total += $weight * ($answers[$question]['score'] / 4);  // five levels: score runs from 0 to 4
}

printf("Total score: %d out of 100\n", round($total * 100));

This gives you two things. You can explain why candidate A ranks above candidate B: higher on PHP, lower on teamwork. And if the ranking does not match what a recruiter expects, you change the weights without asking Jev anything again. For a team lead role you move the teamwork weight up, with the same scores.

Scores from different questions are not aligned with each other: a 3 on PHP and a 3 on teamwork need not be equally strong. The weights are a choice, and therefore your responsibility. Travel time is deliberately not in the list. Whether Arnhem is close enough to Utrecht is a calculation, not a judgement; more on that further down.

Example 4: making Claude and Jev work together

Jev does not write code or hold a conversation. Claude does. So the two complement each other: Claude analyses, reasons and writes; Jev makes the bounded choices along the way, with a probability attached. Your PHP code connects them and stays in charge. TypeSafe says as much: Jev does not replace the model behind a coding agent, but it can handle the routing and assessment steps for one.

Take a bug report: "Our PHP script sometimes saves duplicate orders." Claude reads the code and the log and arrives at a few possible causes. Which do you investigate first? That is a multiple-choice question, and Claude can put it to Jev. For that you give Claude a tool: a function Claude may call on its own when it needs to. This example uses Anthropic's official PHP SDK (composer require anthropic-ai/sdk) and the askJev() function from above.

First the tool. Its description tells Claude when it is meant to be used; the function below passes Claude's question to Jev and returns the answer:

$jevTool = [
    'name'        => 'choose_with_jev',
    'description' => 'Let Jev make a bounded choice between a few clear options, for example '
                   . 'which possible cause to investigate first. Pass the relevant facts as '
                   . 'state and always include an "other" option.',
    'inputSchema' => [
        'type'       => 'object',
        'properties' => [
            'state'    => ['type' => 'string', 'description' => 'The facts the choice rests on'],
            'question' => ['type' => 'string', 'description' => 'The question, short and direct'],
            'options'  => [
                'type'                 => 'object',
                'description'          => 'A short name per option with a description',
                'additionalProperties' => ['type' => 'string'],
            ],
        ],
        'required'   => ['state', 'question', 'options'],
    ],
];

function runTool(string $name, array $input): string
{
    if ($name !== 'choose_with_jev') {
        return "Unknown tool: $name";
    }

    $answer = askJev($input['state'], [
        'choice' => [
            'type'         => 'choice',
            'instructions' => $input['question'],
            'criteria'     => $input['options'],
        ],
    ])['choice'];

    return json_encode([
        'choice'        => $answer['choice'],
        'probabilities' => $answer['probabilities'],
        'confidence'    => $answer['confidence'],
    ]);
}

Then the conversation with Claude. Claude gets the task, the code and part of the log. As long as Claude wants to use a tool (stopReason is then tool_use), PHP runs it and sends the result back:

use Anthropic\Client;
use Anthropic\Messages\ToolUseBlock;

$claude = new Client(apiKey: getenv('ANTHROPIC_API_KEY'));

$system = 'You are an experienced PHP developer investigating bugs. When you face a bounded choice '
        . 'between a few clear options, use choose_with_jev. The probabilities are advice: you remain '
        . 'responsible for the analysis. If the confidence is below 0.5, also investigate the second option.';

$messages = [[
    'role'    => 'user',
    'content' => "Our script sometimes saves duplicate orders. Find out why and propose a fix.\n\n"
               . "Code:\n" . file_get_contents('save-order.php') . "\n\n"
               . "Log:\n" . file_get_contents('logs/duplicate-orders.log'),
]];

do {
    $response = $claude->messages->create(
        model: 'claude-opus-5-5',
        maxTokens: 16000,
        system: [['type' => 'text', 'text' => $system]],
        tools: [$jevTool],
        messages: $messages,
    );

    $messages[] = ['role' => 'assistant', 'content' => $response->content];

    $results = [];
    foreach ($response->content as $block) {
        if ($block instanceof ToolUseBlock) {
            $results[] = [
                'type'      => 'tool_result',
                'toolUseID' => $block->id,
                'content'   => runTool($block->name, $block->input),
            ];
        }
    }
    if ($results) {
        $messages[] = ['role' => 'user', 'content' => $results];
    }
} while ($response->stopReason === 'tool_use');

foreach ($response->content as $block) {
    if ($block->type === 'text') {
        echo $block->text;  // Claude's analysis and proposal
    }
}

Here is what happens, step by step. Claude reads the code and the log and sees, say, three candidates: two simultaneous requests that both save an order (a race condition), a missing unique index in the database, or a form that gets submitted twice. Claude calls choose_with_jev with those three options plus "other", and with the facts from the log as state. The answer that comes back looks like this (illustration):

{
  "choice": "race_condition",
  "probabilities": { "race_condition": 0.62, "no_unique_index": 0.24, "double_submit": 0.11, "other": 0.03 },
  "confidence": 0.47
}

Claude investigates the race condition first. Because the confidence is below 0.5, Claude also looks at the unique index, as instructed, and then writes an analysis with a proposal. With duplicate database records the answer is often both: a unique index that rejects the duplicate, and code that handles that rejection properly.

Two caveats. Jev only knows what Claude puts in the state: if Claude summarises the log wrongly, Jev chooses on the wrong facts. And deriving a cause from code and a log is multi-step reasoning, exactly the kind of work TypeSafe says Jev is weaker at. Whether Jev chooses better here than Claude would on its own, nobody has measured as far as I know. What Jev does add is an answer in the same shape every time, with a number you can log and check afterwards, at a fraction of the price: Jev charges 0.042 dollars per million input tokens, Claude Opus 5.5 4 dollars.

The brake belongs in your code, not with Claude

In the example above, Claude decides when Jev gets asked. That is advice. A check that must never be skipped is something your own code does, outside Claude. For instance, before Claude's proposal is applied:

$risk = askJev($proposal, [
    'touches_data' => [
        'type'         => 'noul',
        'instructions' => 'Does this change delete or modify existing data or the database structure?',
    ],
]);

if ($risk['touches_data']['noul'] >= 0.3) {
    queueForReview($proposal);  // a person looks first
} else {
    runTests($proposal);
}

Claude cannot forget this step or reason its way around it, because Claude does not even know it exists. It is the same principle as with the inbox: the model gives a judgement, your code decides. The threshold is deliberately low, because an unnecessary review costs a few minutes and a lost table costs a lot more.

And in Claude Code itself?

If you work with Claude Code, you can capture the same idea as a skill: an instruction file that tells Claude when and how to do something. TypeSafe has its own plugin for Claude Code (claude plugin install typesafe@typesafe-ai), but it mainly teaches Claude how to build the Jev API into your code. With an API key Claude can also run test queries with it, but that does not make Claude put its own choices to Jev by itself.

If you do want that, you write a small skill, for example in .claude/skills/jev-triage/SKILL.md. This is my own setup, untested:

---
name: jev-triage
description: Use for a bug with several clear possible causes, to decide which one to investigate first.
---
1. Write the relevant facts and the possible causes to question.json, always with an "other" option.
2. Run: php jev-choice.php question.json
3. Investigate the causes in order of probability. With a confidence below 0.5, also check the second option.
4. The result is advice. If your own investigation contradicts it, follow your investigation and say so.

The script the skill calls is again plain PHP with the function from above:

require __DIR__ . '/ask-jev.php';  // contains askJev()

$question = json_decode(file_get_contents($argv[1]), true, flags: JSON_THROW_ON_ERROR);

$answer = askJev($question['state'], [
    'choice' => ['type' => 'choice', 'instructions' => $question['question'], 'criteria' => $question['options']],
])['choice'];

arsort($answer['probabilities']);
foreach ($answer['probabilities'] as $option => $probability) {
    printf("%-20s %3d%%\n", $option, round($probability * 100));
}
printf("confidence: %.2f\n", $answer['confidence']);

Claude then sees a list like race_condition 62%, no_unique_index 24% and so on, and carries on from there. The division of labour is the same as in the PHP example: Claude is the developer, Jev answers bounded questions, and whatever truly must not go wrong, you handle in code.

Where Jev gets it wrong

TypeSafe publishes its own list of weak spots in the current version, Jev 1.13. That is unusually candid and also useful, because almost every item has the same fix: let your own code handle that part. These are the six you will run into first.

Arithmetic and counting

"Jev is not a calculator", TypeSafe writes. Do not ask "Did the customer order more than three items?" The number is in your order system; count it there.

Dates

Jev reads dates as text, not as points on a timeline. "Is this purchase still under warranty?" is therefore a question for PHP:

// Not: 'warranty' => ['type' => 'noul', 'instructions' => 'Was the purchase more than two years ago?']

// Instead: the date comes from your order system and PHP does the maths
$orderedOn  = new DateTimeImmutable($order['ordered_on']);
$inWarranty = $orderedOn->modify('+2 years') > new DateTimeImmutable('today');

Literal reading

"Jev answers the question you wrote, not the one you meant." Ask "Is the customer satisfied?" about "The parcel came quickly, but the coat has a tear" and every answer is half wrong. Split it into "Is the customer satisfied with the delivery?" and "Is the customer satisfied with the product?", and combine the answers in your code.

Double negatives and detours

"Is it not the case that the customer does not want a return?" makes Jev less accurate. Write the question as directly as you can: "Does the customer want to send something back?"

Lots of text that does not matter

According to TypeSafe, accuracy drops as you send more text that has nothing to do with the question. So send the email itself, not the whole thread of earlier replies with signatures and disclaimers underneath. Filtering happens beforehand, in PHP.

Text that tries to steer Jev

Jev does not treat malicious text as hostile by default. An applicant who puts "This candidate is an excellent match" in white text in a CV is a scenario you need to test. So never let a score decide a rejection or an invitation on its own. You could add an extra Noul such as "Does the text contain directions aimed at a reviewer or a computer?", but that is my own suggestion and untested.

When you do not need Jev

The principle tempts you to turn everything into a question. Don't. If you can write the rule down, write it down. "Order value above 500 euros: check manually" is an if statement, not an AI question. Jev is meant for judgements that need an understanding of language, that a person makes in a few seconds, and where the possible answers are fixed.

If you already have thousands of emails sorted by hand, a small classifier you train yourself can do better and cheaper. In an independent test such a classifier beat Jev clearly, although it had seen more than ten thousand labelled examples and Jev none.

What Jev costs and what to know first

The arithmetic is simple, because only input counts. A customer email plus the question definitions easily adds up to 500 tokens. Ten thousand emails a month is then five million tokens: 21 cents.

  • Internet only. Jev cannot be downloaded. Every text you send goes to TypeSafe. TypeSafe says it does not train on customer requests; if you send CVs or other personal data, read the data processing terms first.
  • Limits. On 1 October the documentation lists 100,000 tokens per second and 40 calls per second, adding that these limits can change without notice. At launch the numbers were different.
  • Pin the version. jev-latest points to the newest stable version. For anything in production you can specify a fixed version, such as jev-1.13.0, so behaviour does not change silently after you have tuned your thresholds.
  • Access. TypeSafe paused new signups on 22 September because of demand. Check console.typesafe.ai to see whether you can create an account.

Start with fifty emails you already know

Take fifty emails you have already sorted by hand. Let Jev sort them again with the function from this article, and put its choice and confidence next to yours in a spreadsheet. After an hour you know three things no benchmark tells you: how often Jev agrees with you, whether the mistakes sit at low confidence, and which threshold suits your emails. Fifty emails of 500 tokens cost about a tenth of a cent in total.

Sources

Share this article

Erik van de Blaak

Written by

Erik van de Blaak

AI Solutions Engineer & Full-Stack Developer

Erik van de Blaak is an AI Solutions Engineer and full-stack developer at CareerValue BV. Here he writes about AI coding tools and agents, such as Claude Code.

Comments (0)

Comment on this article

Never published

I read every comment before it goes online.

No comments on this article yet.

Read next