Skip to content
Book a Call → mycocoon.life
Jev Lab

The model that never writes
ask it a question, get a probability back

Everyone has met a chatbot. Jev, the System One model from TypeSafe, is different: it reads a message and returns a typed answer with a calibrated probability, never a paragraph. Pick how you want to meet it, or read what Jev is first.

Live calls to Jev · no code to type · free, one email opens every lab
checking…
Simulated answers. This deployment has no TypeSafe key yet, so the numbers on this page are produced by a small keyword heuristic in your browser and are marked simulated. The requests you build are real and can be sent as they are once the key is set.
Before anyone presses anything, ask the room: what would your software do with each answer? Then send. Point at the shape, not the content: one answer needs a person to read it, the other goes straight into an if.
Stage 1 · The big idea

A chatbot writes. Jev decides.

Send the same customer message to both. Don't read the answers yet. Look at their shape: one is a paragraph a person has to read, the other is three numbers a computer can act on.

💬 Explain it to me

You've probably used a chatbot like ChatGPT. You ask, it writes you a paragraph.

Jev is different. Think of a judge on a talent show: no speech, just a score card held up.

Let's send the same message to both and see.

Your to-do list
  1. Press the green Send to both button.
  2. Answer the question at the bottom of this card.
✅ Left side: lots of words for a person to read. Right side: three numbers an app can use straight away. That's the whole idea.
Edit the message if you want, then press Send to both. Try removing "before Friday" and sending again.
A chatbot (writes)
Waiting…
Jev (decides)
Waiting…
Request sent to Jev
Reply from Jev
The chatbot's answer needs a person. Jev's answer is three numbers your software can act on: which team, how urgent, how upset. No parsing, no "please reply in JSON", no hoping it followed instructions.
Check
Which answer could a computer program use directly, with no person reading it?
Jev returns a typed answer: a label and a probability. Code can compare a number to a threshold. A paragraph has to be read, or fed to yet another model to interpret.
Ask the room what "70% sure" means. Most will say "fairly confident". Slide to 70 and count the envelopes: seven are yes. Then scroll to the real evidence and read it out. It is recorded, not made up.
Stage 2 · Reading the number

What does "70%" actually mean?

Jev never says "yes". It says how likely yes is. Calibrated means the number is honest: take every message it rated around 70%, and about seven in ten really are yes. Not "sounds confident". Actually right that often.

💬 Explain it to me

When a weather app says 70% chance of rain, it doesn't mean "a bit of rain".

It means: on days like this, it rains about 7 times out of 10.

Jev's numbers work the same way. When the numbers are honest like that, we call the model calibrated.

Your to-do list
  1. Drag the slider and watch the envelopes change.
  2. Scroll down to the green Real evidence box. Find the one message Jev wasn't sure about.
  3. Answer the question at the bottom.
✅ Remember: 50% doesn't mean "medium". It means "I can't tell".
Slide the number and watch the envelopes. Then read the real evidence under them: what Jev actually said on messages we checked by hand.
Jev says the message asks for a refund with probability70%
0.5 is not "medium". It means yes and no are equally likely: the model cannot tell from what it was given. Usually that means the message is genuinely ambiguous, or the question was.
Check
Jev says 0.50 for "is this urgent?". What does that mean?
A probability is not an intensity. 0.5 means "I can't tell from this". If you want intensity, that is a Score question with rungs, which comes next.
Run the six tasks as a hands-up vote before tapping. The last one, the lead likely to buy, splits the room. That argument is the lesson.
Stage 3 · The three shapes

Every question is a dial, bins, or rungs

Jev only knows three shapes of question. Pick the shape first, then write the question.

💬 Explain it to me

Jev only answers three kinds of question.

A yes/no question, called a Noul: like a light switch.

A pick-one question, called a Choice: like the Sorting Hat choosing a house.

A rating, called a Score: like giving a film 1 to 5 stars.

Your to-do list
  1. Sort all six example questions. Tap the kind you think fits.
  2. Answer the question at the bottom.
✅ If two things can both be true at once, ask two yes/no questions, not one pick-one.

Noul A dial

Does a condition hold? One probability of yes. Use it for anything you would put in an if.

Choice Sorting bins

Which one of these? Probability spread across options you define, plus a confidence number for how concentrated it is.

Score A ruler with rungs

How much, on a ladder you describe? Each rung is a concrete situation in words. The score is the weighted position.

Six everyday tasks. Tap the shape you would use. Being wrong once is the point.
Rule of thumb: a yes/no you would put in an if is a Noul. A pick-one you would put in a switch is a Choice. A "how bad / how much" you would compare with a threshold is a Score.
Check
A message might ask to cancel AND ask for a refund at the same time. How do you ask?
A Choice picks exactly one bin, so it cannot say "both". When several things can be true at once, ask one Noul per thing. They run in parallel in the same call.
Press Ask Jev, then press Remove "explicitly". The number jumps. Then press Change to "want". Ask the room which wording is right. There is no right one: the model answers exactly the question you wrote.
Stage 4 · Your first question

Two things and nothing else: state and a question

State is what the model reads. A question is a shape, an instruction in plain words, and what each answer means. That is the entire interface. There is no prompt.

💬 Explain it to me

Your turn to ask Jev something.

Jev is very literal, like a genie granting wishes: it answers exactly what you asked, not what you meant.

The person in this message paid twice and is annoyed. But did they actually say "give me my money back"? Watch what one word does.

Your to-do list
  1. Press Ask Jev and look at the percentage.
  2. Press Remove "explicitly". Did the number go up?
  3. Press Change "ask for" to "want".
  4. Answer the question at the bottom.
✅ Same message, different words, different answer. The list under the question shows every try.
One question is ready. Press Ask Jev. Then use the buttons under the question to change a single word, and watch the number move each time.
Add more questions from a ready-made list
Independent questions run in parallel in one call and cannot see each other.
Check
With a chatbot you type a long instruction, called a prompt. What do you give Jev instead?
Jev gets the message and your questions, nothing else. Each question says in plain words what counts as yes and what counts as no. That is why one word in the question can move the answer so much.
Ask the room to guess the severity out loud before you run it. Then show how "Low / Medium / High" smears the answer and concrete rungs settle it.
Stage 5 · The words are the model

Vague rungs get vague answers

Each rung of a Score is judged on its own. The model never sees its neighbours. "Low / Medium / High" describe nothing, so probability smears across them. Concrete situations concentrate it.

💬 Explain it to me

Imagine rating a pizza as "low", "medium" or "high". High what? Nobody knows.

Now try: 1 = burnt, 2 = fine but cold, 3 = perfect. Much clearer.

Jev needs the clear version too.

Your to-do list
  1. Press Run both.
  2. Compare where the bars land on each side, then answer the question at the bottom.
✅ Describe every level like a real situation. Vague words get vague answers.
Press Run both. Same bug report, two ways of describing the rungs. Watch where the probability lands.
Vague rungs
Concrete rungs
Score 1.4 with the mass on rungs 1 and 2 is the honest answer: broken with a workaround for most, blocking for the Safari-only customers. Vague rungs cannot express that.
Check
Your rating keeps landing between two levels, and Jev can't pick one. What do you try first?
The model can only distinguish rungs it can picture. Describe each as a situation, and add an example to the rung people confuse. More rungs with the same vagueness spread the mass thinner.
Read Priya's message aloud. Ask: prioritise her, yes or no? Hands up. Then show the four narrow answers and the one line of code that decides.
Stage 6 · One judgment per question

Split the big question into small ones

"Should we prioritise this customer?" hides four judgments. Ask them separately and let code combine them. Each input becomes inspectable, and you can change the rule without touching the questions.

💬 Explain it to me

"Is this a good phone?" is a hard question.

"Is the battery good? Is the camera good? Is it cheap?" are easy ones. Then you decide what matters most.

Here, Priya buys from a shop every month. She's unhappy and says she's looking at other shops.

Your to-do list
  1. Press Run both.
  2. Read the one line of code on the right, then answer the question at the bottom.
✅ Ask small questions. Let your own rule turn them into one decision.
Press Run both. Left: the broad question. Right: four narrow ones over the same message.
One broad question
Four narrow questions
Check
Next month your boss says: "customers who might leave matter more than big spenders now". What do you change?
The four questions did not change meaning, so their answers are still valid. Only the policy changed, and the policy lives in code. That is the whole reason to split.
Hand a volunteer the sliders. Every move re-sorts the inbox. Keep pointing at the counter: model calls stays at one.
Stage 7 · Code decides

Ask once, keep the numbers, move the sliders

Jev answers three Nouls for eight messages in one call. Your thresholds turn the numbers into lanes. Move a slider and the inbox re-sorts with no new model call.

💬 Explain it to me

A teacher marks everyone's test once.

Then the school moves the pass mark from 50 to 60. Nobody re-marks a single test: people just move above or below the line.

Jev does the marking. The sliders are the pass mark.

Your to-do list
  1. Press Score the inbox.
  2. Drag the sliders a few times and watch messages jump between the boxes.
  3. Look at model calls at the bottom, then answer the question.
✅ Jev was asked once. Every slider move after that was free.
Press Score the inbox once. Then drag the sliders and watch the counters at the bottom.
Human queue if P(refund) ≥0.60
Manager if P(angry) ≥0.70
…or if P(urgent) ≥0.85

Auto-reply 0

Human queue 0

Manager 0

0
model calls
0
slider moves
0
extra calls caused
Check
You moved three sliders. How many new model calls did that cost?
The numbers are already yours. Thresholds are code. Compare that with a prompt that says "reply with a category": every policy change is a prompt change, and a prompt change means re-testing everything.
Take a vote on each prediction before running it. People trust a tool more once they have watched it fail and seen the fix.
Stage 8 · Where it breaks

Predict, then run

These four come from TypeSafe's own published limits. For each: guess what Jev will say, press run, and read the fix. Trusting a tool means knowing its edges.

💬 Explain it to me

A calculator is brilliant at maths and useless at spelling.

Jev is brilliant at understanding words, and bad at counting or comparing dates.

Knowing what a tool is bad at is part of using it well.

Your to-do list
  1. On at least two cards: pick a guess, then press Run.
  2. Read the fix under each one, then answer the question at the bottom.
✅ Let Jev read the words. Let your code do the counting.
Check
You need to know whether a list has more than five items that are fruit. Best approach?
Counting and arithmetic belong to code. Jev is good at the small semantic judgment ("is this a fruit?") and bad at tallying. Split the work along that line.
Ask a Sinhala speaker and a Tamil speaker to read their line aloud before you run it. Then compare the three columns together.
Stage 9 · Our languages

The same three questions in Sinhala and Tamil

Micro LLM's tokenizer cannot read these scripts at all. Whether Jev can is something we find out on screen, not on a slide.

💬 Explain it to me

Sri Lanka writes in three languages.

Here is the same annoyed message in English, Sinhala and Tamil.

Does Jev understand all three?

Your to-do list
  1. Press the green button.
  2. Compare the three columns, then answer the question at the bottom.
✅ The question stays in English. The message can be in any language the model reads.
Run all three. Compare the numbers across the languages.
Check
A customer writes in Sinhala. What do you change in your questions?
If the numbers above matched, the model read the state in its own script. Check on your own messages before trusting it, which is exactly what Build mode is for.
Pairs work best. First pair to 6 / 6 explains their rules to the room. If nobody gets there in five minutes, give the hint: money back, anger, time pressure.
Stage 10 · Prove it

Build a triage that sorts six messages correctly

Six messages, each with the lane it should land in. Choose your questions, set your thresholds, press check. Six out of six earns the certificate.

💬 Explain it to me

Final level! You're running a shop's inbox.

Six messages. Each one should go to the right place: an auto-reply, the human queue (a person answers it), or the manager.

Pick your questions, make your rules, and get all six right.

Your to-do list
  1. Tap the questions you think you need from the list.
  2. Make a rule for each one: If this question is above a number, send it somewhere.
  3. Press Check my triage.
  4. Get 6 out of 6. Stuck? Press the hint button.
✅ Six out of six. You chose the questions, Jev gave the numbers, your rules made the decisions. That's how real apps use it.
Pick questions from the library (or write your own), then set which lane each rule sends a message to. The rule is: manager beats human, human beats auto.
Leave the sheet on screen for questions, or print it. Point everyone at Build mode for Monday: their own inbox, their own questions.
You made it

One page to remember

Everything on this page fits on one sheet. Print it, or email yourself the code for the triage you built.

💬 Explain it to me

You made it! 🎉

Everything you learned fits on the sheet below.

Want to try it for real? Press Use my own messages and paste a few messages you write yourself.

Jev decides, it does not writeState in, typed answer and probability out. No prompt. No paragraph.
Three shapesNoul yes/no probability · Choice pick one bin · Score rungs you describe.
Calibrated means honestOf everything it calls 0.7, about 7 in 10 are true. 0.5 means "can't tell", not "medium".
Words are the modelDescribe every rung and option as a concrete situation. Add an example on the confusing one.
One judgment per questionSplit the broad question. Ask the narrow ones together in one call. Combine in code.
Code owns the policyThresholds, weights, counting, dates, side effects: all yours. Changing a rule costs zero model calls.
Know the edgesNo counting, no date maths, negations read literally, injected text can steer. Keep untrusted text labelled as data.
The whole APIPOST api.typesafe.ai/v1/systemone with {state, questions}. Key stays on a server.
The certificate unlocks when the six-message triage in stage 10 passes.
Take the code

Your questions, ready to run

Ready to run

          
Keys come from console.typesafe.ai/keys. Keep the key on a server, never in a web page.
One email with the code and a link back here.
  • Jev is an API call like any other: API Lab shows what that means on the wire.
  • A decision like this gates a workflow: build the surrounding flow in Automation Lab.
  • The chatbot half of stage 1 is what Micro LLM builds from scratch.
Calibrated means: of everything it calls 0.7, about seven in ten turn out true. Today the room finds out whether Jev is, whether a chatbot's "confidence: 0.9" is, and whether they are.
The duel

You vs Jev vs a chatbot

Same question every time: does the sender explicitly ask for their money back? Slide your probability first, then reveal. Lower Brier score wins (0 is perfect, 0.25 is a coin flip).

Now the room is the state. Everyone sends a real message from their own inbox, and you ask Jev about all of them at once, live, on the screen.
Live collection

Open a room

Phones join with a four-letter code. Nothing is stored past 24 hours and no names are required.

The questions from stage 4. Change them there, run them here.

Inbox

Open a room and the messages will land here.
Duel, with the whole room

Everyone slides, then reveal

Push a calibration message to every phone. They slide, you reveal Jev's number, and the room average joins the duel scoreboard.

Open a room first 0 votes in
Bring your own inbox. Paste real messages, pick the questions that matter to your job, and leave with a table, a CSV and the code that produced it.
1 · Your messages

Paste them, one per line

0 messages
2 · Your questions

Tap the ones that matter, or write your own

cocoon.

Room

Send a real message you have had to deal with at work: a complaint, a request, an email that made you sigh. No names in it, please.

Read more

The Jev series

Plain-language guides to Jev, System One models and what we learned building this lab. Facts about the model come from TypeSafe and its documentation; the experiments are our own.

Copied ✓