Maia Talks About AIMaia Talks About AI

System One Models: What Jev and Laya can do that Claude can't

2026-09-28

By Maia Salti

Illustrated grid of dashboard cards with sliders, toggles, charts and branching choices Generated by ChatGPT

TL;DR: What Jev and Laya have that frontier LLMs don't: speed and low cost. AKA the commoditization of AI classification.


It’s been a while since my last post; I’ve fortunately been busy with my internship at an AI consultancy in Singapore and my remote part-time work at an AI startup in New York, both of which I’ve been thoroughly enjoying.

Just by being around people who are so AI-curious and hearing great discussions around the office, I’ve learned more from showing up to work everyday than all of the articles I’ve read and podcasts I’ve listened to combined. Just another emphasis on the importance of showing up and surrounding yourself with cool, curious people (shoutout to my boss Ned Lowe, go follow him on LinkedIn: he has the greatest takes on all things AI-related. Also shoutout my desk neighbour Sunny Singh, who imparts his CTO wisdom on me. P.S. Sunny, I will be taking you up on your offer to invest in my future startup).

Rent a Crane, Hang a Picture

Before Jev, we used LLMs capable of multi-step reasoning and long mathematical problem solving to classify or label data, score things, and provide confidence scores on their answers. The two main issues with this:

  1. It's bloody expensive.
  2. You can't see where their scores are coming from, and they're not generally reliable values.

It's literally like renting a crane to hang a picture up on your wall.

Then boom. On September 15th, a San Francisco lab called TypeSafe AI came out of stealth with a $40M seed round and a model called Jev. Its launch post sat at the top of Hacker News all day with almost 2,000 points.

Then BOOM (slightly bigger than the first one). Three days later, Convai Innovations put an open-weights alternative called Laya on Hugging Face under an Apache 2.0 licence, and people already have it running inside a browser tab.

Imagine being in stealth for two years and releasing one of your most beautiful creations, just for a startup in India to one-up you with a faster, free version of your product (yay open-source though).

Both of them are what TypeSafe calls System One models. I find them really interesting because they go the opposite way to almost everything else in AI right now. Every frontier lab is racing to build models that think for longer and write more, and these two can't write a single sentence.

This relates to one of my biggest beliefs regarding having an edge in today’s world of constant creation and innovation: perfecting the implementation of a very narrow functionality or workflow in a specific domain. I think you must, or else you’ll be competing with generalists with a lot in the bank. Like Google and Apple.

Thinking, Fast and Slow

The “System One” term comes from Daniel Kahneman’s book Thinking, Fast and Slow, where he categorises human thinking into two modes. System 1 is the fast, automatic one you use to recognise a friend's face or to pick out your name in a noisy restaurant. System 2 is the slow, effortful one you use for long division or decrypting an ambiguous text.

LLMs, especially the long-chain reasoning models, have been getting much better at System 2 thinking. They write out a long chain of thought and check their work before answering. TypeSafe's argument is that a huge amount of what software actually needs is System 1: quick judgement calls like whether an email is spam or which team should handle a support ticket.

The name Jev is a reference to William Stanley Jevons, the economist behind the Jevons paradox. When coal got cheaper to burn, people ended up using more of it. TypeSafe expects the same thing to happen with intelligence, so every time it gets an order of magnitude cheaper, a whole new batch of uses opens up.

TypeSafe's CEO is Diogo Almeida, who at OpenAI, in his own words, "helped build the methods that made language models useful at following instructions." That work became the research behind ChatGPT, which I think makes it quite telling that his next project is a model that doesn't chat at all.

What a System One Model Is

You give it a state, which is whatever you want it to look at: an email or a description of a game screen. Then you ask it typed questions about that state.

The simplest question is a yes or no, which TypeSafe calls a noul. You can also ask it to make a choice between up to 255 options, or give a score on a scale you write yourself.

Here's a real example (from Laya's README, which uses the same format as Jev):

// request
{"state": {"body": "billed twice, refund please or we cancel"},
 "questions": {"dept": {"type": "choice", "instructions": "which team?",
                        "criteria": {"billing": "refunds", "tech": "bugs"}}}}

// response
{"answers": {"dept": {"choice": "billing", "confidence": 0.94,
                      "probabilities": {"billing": 0.94, "tech": 0.06}}}}

Almeida compares a noul question to an if statement, and I think that's the best way to picture it: an if statement you write in plain English that can read messy text.

How They Differ From LLMs

An LLM writes its answer one token at a time (I explain this in my LLM post). A System One model skips the writing and returns all of its answers at once, in a single pass.

Let’s use a high-school exam analogy. The LLM is the free-spirited creative student, great for an English exam and writing long paragraphs. The System One model is the systematic statistician, filling in multiple choice exams and telling you exactly how confident it was on each answer.

Those probabilities are the important part. When you ask an LLM how confident it is, it writes a number, and that number is just more generated text. Jev was trained so its probabilities are calibrated, meaning when it says 90% it's right about 90% of the time.

It's also very fast and cheap. Jev costs $0.042 per million input tokens (output is free) and answers in 70 to 500 milliseconds. TypeSafe says it came out 193.6x faster and 444.6x cheaper than frontier LLMs, but on workflows their own team wrote, so I'd take that with a pinch of salt. In an independent sentiment test, Jev scored 95.4% against Claude Sonnet 5's 95.6%, for under 2% of the price.

The trade-off is that it can't write anything at all, and TypeSafe's own docs say it's unreliable at maths and at comparing dates.

What They Unlock

A model this fast and cheap changes where you can actually use AI.

  • Smart if statements. Routing tickets, filtering spam, moderating comments or flagging a risky payment on every single event, without waiting seconds for an LLM.
  • Real-time loops. At around 100 ms a call, AI can make decisions several times a second, which is fast enough to play a game.
  • Huge datasets. One early tester asked 11 questions about each of 24 documents, and it cost half a cent.
  • Knowing when to ask a human. Because the probabilities are calibrated, you can let the model act when it's confident and hand everything else to a person.
  • Running on your own device. Laya's weights are open, so it can run on a laptop or in a browser tab without your data leaving your computer, and you can fine-tune it on your own data.

Distracted boyfriend meme: a human with classification needs looking at Jev and Laya while Claude, ChatGPT and Gemini look on

For the ML Engineers

When Jev launched, a lot of Hacker News comments asked how this is any different from zero-shot classification, which has been around for years.

A classic zero-shot classifier like bart-large-mnli turns each label into a sentence ("This text is about sports") and checks whether your text agrees with it. That takes a separate forward pass for every label, and the model only ever sees the label name.

Laya's architecture is public, so it's the easiest to compare. It's a ModernBERT encoder with a small decision head. All the options go into a single input with their own [MASK] token, so they're scored together in one pass. Each option can also carry a description of what it means.

The bigger difference is training. Laya is trained with reinforcement learning, where the reward is a proper scoring rule. That's a reward the model can only maximise by reporting its honest probabilities.

An independent benchmark compared Jev with the classic zero-shot approaches:

Jev vs Classic Zero-Shot Classifiers

Accuracy with no training examples (PAWS is AUC). The best score in each row is highlighted.
Taskbart-large-mnliDeBERTa-v3 zero-shotbge-m3 cosineJev
AG News0.6770.7630.7770.865
SST-20.9140.9130.8640.960
Banking770.4280.5790.7220.712
TweetEval emotion0.7490.7600.6500.827
PAWS (AUC)0.7300.9080.6650.936
arXiv, Sept 20260.5500.5890.5540.891
Source: zhuyansen/jev-zeroshot-vs-bert on GitHub. The arXiv row uses papers published after Jev was trained.

Jev wins five of the six tasks. The last row uses papers published after Jev was trained, so it can't have memorised them, and that's where its lead is biggest.

Laya is slightly different. Its base model scores barely above random guessing on the typed-decisions benchmark (0.362, where random gets 0.318), and its model card says it's meant to be fine-tuned on your own task. The fine-tuned version does beat Jev on that benchmark (0.766 vs 0.727).

What I Built With Laya

I keep a Google Sheet of articles/links I want to come back to, sorted by category. Every time I found something I'd have to open the sheet and paste the link in by hand, which got annoying relatively quickly.

So I built a Chrome extension called Resource Collector. You click the icon, click a category, and the page lands in a new row of the sheet.

I realised that eventually, once I have a lot of categories, I won't want to sift through them looking for the right one, and I thought Laya would be an optimal addition. So I added her in. I had to download the model onto my laptop (about 1 GB), then had Claude Code set it up for me, and now it works like this:

Resource Collector suggesting Finance for a Wall Street Journal article Laya is 86% sure a Wall Street Journal markets article belongs in Finance

Laya runs locally on my Mac through a small Python helper that uses the same request format as Jev, so I can switch to Jev in the settings too. Laya answers in about 20 milliseconds (the flap of a hummingbird’s wing btw), and nothing leaves my laptop. Once I have enough articles with labels, I’ll try fine tuning Laya.

I did have some RAM issues, the model takes up about 1 GB of memory every time it runs, and its cache kept growing to around 600 MB on its own until I capped it. I ended up making the helper unload the model after ten idle minutes.

Cool Things People Have Built

Jev and Laya have been out for less than two weeks, and people have already built a lot with them.

Jev plays Doom

Jev playing Doom, with its live judgments next to the game TypeSafe's launch demo

Jev reads a text description of the game and answers questions like "should the player's trigger be held down right now?" about ten times a second.

Jev flies to the Moon

Odyssey, where Jev acts as flight director for a rocket launch Odyssey

Jev is the flight director, from liftoff all the way to lunar orbit. Every decision steers the spacecraft, and the bars show its probability for each option.

jev-mice

jev-mice, a simulation of mice, cats, traps and food jev-mice

It's a browser simulation of mice, cats, traps and food where Jev decides every animal's next move. You can switch to fixed rules and compare the two.

Laya plays Flappy Bird

Laya playing Flappy Bird, showing the state, question and answer probabilities brain function collapse

Laya asks where the bird is relative to the gap and flaps when "below" wins. It makes about 30 decisions a second on a laptop.

Laya in your browser

Laya running in a browser tab with ONNX Runtime Web Laya in your browser by Vishal Mysore

The whole model runs inside a browser tab, so nothing you type leaves the page.

Open copies of Jev also showed up within days, including Nokia's AnyJev, Jared Palmer's Kev, Bespoke Labs' Nimble and CUA-S1 for computer use.

I think the coolest part is that most people will never knowingly use a System One model. Similar to how most people don’t know how their car works, we trust the engineers of the world (previously mechanical and now software) to create models like Jev and Laya that will be hidden inside apps making thousands of small decisions for us every second.


If you work on classifiers or calibration and think I got something wrong, please email me at maia.salti@gmail.com. I'd love to learn as much as I can.


Sources

Data research run by Claude Code

Get new posts in your inbox
I'll send you a short email when a new post is up. No spam, just the link.