A chatbot answers by writing, one small piece of text at a time. A new kind of model skips the writing. It reads a situation and returns an answer, such as a yes or no with a confidence number, in one pass through the model. On October 7, Liquid AI released two open-weight versions. Around the same time, AutoTrust AI published a smaller, faster build of its own. The idea is real and useful. But most of the numbers still come from the makers.
How a Model Decides Without Writing
A normal language model writes an answer token by token, and every token needs another pass through the model. That takes time. Liquid AI’s d1 models work differently. You give them a state, such as a customer message or a photo, plus a list of questions. They return typed answers in one pass, with zero output tokens. The question types are yes or no, pick one option, or rate on a scale.
The answers come with probabilities. A “calibrated” model is one whose 70% really means right about 70% of the time. That is Liquid’s claim, and it has not been independently tested here.
d1-3B has about 3.1 billion parameters and reads text and images. It is built from Liquid’s own vision-language model. The smaller d1-omni-600M handles text with images, or text with audio, and Liquid calls it experimental. Liquid says plainly that these are not chat models. It suggests uses such as routing, moderation, classification and guardrails for AI agents.
What Liquid AI Claims and What Is Checked
Liquid reports that d1-3B scores 48.57 on the Decision Index v0.2.1, a public leaderboard hosted on Hugging Face. The figure comes from the d1-3B model card and Liquid’s blog. Liquid says this is ahead of every model under 10 billion parameters. Two cautions apply. Liquid scored its own models with the official scorer instead of submitting them to the leaderboard. And the comparison covers only the models Liquid chose, so a larger model may still score higher. The 600M model scores 15.95 on the same index, according to its model card.
Speed is the main selling point. Liquid reports 8 milliseconds for one question on an RTX 4090 graphics card, 16 ms on a Jetson AGX Thor and 50 ms on the small Jetson Orin Nano. Read the fine print. The 8 ms figure uses a compile setting, and without it Liquid’s own card lists 16 ms. These are warm runs, one request at a time. On the Thor, three questions over one input took 20 ms against 16 ms for one. No independent replication was found in this research.
Licensing also needs a careful read. Liquid’s blog says users can deploy the models without restrictions. The LFM Open License v1.0 says free commercial use ends once a company’s annual revenue passes $10 million. It is open-weight, but not open without limits.
A Second Release: AutoTrust’s GEV-26B-Decide
AutoTrust AI, a Singapore lab, released a 4-bit NVFP4 build of its GEV-26B-Decide model. The weights shrink from 51.1 GiB to 17.1 GiB of GPU memory, a cut of 66%, and the card says one GPU with 24 GB or more can run it. The model sits on Google’s Gemma 4. A fast first-pass answer comes first, and harder cases go to slower step-by-step reasoning.
The headline claim is 95% success on 60 browser tasks at about 85 ms per click. The tasks were random, multi-step jobs such as shopping, settings and mail, each needing 3 to 7 clicks, in a test browser. It is not one 60-step task. The model card also shows the limits. The 95% needs the text of each page element, like an accessibility tree. With only numbered boxes on the screenshot, the same model scored 15%. In a robot-arm test it finished 40% of scenes, against 75% for a larger sibling. These are the authors’ own results.
Why It Matters and What Comes Next
AI agents make many small choices: is this safe, which tool, which route. A chat model that writes a paragraph for each one adds delay and cost. A decision model gives a probability that ordinary code can act on, such as escalating when confidence is low. Many projects are now trying this, including Decider, OpenJev and NanoJev.
In the near term, watch for independent leaderboard entries and real deployments, which will test whether the confidence numbers and speeds hold up. The long-term idea, that decision models could replace chat models inside agents, is a prediction. One limit is already clear. A model that returns only a number can be confidently wrong, and it explains nothing.