Does my business need a full AI chatbot model or a small decision model?
Last updated 2026-10-02
If the AI only has to pick a label, answer yes or no, or give a score, a small decision model is usually enough and far cheaper and faster than a chatbot model. Keep a full language model for writing, conversation, and reasoning. Decide by measuring accuracy on a few hundred of your own labelled examples.
Many businesses paying for a chatbot model only need a decision model. Routing a support ticket, flagging a refund request, spotting spam, or deciding whether two customer records are the same person are all classification: the software needs a label, a yes or no, or a score, not a paragraph. In September 2026 a startup called TypeSafe launched a model built only for those answers, and open-source copies appeared within two weeks.
- ChoiceWhich team handles this ticket: billing, technical, or account?
- Yes or noDoes this message ask for a refund?
- ScoreHow frustrated is this customer, from 0 to 2?
What is a small decision model like Jev?
A decision model reads text or structured data and returns a typed answer: one label from a fixed list, a yes or no with a probability, or a number on a scale. TypeSafe's Jev, launched on September 15, 2026, returns those answers with a confidence figure in 70 to 500 milliseconds and cannot produce free text.
TypeSafe calls these "System One" models, after the fast, intuitive thinking described by the psychologist Daniel Kahneman. Its documentation names three question types: Choice, Score, and a yes or no type.
The launch thread on Hacker News drew 520 comments. One developer wrote that they already used language models in production exactly this way, narrowing every call to a small structured choice.
Which business AI tasks are really classification tasks?
Most back-office AI tasks are classification: sorting inbound email, routing support tickets, flagging refund or cancellation requests, filtering spam, scoring leads, and matching duplicate customer or supplier records. Each one ends in a label, a yes or no, or a number. Writing replies, summarizing calls, and answering open questions are not classification and still need a language model.
| Task | Answer shape | Decision model enough? |
|---|---|---|
| Route a support ticket to a team | Choice from a fixed list | Usually |
| Flag a refund or churn risk | Yes or no, with probability | Usually |
| Filter spam or phishing | Choice | Usually |
| Match duplicate customer records | Yes or no | Often, with rules for edge cases |
| Draft the reply to the customer | Free text | No, use a language model |
A commenter in the thread on whether OpenAI will copy Jev put it bluntly: many tasks now done with generative models are "classification in disguise."
How accurate are the cheap open-source Jev copies?
The open-source copies of Jev are close on their own benchmarks and less reliable on real customer data. Jeff, a free model trained in about two hours on one workstation GPU, reports 79.1% on five benchmarks against 83.0% for Jev. One developer testing it on their own classification work measured 70% against 94% for Jev.
The Jeff project page lists answers in 22 to 28 milliseconds on a single GPU for its smallest model, and a larger 2B version at 83.1% on the same benchmarks. The 70% figure comes from a commenter in the Hacker News thread on Jeff, who called it unacceptable for classification.
Ollaya, a free tool that runs Jev-style models on your own machine, publishes its own comparison: one of its models at 0.722 accuracy in 89 milliseconds against 0.738 for Jev's API. In the Ollaya thread, one user reported that its Laya model performed significantly worse than Jev on complex queries. A builder's own benchmark is a starting point; your data is the test.
What does a decision model cost versus a chatbot?
A decision model costs a fraction of a chatbot model per request. TypeSafe prices Jev at $0.042 per million input tokens with output free. OpenAI's cheapest current model, GPT-6 Luna, costs $0.10 per million input tokens and $0.50 per million output. A free open-source model on your own hardware costs only the server.
At small volumes those differences are pocket change. The gap starts to matter at hundreds of thousands of decisions a month, and speed matters sooner: a 30-millisecond answer can sit inside a checkout or a live chat, where a multi-second model call cannot. Prices are from TypeSafe's launch post and OpenAI's API pricing page.
How should an owner test a decision model first?
Test a decision model on a few hundred of your own past cases whose correct answers you already know. Run the same cases through the decision model and through the chatbot model you use now, then compare accuracy, cost per thousand decisions, and response time. Pick the cheapest option that meets the accuracy the business actually needs.
- List the tasksMark which ones end in a label, yes or no, or score
- Label 200 to 500 casesPast tickets or records with the answer staff gave
- Run both modelsSame cases, decision model and chatbot model
- Compare three numbersAccuracy, cost per 1,000 decisions, response time
- Set a thresholdLow-confidence cases go to a person
Decide the acceptable error rate before looking at results. A spam filter that misses 5% of junk is fine. A refund flag that misses 5% of angry customers may not be. The confidence score is what makes a decision model safe to use: act automatically above a threshold and route the rest to staff.
What does the Jev copy wave teach about AI moats?
An AI feature without proprietary data, distribution, or service behind it can be copied in weeks. Jev launched on September 15, 2026. Ollaya was posted to Hacker News on September 25 and Jeff on September 28. Owners should expect the same speed of copying for any AI feature they pay a vendor to build.
A commenter in the Ollaya thread asked what the end game is for a startup copied by open source in about two weeks. For a buyer, the lesson cuts two ways. Do not pay a premium for a feature a free model will match next quarter. And if AI is part of what you sell, the lasting value is your labelled data and your customers, not the model. For an independent read on which parts of your plan need a language model, a second opinion before the build costs less than the build.
Related guides
- Before hiring an AI consultancy, find the part that actually needs AI
- Where should AI actually fit in my product?
- When is an AI API wrapper enough, and when do you need more?
- Build custom software or buy off the shelf?
- TypeSafe: System One models documentation
- Jeff: open-source Jev-compatible decision models and benchmarks
Key takeaways
- Ticket routing, refund flags, spam filtering, and record matching are classification and rarely need a chatbot model.
- Decision models like TypeSafe's Jev return a label, yes or no, or score with a confidence figure in milliseconds.
- Open-source copies score close on their own benchmarks, but one developer measured 70% against Jev's 94% on real work.
- Test any model on 200 to 500 of your own labelled cases before trusting a published benchmark.
- AI features without proprietary data behind them were copied in about two weeks, so do not pay a premium for one.
Frequently asked questions
What is the difference between a decision model and a chatbot model?
A chatbot model, or large language model, writes free text: replies, summaries, code. A decision model returns only a typed answer, such as one label from a list, a yes or no probability, or a score. Decision models are faster and cheaper, and they cannot write a reply to your customer.
Can a small decision model replace ChatGPT in my business?
Only for the tasks that end in a label, a yes or no, or a score. Sorting, flagging, and matching are good candidates. Drafting, summarizing, and answering open questions still need a language model. Many businesses end up with both: a decision model for routing and a language model for writing.
How many examples do I need to test a decision model?
Two hundred to five hundred past cases with known correct answers is enough for a first comparison for most small and midsize businesses. Run the same set through every candidate model so accuracy, cost, and speed are compared on identical work.
Are free open-source decision models good enough for business use?
Sometimes. Projects such as Jeff and Ollaya run on your own hardware at no licence cost and publish benchmark scores near Jev's. Users on Hacker News reported lower accuracy on their own data. Measure on your cases, and keep a person reviewing low-confidence answers until the numbers hold.
About the author
Giacomo Balli is an independent technology advisor in San Francisco. He has built software and mobile apps since 2010, runs a portfolio of more than forty live apps of his own, and reviews software, AI, and vendor decisions for owners before they commit the money.
Disclosure
Giacomo Balli sells fixed-fee independent reviews of technology decisions, including AI initiatives. He does not build or resell software and takes no referral fees. No company or project named on this page paid to be mentioned.