Clef: Open-source decision models, and new RL fine-tuning platform (blog.cloudflare.com)

348 points by jasondavies 5 hours ago

manlymuppet 4 hours ago

Am I hearing this right, that they made a decision model based on Typesafe's new paradigm, and actually made a model better than Jev based on Typesafe's own ranking?

And it's only been a few weeks.

slopnt 4 hours ago

They have to have decision models already in production. Part of their business is detecting bots, DDoSers and spammers.

alightsoul 2 hours ago

Yeah that's a decision tree, random Forest or some other machine learning classifier. They have existed for a long time

smallmancontrov 2 hours ago

TeMPOraL 4 hours ago

It's not a "new paradigm", it's a low-hanging fruit that's been lying around for years; Typesafe were the first to bother to stop and pick it up, and market the shit out of it. But it was still a low-hanging fruit.

There are many, many of those left around, because AI frontier is moving forward so fast, everyone is racing ahead. Which is why I laugh when people say AI is not transformative and LLMs are a dead end (and my favorite, "what are we going to do with all those GPUs when the bubble pops?"). Even if SOTA LLMs hit a hard capability limit tomorrow and never advanced again, there's a good decade of growth and advancement to be extracted just from all the low-hanging fruits that were left unpicked along the way.

seizethecheese 4 hours ago

Name a few of these low hanging fruit left around.

TeMPOraL 4 hours ago

sarkarghya 2 hours ago

hobofan 4 hours ago

murkt 4 hours ago

CamperBob2 2 hours ago

vulture916 2 hours ago

Jev = $0.042/m input, output free Clef = $0.24/m input, no output price listed

At 300 tokens per call, you'd get:

One million decisions on Jev cost about $12.60. One million decisions on Clef cost about $72.

Would probably make sense to self-host Clef, if you have the capability/resources. If not...

scronkfinkle an hour ago

> no output price listed

It's weird to think of these kinds of models as having "output tokens". Cross-encoder approaches like Laya add a [MASK] marker per option, but nothing is generated the way an autoregressive transformer generates. It's one bidirectional pass over your input, then a small head scores each option, so you wouldn't really pay for output as much as only input

jampekka an hour ago

Output for decisions have so few output items (not really tokens here) that they are negligible anyway. Jev hyping "free output" is almost lying by omission.

buildbuildbuild 4 hours ago

Open weights, not open source.

The weights have permissive licensing, but the data and training pipeline are not published to reproduce them from their proprietary Qwen starting points. Weights are not "source."

jMyles 4 hours ago

Came directly to comments hoping not to see this one.

<sad trombone sound>

Surely someone will soon do what the title of this post makes it seem like cloudfare did. Truly modular open source training and inference logic, along with a totally open corpus and weights, will eventually out-compete the closed ecosystem.

ainch an hour ago

There are some groups doing it for LLMs - like the Allen Institute for AI's Olmo models and Eluether AI's Pythia.

ssiddharth 5 hours ago

Pricing is $0.24/million input tokens which is ~6x compared to Jev. Clef-flash is at $0.09 which is way more competitive.

CBLT 3 hours ago

Yeah I also thought it was strange their pareto frontier didn't include cost.

bityard 4 hours ago

Clef is based on Qwen3.8-27B and Clef-flash is based on Qwen3.8-9B (edit: actually Qwen3.5-9B). So, similar in spirit to Kev by my understanding, but based on a newer model.

NitpickLawyer 3 hours ago

> and Clef-flash is based on Qwen3.8-9B

There is no official qwen 3.8 9b

From the model card:

> Clef-Flash is post-trained from Qwen/Qwen3.5-9B. See Clef for the larger variant.

bityard 3 hours ago

Thanks, I missed that. Fixed my comment.

ddarolfi 3 hours ago

It's based on Qwen3.5-9B, maybe a typo

okpatil 2 hours ago

Atom is 60M Param (around 133x to 400x smaller).

16ms latency. And locally run.

https://at0m.pienomial.com/

Why go big when you can go small ?

kamranjon 2 hours ago

Cause it's not open?

okpatil 2 hours ago

mrkn1 3 hours ago

For smaller scale decision model that runs on CPU, check https://news.ycombinator.com/item?id=49923223

fastball 2 hours ago

tbh saying all these dumb decision models are similar to Jev is like saying markov chains weren't far from GPT-2.

The value isn't really in the I/O shape, it is in the intelligence combined with the output shape. Every extra ounce of intelligence in these models unlocks additional use-cases. But the converse is also true: a dumb decision model is going to be less useful than using a more intelligent standard LLM.

That is the appeal of Jev: for certain usage it has more intelligence than some small SOTA LLMs. It is the first decision model that actually feels intelligent (to me).

amluto 2 hours ago

I’ll go out on a limb and suggest that I don’t think a Jev-like model is particularly useful unless you can fine tune it. The Jev API has zero ability to pass in a prior [0], and, if you can neither pass in a prior nor fine tune for your system, you will get an output that may be almost meaningless.

I’d love to see someone build a model of this sort that can actually accept priors and do something intelligent with them.

[0] You can feed Jev a prior as text. I’ve tried it. It works poorly.

mikeocool 2 hours ago

It seems like jev's major advantage over existing classifiers is that I dont have train it.

If I have to gather and tag data to fine-tune Jev, I can probably just train an "old school" classifier model and make it even cheaper, faster, and just as accurate.

sheepscreek 2 hours ago

Also one of the more interesting features of Jev is the confidence rating that hardly any Jev-cc talks about.

amluto 2 hours ago

It seems interesting to me only in the sense of being useless. From the horse’s mouth:

> Confidence is derived from the probabilities

https://docs.typesafe.ai/confidence

(Why is it much easier to find AI-slop websites quoting this than it is to find the actual documentation?)

My inner Bayesian would like for Jev to provide something resembling “evidence”, although I admit that one might ask Jev questions that are somewhat awkward to treat as typical Bayesian questions. If I ask “will this PR be merged”, it’s kind of strange to contemplate the probability of a PR conditioned in that PR being merged in the future. But I bet there is a way to formalize a prior-free classifier in a way that makes Bayesians and non-Bayesians happy, possibly involving actual learned probabilities and confidence levels. If you read the literature on scoring rules, you will find that classifier scores do somewhat naturally decompose into a few interpretable terms.

brokensegue 2 hours ago

I think better than priors would be a closed loop where you tell it what the right answer was (or some signal) and they monitor and fine-tune for you

okpatil 2 hours ago

We were able to completely automate 20,100 token prompts with At0m[https://at0m.pienomial.com/].

We believe entire compliance workflows (even multilingual) could be automated.

Would you like to get a demo ?

sheepscreek 2 hours ago

You’re coming on a bit strongly - a couple of comments with a link is sufficient. Before trying to sell, try to genuinely further the conversation, provide some useful knowledge in return for the reader’s attention.

okpatil 2 hours ago

fooker 3 hours ago

This is awesome.

I bet the competition will result in research into how to make these decision models several more orders of magnitude faster and cheaper.

Here's a challenge problem - look at a 1M context window and produce N decisions (different queries) from it in 50-100ms.

yipinwong 4 hours ago

A question someone not trainined in AI/ML field, Is a decision model that easy to crete that there are floods of these JEV alternatives already?

Or are companies/people already building this based on say an arXiv docs? n

---

The pricing is ... hm more expensive but not at the point I won't give it a try due to the embeded vision encoding

TeMPOraL 4 hours ago

Yes, it's easy. The thing people are missing (especially those believing AI is a "dead end" and "not transformative") is that the field has been advancing so fast in the past few years, that there's lots of such unexplored avenues, unpicked low-hanging fruits, that everyone just raced past. We've barely begun exploring the capabilities ML brought us - patterns, applications, and architectures.

Now that we're hitting against the hardware supply limits of global economy, I expect more people to go back and revisit the things left along the way in the mad rush to "just throw more compute at it / make a bigger model" - and thus many more cases like Jev to show up in the next few years.

nico 4 hours ago

The basics are pretty simple. And depending on what your specific need is, the model can be really really basic, fast and super effective (ie. run on a mobile device and process thousands of requests in <100ms)

I've been playing with this for the last year or so. Started with a personal email classifier, also did benchmarks with some public datasets, then created a couple classifiers that could play Doom, and now I've been trying out some other experiments, like a request proxy/router to automatically choose a classifier and fallback to LLM to handle unseen requests

Jev did a great job at creating hype, but also at shaping the concept and space of "decision engine" or "decision model". People were already doing this with LLMs, which is very inefficient for most tasks like that, and the Jev guys figured there was a market there. It seems like they were right, and now there's a rush to flood the space, taking advantage of the hype window

calebkaiser 4 hours ago

There is a bunch of stuff to tease apart.

In general, training a general purpose classifier is something lots of people have worked on for a long time. Large Transformer models themselves are typically "generalists" already, so structured generation and constrained decoding have given you the ability to use an LLM as a general classifier for years. It's an incredibly common pattern for working with LLM judges or any sort of branched decision making workflow.

A lot of people who are a bit less familiar with the field saw the hype around Jev and presumed that the reason it was so exciting was that it was a fundamentally new interface for working with an LLM. And that additional excitement drove even more attention to Jev. But fundamentally, TypeSafe's announcement was that they found a particular architecture/training paradigm that resulted in a model for this particular interface that had incredible accuracy, very low latency, and for which they could offer inference at a super low cost.

I've not kept up with the flood of Jev clones that have been released, but I think this is just typical for any new component in deep learning that gets popular. There are an absurd number of open source autoregressive LLMs and fine tunes you can use. The thing that makes one more popular than the other is typically the general performance of the individual model.

But training a model for this purpose, or emulating the procedures described in Jev's papers, isn't something that would be beyond the capabilities of any lab. It's not an entirely alien architecture or approach.

The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

tomrod an hour ago

> The bigger question for TypeSafe as a company would be if other teams are producing Jev-like models that win on performance or cost. Like I said, I haven't followed the reports super closely, so no idea if that's the case or not.

If CF's benchmark is representative and sufficient, Clef outperforms Jev!

Models by themselves don't guarantee market capture. Rather, its how they integrate. I think a lot of folks are burned by the closed nature of many models.

conmod278 4 hours ago

Live coding Jev from Scratch | Understanding Qwen architecture

https://www.youtube.com/watch?v=AzxoU7kxjig

orbital-decay 4 hours ago

Yes it's easy for an established shop, all they need to do is to tweak the post-training workflow. "Decision model" is the same kind of marketing as "LRM" attempted by OpenAI when RL CoT was new (to hyped up crowd). It's still fundamentally a classifier used for "decision making", games and RP were using generalist models and constrained outputs to do what the DOOM demo does for years.

janalsncm 3 hours ago

The interesting part is also the easy part. The model and architecture are not hard for an experienced machine learning engineer to build.

The hard part is the data and evaluation. Sure, it’s not that hard to build a fast model with good predictive power. But fast at doing what? You probably don’t care about classifying whether a hotdog is a sandwich (which is the Jev demo).

XCSme 4 hours ago

You can make a basic one in minutes based on existing open-source models.

Latency won't be that good, but could still work similarly. Simply force the structured output of a LLM to the given schema.

Probably also easy to train because we can use stronget LLMs to generate input/output data, or even synthetic data is easy to generate.

It's not really a new technology, it's more like a new use-case.

sigbottle 4 hours ago

What even are these new "decision models?" Take an existing LLM, feed it a prompt, force it to pick a choice; decode is 1 token (or rather, the whole logit set for only that last token; token implies selecting one logit) so you made a choice. That's it?

redox99 3 hours ago

orbital-decay 4 hours ago

popinman322 4 hours ago

redox99 4 hours ago

Yes it's very easy if you have fairly basic ML knowledge.

cakoose 2 hours ago

> This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed.

1. Humans are already not in the loop for lots of LLM agent actions. Isn't that just a function of how much you trust it and not some completely new paradigm? Am I missing something?

2. How can it gather context if it just outputs a single decision?

One guess: Maybe it's decision can be "gather more context and re-run me"? But an LLM can be much more expressive about what context it needs.

open592 5 hours ago

2 years in stealth...

ranyume 3 hours ago

I found the paragraph about how much networking data they have weird. I mean, if you already have all that data why didn't you train your models already using that? Why did you need clef to begin with?

ksymph 4 hours ago

With all these new Jev-like models popping up, has anyone actually started building anything with them yet? It's odd how quickly they've multiplied despite being relatively niche in their use cases, as far as I can tell. I suppose they're simple and cheap enough to make that it's a sort of 'why not' thing for a lot of these companies.

croemer 3 hours ago

Is there a Jev-like model I can run on my Mac? Something like Ollama? Or what's the best way to play with it? Is there a cheap/free API service eg on OpenRouter?

handfuloflight 3 hours ago

croemer 3 hours ago

Thanks for the tip! Set it up locally and it works!

okpatil 2 hours ago

We got you.

At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision

https://at0m.pienomial.com/ https://news.ycombinator.com/item?id=49920350

mpolichette 3 hours ago

Jumping on this, what about on-device?

I'd love an privacy first on-device model i could use in iOS.

okpatil 2 hours ago

We got you. How would you use it though ? Through Apps ? Would love to discuss.

At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision

https://at0m.pienomial.com/ https://news.ycombinator.com/item?id=49920350

okpatil 3 hours ago

Why give cloudflare your data ?

At0M: A 60M local Jev at 16 ms latency and 79% accuracy on Typed Decision

https://at0m.pienomial.com/ https://news.ycombinator.com/item?id=49920350

kamranjon 2 hours ago

This account was created only a couple weeks ago and seems to be spamming this closed model, seems it might be a bot?

okpatil 2 hours ago

Actually it was created around 12 years ago. We just launched the model around 3 days ago.

Apologies if it is too much of a bother.

afzalive 3 hours ago

Well, for one, that one's not released as far as I can tell.

okpatil 3 hours ago

It's open for benchmarking. V1 release next week.

Businesses are built on outliers. It doesn't make sense throwing your hard earned insights while paying them money to steal it.

Also, cloudflare https://robindev.substack.com/p/cloudflare-took-down-our-web...

Transformanshen 2 hours ago

It hasn't been released yet, and local versions aren't always convenient

okpatil 2 hours ago

The same engine is is available as an evaluation API.

curl -s -X POST https://at0m.pienomial.com/decide/v0 \ -H 'Content-Type: application/json' \ -d '{ "state": "Charged twice for the same card payment this morning.", "questions": { "queue": {"type":"choice", "instructions":"Which team should handle this?", "criteria": {"billing":"invoices, charges, refunds", "technical":"outages, bugs, deploys", "fraud":"unauthorised or suspicious activity"}}, "urgent": {"type":"noul", "instructions":"Needs action today."}}, "email_id": "you@example.com"}'

If it fits your use case, you are welcome to use it.

When it is a rust standalone rust executable, as it is powering the API, it becomes just plug and play. No dependencies needed.

cootsnuck 3 hours ago

If it's 60M local, is it open source as well?

okpatil 3 hours ago

It is a single rust executable, with model embedded inside of it. Current evaluation API is being run by the same.

We wanted to stress test the system before the V1 release.

dcastm 3 hours ago

What’s the context window?

okpatil 3 hours ago

32K

jasfi 3 hours ago

Related: an intelligence cache for decision model data: https://cachev.dev

I built this for my own needs, and thought others might find it useful too.

alex7o 3 hours ago

Oldy enough I tired this 2h ago as I was testing jev on cf and was like oh this should be a better replacement but it takes 3s which is useless to me

nikcub 2 hours ago

in a quick mini-bench here n=250 of clef vs jev, clef came out 5.2x more expensive, a lot slower (p50 of 350ms vs 1.9s) with only marginally better results (78.6% vs 79.8%)

6thbit 4 hours ago

I wonder if a good usecase for this would be cloudflare's WAF rules. Give broader request context to the decider and let it pick type of challenge/block traffic directly.

Perhaps that may be too costly atm

aryabakh 5 hours ago

it's great to see Cloudflare releasing consumer edge level models.

warkdarrior 5 hours ago

Can someone explain how so many folks managed to build decision models within days or weeks after Typesafe came out with Jev? Is this concept of decision models been in the works for a while? Is it easy to copy?

petercooper 5 hours ago

Smaller models have been able to do these sorts of tasks, but a little slower, for a while now. Give a small Qwen 3.8 model a classification task and force a structured output, and it'll do a good job. I've used Qwen 0.8b for basic image classification in <500ms on my local machine for a while now.

There are a few technical details that can reduce the latency significantly (covered in the post) but the real insight has been from watching the reaction to Jev and seeing that there's enough of a market interest to offer it as a distinct thing. The underlying concept/approach was already there.

theapadayo 4 hours ago

Not just structured output. Dropping down to logprobs, prompting the model to emit one word as the answer, and then ranking the output tokens to pick your answer works great on small Qwen & Gemma models.

The fascinating part to me is that Jev seems like this technique plus post-training to get multiple independent confidence values for each possible answer.

woah 5 hours ago

Transformers output a set of probabilities over outputs. For ChatGPT etc, those are predictions of what the next token will be. But it can also be a structured list of options or classes. Jev mostly innovated on the interface, API, and product concept around this, and made it click for a large number of people. Unfortunately for Jev, it's very easy to copy an API, and any pretrained LLM can be adapted to work in this way.

ford 5 hours ago

I think Jev also innovated on data & algorithms, but it remains to be seen if it's enough to be meaningfully better than traditional LLMs + a few tweaks.

ramoz 4 hours ago

Jev created accessible/programmatic ergonomics around a general purpose classifiers, and did it very well; ie intuitive api and structured data approach.

Anyone can copy that and apply to an array of models - stripped down LLMs or already slim/highly performant traditional classification architectures (just wrap inference with an api that inputs/outputs the same structured data).

Jev, I think, would say their advantage is the intelligence of their models and training data including calibration: https://medium.com/code-applied/calibrated-classifiers-makin... (which i still struggle with in the general application... there's no free lunch with these things).

nico 4 hours ago

Most answers explain the LLM-based approach to these models, which is also what Typesafe did with Jev. However, depending on what you need, there are far simpler classification models, and for a lot of use cases, these models can be way faster and more accurate than Jev

But, for these adhoc models, you need to understand the task more, collect some data and train the model (on CPU, no need for GPU). So Jev-like models are a great way of getting a hosted general decision model, but if you have a very narrow task or set of tasks, you might be better off with some more basic models that you can run on the same server you run other things or even on your laptop

didibus 5 hours ago

You can use already trained large transformer models to make one, so it doesn't require the kind of high-scale compute, high quality data, data cleanup, reinforcement, and so on training that say an LLM does.

233mhz 4 hours ago

What's new is "smart" decision models than you can supposedly use on anything without additional training.

If you have a very narrow use case you can train a BERT based decision model on a laptop an hour if you have good data to train it on. It'll answer faster than the roundtrip to clef/jev and use <1gb memory

conmod278 4 hours ago

If you have a very intelligent swiss army knife like hammer, that hammer will adapt to almost any nail, which is a good thing.

zitterbewegung 5 hours ago

You just have to fine tune an LLM like Qwen on some synthetic data to do so. There was even someone that had a model that was exactly like Typesafe and published their work a year before Jev (but wasn't marketed as heavily since it was academic).

pizzafeelsright 4 hours ago

The question of AI in automation is "can it make decisions in a consistent and predictable manner, with near 100% determinism?"

Many people seem to have run into the same question and started working out the answer.

kerenskiy 5 hours ago

The concept existed a year before Jev or so. See Laya

XTXinverseXTY an hour ago

Laya came after Jev, the original post is clear about this much [0].

Moreover the specific prior art claim is absurd (self-plug) [1]. GLiClass[2] is at least a coherent precedent.

[0]: https://laya.convaiinnovations.com/

[1]: https://xtxinversexty.com/layas-prior-art-claim-is-absurd/

[2]: https://github.com/knowledgator/gliclass

giancarlostoro 4 hours ago

It's not a new concept, it just took someone adding on to the approach and refining it. I never deep dove it, but I assume JEV is sort of like how Sora works? They had a blog post about how it has a sort of tiny LLM, which OpenAI's small LLMs are insanely good and well defined. I think any lab tackling this with a from-scratch model could yield affordable alternatives that are highly competitive.

It seems insanely obvious at least to me, that JEV is the new hot thing for the AI field since they give you stronger output that isn't... flat out wrong, that alone is impressive.

porridgeraisin 4 hours ago

They are not too difficult to train if you already have infra to train regular LLMs. You can typically replace a few layers train them alone and you're off to the races.

Getting training data that works well for calibrated classification objectives is difficult.

I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.

But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.

So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases. People mostly only care about accuracy, not confidence.

damsta 3 hours ago

Competition in this area is great and kudos for releasing something that we can try out today.

gitghxst 2 hours ago

that's good. I believe decision focused models will explode in the next few months

schainks 4 hours ago

AMAZING, thanks, Cloudflare!

swingboy 4 hours ago

It allows image input. Nice!

ttul 4 hours ago

That's probably driven by their own internal need to show the model images of emails and webpages to detect phishing, despite obfuscation of the underlying HTML.

wakeywakeywakey 2 hours ago

jev can ask jev if they should return investor money and shut down jev

swe_dima 4 hours ago

Would love to see benchmarks on visual tasks.

DesaiAshu 5 hours ago

brb while I build my entire cloud stack on Cloudflare

ralusek 3 hours ago

Just tested clef-flash vs jev:

- Jev/TypeSafe: 230 ms median, 254 ms mean

- Jev/OpenRouter: 237 ms median, 267 ms mean

- Clef Flash: 661 ms median, 806 ms mean

What gives?

MisterMunchkin 4 hours ago

Imagine making your whole company on one model and then being cucked by everyone within a week. I don't think I've ever seen anything like it.

hansonkd 4 hours ago

Yeah, these AI companies have some weird paradox that if they actually had a model that was super efficient and could arbitrage cost/intelligence of other inferior models, they would keep everything about it secret. If an intelligence research group had something groundbreaking, they would just dump their own money into the magic money machine.

Instead, to make up for the lack of economic viability of their models, they are forced to release publicly to get marketing to get others to pay based on hype.

globular-toast 4 hours ago

If they were actually useful they'd just make money doing the useful thing and wouldn't even talk about models or AI.

RGS1811 4 hours ago

If everyone else can spin up their own version of your product in under a month, there probably wasn't much product there.

zwaps 4 hours ago

No mention of calibration. Is it just another llm finetune?

kflansburg 4 hours ago

> Our post-training utilizes label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration.

esafak 4 hours ago

I feel bad for the Jev guys. I wonder if they anticipated this much competition?

hbcdbff 5 hours ago

“Urgency” of “yes”?

mococa 4 hours ago

I bet this's 100% slop.

selfawareMammal 4 hours ago

Horrible name

johnecheck 5 hours ago

Wow, Cloudflare is definitely buying some goodwill from me. Just consistently interesting new releases alongside and solid products at great prices. Seems nearly too good to be true.