Ollaya – Ollama for open-source, Jev-style decision models (ollaya.dev)
231 points by Ardakilic 4 hours ago
alex7o 2 hours ago
Guys I have a real q, what is the difference between an instruct based re-ranker and laya/jev I just don't see it.
Edit: One is that jev/laya are tuned to have better probabilities, but a reranker can be fine tuned to do that as well. And jev/laya use RLCD?
Swizec 2 hours ago
> difference between an instruct based re-ranker and laya/jev I just don't see it
Main difference is that laya/jev/et-al give you a zero-shot classifier that requires no training. You can prompt engineer your way to a quick fairly reliable cheap enough decision engine that you can use to iterate quickly (by prompt engineering).
Right now a lot of people are doing this with LLMs and it's too slow and expensive.
Imo the right iterative approach to productionizing these systems is something like:
1. Build it with an LLM. Iterate on the prompt
2. Start building a real-world dataset
3. When the prompt works, turn it into a clear rubric for Jev or similar
4. Keep iterating until desired accuracy achieved
5. Use the real-world evals you've built to train a custom classifier fine-tuned to your needs
You now have a system that has produced useful results in production from the very beginning and by the end it's a reliable super cheap classifier that can make thousands of decisions per second.janalsncm 30 minutes ago
I don’t think that’s it. I sincerely doubt most developers are doing side by side comparisons of calibration quality.
OpenAI has a section on their embeddings model api page for zero shot classification. Of course you can choose an open weights embedding too if you’d like.
https://developers.openai.com/cookbook/examples/zero-shot_cl...
I think Jev wins on marketing and convenience. Most SWEs don’t want to talk about embeddings, cosine similarity, or precision/recall tradeoffs. They want something which plausibly works and is easy to use.
kakugawa 25 minutes ago
Jev's value becomes more apparent when the task is a moving target. eg an auto-mode classifier.
avereveard 2 hours ago
Calibrated probability across multi task with zero shot I guess. A reranker is single task and tuning it make it even more narrow. And I guess some piping to make multiclass efficient since you cannot mask logprob for independent questions in the same output space without throwing calibration away.
solaire_oa 29 minutes ago
I installed it, I tried the examples, it works.... But forgive my lack of imagination... what is this useful for?
Like, their example is of classification for a support interface.... `refund_requested`. Pretty convenient bool given the example is about a refund- what if 99% of submissions don't ask about a refund? Also, is that user not a `churn_risk`? What could possibly qualify as a churn risk if not a user asking for a refund?
https://ollaya.dev/library/laya The examples suffer the same problem of why I'd prefer to use a string column vs an enum. Changing an enum means you need to update the db, using a string you can do whatever.
I'm not trying to be negative, I genuinely want to know about some practical examples (that don't require tons of backwards maintenance).
devttyeu 9 minutes ago
I have a lot of semi-practical examples of how you can use this model wrapped in unix-ish tools - https://github.com/aurorainfra/grev (readme links to docs of each tool with some more or less practical examples)
Really I think "smart grep" is a pretty good one ('look for an error looking vaguely like this'). Also I think sql-based shell history + decision model is quite good to make the last 'which one of those choices is best fit gives users past few commands' etc.
spaniard89277 6 minutes ago
Isn't it better to use an LLM to train modernbert or xgboost et al?
devttyeu a few seconds ago
motoboi 6 minutes ago
Is for when you want an AI to make a decision. If you have been using gpt or claude or open source models for that, than it’s a way cheaper alternative.
And if you have not been, it’s for when you have to extract the context from text. When you have numbers or fixed options, it’s just a matter of code.
So if you find yourself having to decide if a given user comment is a refund_request, that’s for that.
It’s not perfect, you still have to fine-tune (or calibrate) using examples you have (and keep those examples updated over time). But it’s way better than trying to parse text with regexes.
george_max 3 hours ago
Has anyone actually seen better or the same results with Laya compared to Jev? From my experience, Laya performs significantly worse. It's less confident and often makes wrong decisions with more complex queries.
jonmagic 3 hours ago
I've been following jevbench twice a day for the past week and that's been a lot of fun. Latest update:
Rank System Score Public / sealed accuracy Evidence
1 decider-4b v2 64.13 83.5% / 34.7% Evaluator-run, offline
2 Jev 1.13 63.29 86.6% / 36.7% Evaluator-run API
3 JevK5 v0.2 62.04 85.3% / 33.1% Evaluator-run
4 Cygnet 12B 61.76 87.9% / 33.8% Evaluator-run, offline
5 Hopper 59.43 82.3% / 34.1% Evaluator-run
28 Kev 4B 36.14 66.2% / 22.4% Evaluator-run
41 Laya 421M 30.25 58.4% / 30.8% Evaluator-run
Havoc an hour ago
Amazing - was looking for some benchmarks around this earlier
philipodonnell 2 hours ago
What the best way to see how a homegrown version compares?
scronkfinkle 3 hours ago
Yes. JEV generalizes better because they probably have an enormous corpus and trained on it for a long time. Laya's out of the box model is much weaker. However, in the age of LLM's it's incredibly easy and cheap to generate large datasets to fine tune laya for your task, and the training loop is pretty quick and cheap too.
It's so easy that I question why I would ever pay for JEV when eventually I'll have done enough random things that I will also have a large corpus and likely a general model as well.
mtkd 3 hours ago
Isn't the point of Jev that it generalises better?
It's a fast classifier you can use out-the-box, ~1.5bn tokens is about $40 (I've been hammering it)
It just works ... a whole bunch of low-level/low-importance workflow stuff that was getting farmed out to small/fast LLM models now has a competitive alternative ... and bits that hadn't even been considered to go into some external descision/classifier service can be tested/deployed at ~$0.00003/req
I don't get this wall of negativity on it, it's genuinely innovative/useful tech ... would expect HN to be more positive, regardless of whether it's the absolute best execution
digitaltrees 2 hours ago
shepardrtc 3 hours ago
not_a_bot_4sho 2 hours ago
cobanov 3 hours ago
Developer here. You're right, Laya is a lot weaker than Jev, especially on harder queries. It's a small model, so it's fast, but that's the trade-off. The open models that get close to Jev are much bigger, and running those is what I'm working on next.
adinb 26 minutes ago
It doesn’t to be a ton bigger, 16k and reliable 8k would be a godsend. (I run at 2k)
mikodin 3 hours ago
What are the models? I am super curious in these as well
simcop2387 2 hours ago
jasonjmcghee 31 minutes ago
In my experience it's not close and the benchmarks I've seen don't reflect my experience at all.
But I'm guessing people will find the right training regime and data mix soon to close the gap.
But big things I see are instability and inaccuracy - like pick a random problem.
verdverm 2 hours ago
one day, perhaps people will click through to the laya author's arxiv paper content and the why may become clearer, you won't have to read it, a skim will suffice
iamflimflam1 3 hours ago
Nothing yet. Unfortunately it sometimes feels like our industry has been overrun by grifters and chancers.
I’m sure this has been a gradual and long decline. Maybe it even started with the dot com boom and accelerated with crypto. With AI it seems to have got worse.
pradn an hour ago
I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly (what, like 2 weeks?). There's "consumer surplus" for everyone, to borrow an economic concept. But we do ideally want some of the surplus to flow to the innovator, too. I know there were precursors, but that's fine - it's hard to have a totally novel idea in such a popular field. I don't know what the end game is for TypeSafe - they'd need to demonstrate perpetually better results, or compete in another axis: UX, support, custom solutions, etc. So much of the time, someone proving a concept, or it simply getting enough publicity, is enough for a "Cambrian explosion" of follow-ups and copies. Famously, that was true for "Attention is All You Need", and the general idea of "next-token prediction" being so powerful.
We've stumbled into general differentiable models..
redox99 6 minutes ago
Because what they did is kinda trivial. Its basically like the Dropbox comment really[0], except here you don't need petabytes of storage and infinite VC pockets.
After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.
To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.
The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.
totetsu 32 minutes ago
Are you saying laya copied from jev, and released in two weeks? If so I don’t thinks it’s quite as simple a story as that. https://xtxinversexty.com/layas-prior-art-claim-is-absurd/
janalsncm an hour ago
Presumably the training recipe and training dataset itself cannot be easily copied in a week or two. So if they want to shut down these competitor models they need to make it obvious how they are better than them.
ranyume 3 hours ago
>Run decision models locally.
>example is a text classification task instead of a decision
hbrn 3 hours ago
"Decision model" is just marketing jargon.
decision model = classifier
system one model = small non-reasoning LLM
noul = boolean
confidence = f(probabilities)
It's sad to see how gullible engineers are today.
verdverm 2 hours ago
> how gullible ... today
that laya is even a thing is further evidence, people took that author at face value, the paper contents are incomplete and describe something that does not sound like Jev at all
this was the period of arxiv history that led to the new vouching system, laya author contributed to that imo
hbrn an hour ago
OgAstorga 3 hours ago
text classification is equivalente to decision. This is exactly the same thing Jev does.
ranyume 3 hours ago
If it has four legs, a tail and barks why not call it a dog?
gchamonlive 3 hours ago
seemaze 2 hours ago
ricardobeat 3 hours ago
It is not. In a benchmark with actual decisions - navigation, traffic, waypoints - laya does only slightly better than a small classifier.
abirch 3 hours ago
Jev does it more efficiently because it doesn't use an LLM https://typesafe.ai/blog/introducing-system-one-models-and-j...
rockinghigh 2 hours ago
cobanov 3 hours ago
Fair point, that example is basically classification. I'll change it to something that looks more like a real decision.
nacs an hour ago
It would be good to list 1) zero-shot accuracy and 2) latency on the models page . The LLM-based models' latency is probably much higher than the BERT approaches I would assume.
Also curious, it seems from looking at the accuracy scores you gave that it seems to be NLI > Gliclass > Laya (for Bert types)? Why do you seem to feature/recommend Laya more - is Laya better in some way?
nickstinemates an hour ago
Laya is pretty easy to set up on its own without ollaya. I just did that and replaced my current jev API usage to laya running on a GTX 970 with 4GB of vram.
Very small context window, but for some existing small llm work I was doing, it was a drop-in replacement and it makes me happy I can get use out of old hardware I have running.
mococa 3 hours ago
It would be really cool to have LLMs and System One in a single tool - in this case, if Ollama implemented it.
verdverm 41 minutes ago
next vLLM release will have this
if you use gateways, GoModel support the S1 endpoints, my favorite feature is the virtual models, stable name, I can swap out the backing model(s)
https://gomodel.enterpilot.io/docs/getting-started/quickstar...
(the "kev" in the docs is my fault, I should have said Jev / System1 in my feature request)
thih9 an hour ago
FAQ[1] says:
> It is an independent project, not affiliated with Ollama.
handfuloflight 3 hours ago
Sounds good on latency but how is its actual decision quality vs. Jev?
cobanov 3 hours ago
Depends on the model. The small ones I support today are well below Jev on harder queries, but fine for simple, well-defined questions. The open models that get close to Jev are bigger, and I'm adding support for those next.
qurren an hour ago
Would be great if you supported CUDA 12; I don't feel like paying $15K to upgrade my GPU right now
verdverm 42 minutes ago
wait another week or so for vLLM's next release
datadrivenangel 4 hours ago
Are there many models that are comparable to Jev for generic decision making?
Smarter move if you have an eval set is to just train a classifier and call it a day.
rgbrgb 3 hours ago
there's this thing with a bunch of similar models https://huggingface.co/spaces/multimodalart/jev-decision-ind...
top open one is trained by perplexity cto for $3k, kinda cool https://x.com/denisyarats/status/2102252088067850507
physicallyIllfr 3 hours ago
<<<"i was curious to see if i could train a competitive Jev-like model completely autonomously with a swarm of agents using our internal system."
Bro is writing off the H200 lol
On a sidenote I really can't stand the term "swarm" and definately plays into AI doomerism.
cobanov 3 hours ago
The link rgbrgb posted is a good overview. The best open ones are close to Jev now, but they're big models. And I agree, if you have an eval set for a fixed task, a trained classifier is the better choice.
oguzhankayan an hour ago
Nice work! Making open models easier to run locally is valuable on its own. Keeping the API compatible with Jev is a thoughtful touch, too.
emmettbt 3 hours ago
Cool... but this does seem undermined by the fact that Ollama can add support for decision models at any time.
cobanov 3 hours ago
Fair, and I'd be happy if they did. Ollaya uses the same API as Jev, so your code isn't tied to it either way
accountrequired 3 hours ago
and that ollama is go-llama and not rust, so it's not really the ollama of anything
vorticalbox 2 hours ago
Does anyone know what laya multi lang is faster than laya en? I would have thought focusing on a single language would be faster.
gauravsapkotanp 3 hours ago
I have also tried this and its really awesome
george_max 3 hours ago
I am fairly confident if Jev-style decision models are seen as prominent (which, they seem to be), Ollama will support them. Surprised the team hasn't implemented this already.
eserozvataf 3 hours ago
great project for empowering open-source alternatives.
rkovashikawa 3 hours ago
open-source is the only way for safe AI development. whoever doesn’t share the weights/code will lag behind.
cobanov 3 hours ago
Thanks!
pishpash 10 minutes ago
Why do you need another model-type specific Ollama? Can't Ollama be made to support these models?
adityamwagh 2 hours ago
Hey Claude, make ollama for Jev like models. Make no mistakes /s
verdverm 2 hours ago
hey Claude, download and run vllm nightly for me
(already merged)
GoModel (gateway) already supports Jev like endpoints too
amar-laksh 2 hours ago
This inference engine is soooo much faster btw: https://github.com/tamnd/kime