Typesafe-computer-use drives a Mac toward a goal for 1/50th of a cent per step (github.com)

73 points by rahimnathwani 2 days ago

jdkoeck 3 hours ago

I trailed off a few lines into the README. No human ever edited any of this. « LLM detected, project rejected ».

antoniojtorres 3 hours ago

Every third post in HN has this same complaint comment. Everybody knows.

petterroea 2 hours ago

It's better that everyone is loud about it then everyone giving up and being silently irritated. At least if people complain it's possible to read the room

Rexxar 18 minutes ago

I prefer when I see other comments to confirm or infirm my opinion.

flohofwoe 16 minutes ago

Tbh, I still wonder why people even bother releasing such tools that were shitted out by an LLM in an hour or two.

Everybody else can do this now themselves, and they get a tool that's more personalized to their tastes. And why have a readme at all when it's just LLM mumbo-jumbo - that sort of mumbo jumbo is written for LLMs to consume, not humans.

zahlman 2 minutes ago

slacktivism123 25 minutes ago

>Everybody knows.

Awesome!

I think these repo owners should be required to film themselves reading out their Claudemade readmes with a straight face.

Every honest caveat. Every seam. Every "Frames lie, so the walk prunes hard".

jdlshore 2 hours ago

I appreciate it when people say something is slop. Saves me from wasting my time looking at it.

How many slop ideas have come to the front page, never to be heard from again because the execution isn’t actually any good? I’m guessing most of them.

ithkuil 37 minutes ago

theturtletalks 2 hours ago

IshKebab an hour ago

Clearly the people writing these annoying articles don't know.

I wish they'd just use Astra instead. It doesn't write this awful prose.

soulofmischief 2 hours ago

And every comment on HN calling out vibe-coded slop has this same complaint comment in turn. OP has an actual complaint, your complaint is just "stop complaining".

antoniojtorres an hour ago

alex_suzuki 3 hours ago

I stopped reading at “The honest caveat”…

undefined 2 hours ago

[deleted]

undefined 2 hours ago

[deleted]

ako 2 hours ago

I thought it was a fine informative readme, starts with the problem, outlines the core of the solutions, and some limitations. Everything I want to know in the first few paragraphs. No need to spend human time to improve it.

jimmySixDOF 2 hours ago

not sure how this is innovative they show the System-1 model can play Doom right in the announcement [1] :

>Doom >We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI. The engineer behind it was worried about making 10 queries a second (which ends up costing ~$7/hour), but the rest of us agreed that was lower than expected! This is so fun we intend to not only release an in-depth walkthrough, but also host some events to hack on this.

[1] https://typesafe.ai/blog/introducing-system-one-models-and-j...

orbital-decay 2 hours ago

Well nothing about the Doom demo or this entire model is new new either, is it? I don't even think Typesafe themselves are claiming anything novel, they say that they're focusing on practicality instead of chasing big numbers and AGI. Classifiers are older than generative models and are used everywhere. Fast classifiers are used in sampling machinery of every big model and for automation in agentic game plugins for years, except they're usually small and finetuned for the task, not general-use.

I think many people wondered why non-generative models are so underused on a big scale, well here's a long overdue attempt to market that which evidently goes well with people being interested in this again. The field has been captured by the vibe coding and valuation-goes-up hype and a bit. AI has a ton of low hanging fruits that are much more practical than using one tool that gets most attention for everything.

altmanaltman an hour ago

So it's kind of like AI hype growing up and rediscovering its roots because the future isn't futuring soon enough.

mmastrac 3 hours ago

I'd be interested to see if using DiffusionGemma-as-Jev helps as you can feed the image directly into the model and it'll make decisions based on the image embeddings.

nowittyusername 3 hours ago

I had a long talk with chat gpt about this today as well. I think its duable and prolly not too hard either, also you could do lotsa funky stuff with stitched frames of a video in one 4x4 grid for example and send that as one image for analysis. that way temporal understanding can be had for fractions of a second by jev... also because vlm works in pixel space you can get around the whole state machine issue as well, so many possibilities...

zjy365 2 hours ago

[flagged]

Zaraif13 4 hours ago

How does it do on OSWorld-verified? Recently read that even Fable 5 is just at 85% .

jgilias 3 hours ago

This probably throws a spanner in the wheels there:

> Every piece of reasoning the frontier model does for free has to be rebuilt here as deterministic state.

EDIT: Not to shit on this though. I totally believe that some smart mixture of LLM-reasoning + Jev-style + determinism is going to be pretty amazing.

aruss an hour ago

It seems like the more honest comparison would be to OCR the screen and send that as input to the LLM?

john_minsk 4 hours ago

Super cool. Hope waitlist will move soon. I have a use case for it too.

are you the author? If so - what are your notes on using Jev in this scenario?

jasonjmcghee 4 hours ago

Did you do the follow up questions? I was invited within a few hours of joining today.

It's also now on openrouter and cloudflare

baxtr 4 hours ago

Same here. 8h hours later I was in.

sroussey 3 hours ago

Curious about people’s experience here. I am working on a small model, verify by jev, and escalate to big model. Some cases, the small model is not a model but some regex.

cheap-confirm-escalate

Using jev as the confirm step.

Surac 2 hours ago

I did not understand what this is all about. Anyone with more brain than me can explain please?

orbital-decay an hour ago

It's a quick proof of concept of computer use powered by Jev, a new general-purpose classifier model. Here it takes the description of the interface (it can't do images yet) and outputs commands. It's faster and many times cheaper than using frontier generative models like GPT to do the same.

smartbit an hour ago

curtisblaine an hour ago

"The honest caveat"

pulvinar 3 hours ago

[dead]