Whistle: Speech to Text in 16.9 MB (cactuscompute.com)

380 points by gmays 5 hours ago

skolos 2 hours ago

Interesting that this is here. I used whistle (and bunch of other things) to take ownership of my echo show. It now doesn't dial to Amazon at all - it does all processing locally with its own CPU and connects to my homeassistant for home automation. My initial setup involved qwen asr (1.7b model) running on rtx 5080. Compared to that, whistle was really bad (out of 170 messages, qwen recognized correctly 168, whistle - 70), but I adjusted whistle to work like jev - instead of free form transcription it recognizes only select set of templates (I trained tiny network with 10,000 generated utterances to translate whistle final state to probabilities within templates). The precision went up to 164/170 - almost matching qwen. By the way - I'm speaking with heavy accent.

stronglikedan an hour ago

> By the way - I'm speaking with heavy accent.

I chuckled at this because my inner voice had an accent as I was reading your comment, due to your writing style.

ASalazarMX 9 minutes ago

Excuse me, could you write slower? I couldn't understand you.

yuchi 41 minutes ago

Sorry, curious non-native speaker here. Which accent? And which telling patterns made you think of it?

Nition 30 minutes ago

Zacharias030 28 minutes ago

schappim 21 minutes ago

Have you done a blog or YouTube about this ?

skolos 17 minutes ago

It is all custom made and not very reproducible. I'm working on reproducible setup and once it is done will publish it here: https://blog.kvit.app

mrguyorama 10 minutes ago

If you are purposely limiting yourself to select templates, even fairly complicated templates, then classic voice recognition is perfectly sufficient.

With a restricted grammar, built in Windows voice recognition, all on device, has managed this exact use case quite well for over a decade. I used it to try and build a clone of the various paid apps that allow you to issue orders to Arma soldiers with voice commands

INTPenis 4 hours ago

I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.

I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.

But every single sound he makes with his mouth ends up on the page too.

ComputerGuru 4 hours ago

Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.

Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...

cgbur 4 hours ago

For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.

lnenad 31 minutes ago

apitman 30 minutes ago

dv35z 3 hours ago

flockonus an hour ago

Imustaskforhelp 3 hours ago

nvtop 4 hours ago

I've been using Gemini Desktop App purely for dictation. It's a miracle! For the first time in my life I'm blown away by the quality of my (heavy accent) speech recognition. Just be sure to disable the "speak to window -> reasoning" option to make it purely dictation and stop from writing whole emails for you.

islewis 2 hours ago

> I don't think the challenge with speech to text was size of the binary

The usecase for small models like this is making on-device STT/TTS more accessable. This is important if your usecase is sensitive to either privacy or latency, but this comes at the cost of quality.

My experience has been that these small TTS models are unexpectedly good if your audio is in distribution (western accents, higher quality audio, common vocabulary), but pretty quickly degrade as you move outside of that. They often dont support more complex features such as diarization, multilingual, or realtime streaming either.

boplicity 4 hours ago

There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.

thayne an hour ago

One of the most common things I did with google assistant was tell it to remind me to do something, and it would reliably create a reminder on my calendar. With Gemini, it is very unreliable. Sometimes it does a google search. Sometimes it just opens a Gemini chat where it parrots back a (sometimes garbled) paraphrase of my request. Etc.

It's also really terrible at recognizing names of my contacts, probably because those names are not represented in the training data.

xp84 3 hours ago

This reminds me of how the thing iPhones had pre-Siri (so we're talking pre-2010), which was entirely offline, did a better job than even the most modern thing at "Play [one of the finite set of songs in my library]." I sometimes get absurd matches from bands I've never heard of, when the right answer is something right there in my library.

raddan 2 hours ago

hbn 2 hours ago

yu3zhou4 3 hours ago

I hope you will find a solid solution for your father.

I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.

apitman 22 minutes ago

This is a pretty cool one I saw recently: https://youtu.be/_j806JHhCRo?si=tr8ENJo_VlJF-b4y

Joel_Mckay 2 hours ago

Have you tried playing music from his childhood for 20 minutes a day before his writing sessions?

In some cases, this may improve function for a few hours. Best regards =3

yymir 4 hours ago

i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes

testycool 4 hours ago

Unrelated: I love your username.

albert_e 4 hours ago

What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.

raddan 2 hours ago

What do streaming implementations do when a bigger context reveals a different interpretation/parse? When I use Whisper in the terminal, I can see it going back and correcting itself. Are corrections off the table for a true streaming transcription?

solarkraft 3 hours ago

My mind is boggled by how many implementations miss this.

Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!

nicksaroha 2 hours ago

Yeah! Streaming is crucial for any real time use.

iforgotmypasswo 3 hours ago

This is the main reason I lean on Deepgram over local services.

paynedigital 2 hours ago

You can absolutely do high quality, low latency, even multilingual local streaming nowadays. As the commenter above says, Nemotron 3.5 Streaming is awesome. We make heavy use of it in our transcription app.

zimpenfish 3 hours ago

Tried it on a random TV episode and it seems to get stuck sometimes where it just outputs "Thank you." as a default - at one point emitting that for 60s of dialogue (and no, the episode does not have someone repeating "Thank you." for 60s.) Happens several times during the transcription.

aqfamnzc a few seconds ago

Reminds me of "foreign" and "[Applause]" frequently appearing on YT auto-subtitles.

jwr 3 hours ago

It's funny how so many models tend to generate "Thank you. Don't forget to subscribe" or "Thanks for watching" if there is silence. Shows you what they've been trained on :-)

wkcheng 4 hours ago

How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)

This definitely seems lighter and faster. How does accuracy compare?

jwr 3 hours ago

People keep praising Parakeet, but I've found it to be worse than Whisper Large. Yes, it is much, much faster and smaller, but accuracy matters a lot if you are to use dictation regularly and seriously.

I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).

weitendorf 2 hours ago

It’s really domain dependent I think. If you are doing anything conversational interfacing with less AI-familiar users, latency matters a lot.

If you’re feeding the results into a very smart LLM, it will figure out what you meant (but crucially ONLY if you warn it or tell it to do so, in some cases!). If you’re writing code directly or creating something for public consumption, you can’t tolerate mistakes. If you’re taking notes for yourself you just want it to work cheaply.

If you are ok with the complexity you can run both, and a Meta/Google open model with native audio, and let a smart LLM doctor it up. If performance really matters you can pay for a proprietary model or train one yourself. Until you get to that point, I think it probably doesn't matter much either way. Voice is just too easy to fiddle with

theturtletalks 3 hours ago

Parakeet is the gold standard. With models like moonshine and koroko (TTS model), it’s more about embedding the model in the application itself. If you’re using Parakeet, embedding it in the application is not feasible.

I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.

andy_ppp 5 hours ago

Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!

I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!

joewhale 4 hours ago

I initially read this as whistle to text, which would be way cooler.

mejutoco 4 hours ago

A good project for training an llm

https://en.wikipedia.org/wiki/Silbo_Gomero

jasonwatkinspdx 4 hours ago

I've met folks that descend from the Zapotec in southern Mexico, and they still use whistling language to talk to each other across mountain valleys.

stymaar 3 hours ago

charv 4 hours ago

Just like Marvel's Yondu!

amelius an hour ago

Let's say I want to build a hardware product now, voice-controlled, with voice feedback, so STT, LLM, and TTS. All local. What are the best libraries to do this now, say with 8GB of GPU memory available?

properbrew 2 hours ago

Might look into embedding Whistle into Whistle if it can make it 30x smaller (https://play.google.com/store/apps/details?id=com.blazingban...) - It's a shame there isn't as many languages supported though, I'm surprised at the amount of non-english downloads (I really shouldn't be, of course non-english speakers want dictation) of the app there is.

TomGarden 2 hours ago

The Qwen STT model that's leading the open weight leaderboard right now is excellent. Parakeet v2 is blazing fast and accurate enough on English. I feel like STT gains from this point on will be marginal, especially given that you can do a quick LLM pass afterwards with a small model

sfpk 3 hours ago

If you need speech2text try this, https://ccoreilly.github.io/vosk-browser/

jakobov 2 hours ago

ZWhispr seems to be the best for those who care about accuracy as it uses three SOTA models.

https://zwhispr.com/

alasano an hour ago

I find the whole category of "added value" STT apps funny at this point.

Built my own in a day that blows them out of the water. You can quite literally pick any sufficiently good local model or API provider and combine it with Cerebras for cheap and very fast AI post processing and formatting.

Lebenita an hour ago

macOS only (like so many in this space)

MisterMunchkin 3 hours ago

For something you could plausibly ship inside a webapp, it’s very good. I can definitely think of some cool uses for this.

e12e 4 hours ago

Hm. I saw language=detect and tried some Japanese - which (given the actual list of supported languages) unsurprisingly turned into some mangled Spanish.

Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).

So, reasonable, but limited?

armcat 4 hours ago

Those are insane benchmarks at this size. Well done!

kamranjon 4 hours ago

Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.

rafaelm 3 hours ago

Huh, this was really confusing. I already had an STT app called Whistle on my Android phone.

rshemet 2 hours ago

hey, Roman here from Cactus, thank you for the feature!

opening this thread for questions/feedback if you have any

jayshah5696 4 hours ago

This is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box.

Centigonal 3 hours ago

This is quite good, especially given the size and the fact that it runs on the CPU.

lab14 an hour ago

Tried it a few times with English, Spanish and French and the quality/accuracy is pretty "meh". If the model doesn't really work, it doesn't matter if it fits in 1MB.

mo2art 4 hours ago

RuntimeError: audio limit is 30 s

jjice 4 hours ago

Did you record over 30 seconds?

mrkn1 4 hours ago

love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free.

mrkn1 4 hours ago

pzo 3 hours ago

tested in polish and unless you speak very loud, clear and slow is not that good, parakeet definitely better.

saturn8601 4 hours ago

Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.

MayeulC 4 hours ago

Speech to text I assume? Maybe it has to do with your a accent or pronunciation? You could contribute a bit to Mozilla's Common voice, if that's the case. I assume it is part of every STT training corpus.

saturn8601 3 hours ago

Yes sorry Speech to Text. I have a standard US East coast accent but sometimes I speak a little mumbly. When I made an effort to speak more slowly and with a cleared throat there was some improvement but still not writing all words.

aidotguru 4 hours ago

eager to see if working in android phones

rpdillon 4 hours ago

FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model. I've used it for years now and love it.

https://futo.tech/

Edit: As others have pointed out, this is not actually open source. It's source-available, which is quite a bit different because folks can't fork and distribute it as easily. The license also appears to be revocable and non-transferable, which makes it different from open source licenses.

robertlane0 4 hours ago

I'd classify it as "source-available" given the noncommercial clause in the license.

https://github.com/futo-org/android-keyboard/blob/master/LIC...

rpdillon 3 hours ago

lrvick 4 hours ago

Futo only produces source-available proprietary software. They most certainly are not Open Source, though they unfortunately lied about this a lot before they got called out enough times.

https://github.com/futo-org/voice-input/blob/master/LICENSE....

rpdillon 3 hours ago

helterskelter2 4 hours ago

I've used FUDO keyboard for a long time, but I never tried using the voice input, so I'm testing it now. Let's see how well it transcribes everything.

...Okay that was pretty good.

paaloeye 2 hours ago

RIP Wispr Flow

tecleandor 5 hours ago

Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...

kaoD 4 hours ago

Spanish from where? Here (Castilian Spanish) it seemed to work fine.

tecleandor 4 hours ago

Madrid. But it will only work properly if I'm clearly dictating with a very regular rhythm (ViaVoice dictation, if anyone remembers...). If I use a more natural/conversational rhythm (no slang, no abbreviations...) it easily confuses words.

chilicuil 4 hours ago

Mexican and venezuelan aren't detected correctly

contingencies 2 hours ago

For speech to text UX I currently use https://handy.computer/ as it's cross platform and open source. With that I am currently using Parakeet Unified EN 0.6B and finding it excellent. Often I use it to talk to AIs without giving them audio, which works very well. Honestly, I would never go back to typing now. Promised since ~Y2K, the tech is finally here. You really notice it when you wake up at 2AM and don't want to wake people ... it can get really annoying reverting to key-tapping. My long-gnawing fear of losing my hands to RSI is no longer a thing, and I can focus on losing them to another hobby: like sailing or machining! Just bought a band saw...

try-working 2 hours ago

I built an STT plugin for DeepSeek Harness that uses this Whistle model as well as a larger one from Desert Ant Labs: https://github.com/try-works/dsh-stt

thomkaar 3 hours ago

it can't even tell when i say "my booty thick"

thomkaar 3 hours ago

i can't even tell when i say my booty thick

agilek 4 hours ago

Can we have more languages?

mrkn1 4 hours ago

yapsnap has 10 languages