Transcribe.cpp (workshop.cjpais.com)

746 points by sebjones 2 days ago

rmunn a day ago

Looks very cool. One thing I have been looking for, which this doesn't seem to cover (at least I didn't see any mention of IPA in the model documentation), is a way to transcribe unknown languages phonetically, using the International Phonetic Alphabet to spell them (sound-based spelling rather than meaning-based spelling). I know several linguists doing research on minority languages (fewer than 10,000 speakers in some cases), which are small enough that they will never have enough effort made towards training language-specific models in that language.

Are there models I'm not aware of that are trained for this task? Taking audio in an unknown language, and rather than identifying the language, just transcribing the sounds to IPA? That would not be useful to most people, but it would be a Godsend to many, many linguists working with minority languages around the world.

shenberg a day ago

We take a lot of shortcuts when speaking, it's actually much harder to transcribe phonemes than to transcribe words, even when aware of the language being spoken. Some models have been trained for the task (e.g. look at https://huggingface.co/spaces/KoelLabs/IPA-Transcription-EN ), but the error rate is really high.

rmunn a day ago

There are, broadly, two kinds of audio recordings that linguists want to transcribe. One is native speakers telling traditional stories, where they're speaking naturally and taking the natural shortcuts (such as "wanna" and "gonna" in English). The other is native speakers reading words (or short example sentences) very carefully and distinctly, so that the linguist can listen to the recording over and over to learn how to pronounce the word right. In those recording, they'll say "want to" and "going to" rather than "wanna" and "gonna".

Thanks for the pointer; I'll check out that model and see if it handles the "slowly and carefully" type of recording better than the "natural speaking" type. (And depending on what kinds of errors the model makes, even the recordings where it makes errors can prove useful: for example, a linguist studying regional variations in speech would want the model to produce the IPA for "gonna" rather than "going to").

uoaei a day ago

Dialects degenerate phonemes that would otherwise occupy identity relations between different utterances of the same "word" (word/concept mutable hyperobject as is the standard in any socially-relevant spoken language) which would be a bit of irony in this thought experiment since common knowledge dictates that more data samples must be present in the dataset (not less, as in rarely-spoken languages) to associate separate pronunciations of utterances representing the same underlying concept. However very-rarely-spoken languages probably don't have distinct dialects since so much focus is put on mutual intelligibility with the few members of the group that remain fluent in that language. It's not outside the realm of possibility that small speaking communities nonetheless fractionate into dialectical specialities but that seems increasingly unlikely as the fervor for preserving/recognizing dying languages increases, and global instant communication continues to become more commonplace.

Example: Schwabisch is wild and would be phonetically transcribed very differently from Hochdeutsch which is its ostensible language progenitor (technically more a cousin than an ancestor in the lineage of language evolution), but if the goal is merely to focus the model purely on phonetic transcription then you can add additional post-processing layers which map sounds to core concepts shared across dialects for actual translation. But I like your idea of interacting with the intermediate elements to familiarize yourself at least with the phonetic patterns, we humans are still thinkers enough to infer patterns of grammar and semantics from these building blocks just as we have done for the entire history of the species/lineage before written representations of language came along (relatively late -- evidence of script cropped up only once civilization had centralized to a sufficient degree to make economics non-local and non-trivial).

tl;dr the big words: it's not til you collect enough spoken samples of the dead(ish/dying) language being spoken that the local idiosyncracies are discovered, luckily linguists are smart enough to probably anticipate and certainly post-process language snippets to grasp the common structures for this or that given language.

rmunn a day ago

akreal 3 hours ago

This field got reactivated recently, some new models were released (in the order of release, but the last one seems the best at the moment):

1. ZIPA https://github.com/lingjzhu/zipa

2. POWSM https://huggingface.co/espnet/powsm

3. PhoneticXEUS https://github.com/changelinglab/PhoneticXeus

I would be curious to know how to help people to use these models, or what kind of tasks they could be applied to.

verst a day ago

I would love such a model. My wife's family is Iu Mien which is a sub group of the Dao/Yao Chinese ethnic minority. Mien is its own language but most speakers are essentially illiterate. I'm good with language but there simply isn't a course or any books for learning the language. Not much in writing to begin with given the high illiteracy rate. I would love to build a translation system - project Hail Mary style :)

yorwba 20 hours ago

FWIW someone wrote their PhD thesis on the Iu Mien language: https://opal.latrobe.edu.au/articles/thesis/An_Iu_Mien_gramm... It's not a textbook, but that doesn't mean you can't try to use it as a textbook. For example, there are some thoroughly-analyzed example texts. Also, the acknowledgments section mentions the Iu Mien Literacy Project, which has published a short course on the very basics, as it turns out: https://iumienliteracyprojects.com/resources/ I suspect the 925-page PDF contains some more hints to additional resources like that. :)

verst 11 hours ago

sipjca a day ago

Largely this is out of scope for the library, mainly because I’m not aware of many models supporting this. but if there are models which support this would be happy to support

simsla a day ago

Automatic Phoneme Recognition (APR). There are some models that do this, but they're only so-so.

msm_ 18 hours ago

It sounds like the only way this would make sense is if such model knew the range of sounds it expects to "hear". There's a lot of possible sounds that IPA knows about, but world languages only use a fraction of them at once. Think English dark and light "l" (ball/light) or aspirated "p" (pin/spin) - some languages contrast them, while in english the difference is not meaningful.

Or maybe linguists are actually interested in having maximally faithful IPA representation and manually normalizing it? You are clearly way more knowledgeable about that topic than I, so I'm curious what you think.

rmunn 14 hours ago

The linguists I know are not necessarily a representative sample... but they're mostly interested in just that: maximally faithful IPA representation. They want to know if the speaker switches back and forth between aspirated p and unaspirated p on the same word, because that tells them something about the language — that aspirated consonants are not meaningful.

Linguists studying the sounds of a language, its phonology, often want to find "minimal pairs", words that differ by only a single sound. For example, din and tin in English. You record a native speaker saying both words, and telling you their meaning, and then you play back either recording A or recording B to other native speakers and ask them which word it is. If they can identify the word every time, then you've found two sounds that are meaningfully distinct in this language. (Some languages don't distinguish the d and t sounds, but English does). But if the native speakers go 50/50 on which word it is, or ask to hear it in a sentence for clarification because it could be two or three different words, then you've found a pair of sounds that this language does not distinguish. (Note that you're playing the words in isolation, because sentence context might make it obvious which one it is, e.g. you can't tell if an English speaker is saying their or there until you hear more words of the sentence).

So yes, the linguists that I know (who, again, are not necessarily a representative sample) are interested in as faithful an IPA representation as they can get, because that inconsistent transcription will give them many clues about the language. It still all has to be checked, because that switching back and forth between aspirated and unaspirated p (for example) could have been an artifact of a poor-quality microphone not picking up the aspiration, or a windy day causing aspiration sounds that the speaker never said, or the speech-to-text model making a mistake. But I watched my wife listen to the same two-second recording on loop over and over, trying to be certain of which sound she was hearing in the middle of the word. Double-checking the output of the model would (in most cases) only require listening to the audio once or twice, not half-a-dozen times like she typically did while researching her thesis. At the time, LLMs were not really a thing yet, but if she were doing her thesis today I bet a speech-to-IPA model would have saved her quite a lot of time — but only if it output every distinction, even the ones not meaningful in the target language. The "maximally faithful" representation, as you put it.

joshspankit 13 hours ago

KingMob a day ago

I learned IPA for Thai, and as part of that, I also read that a lot of professional linguists still find IPA too limited.

IPA seems very comprehensive from my amateur perspective, but apparently a lot of modern linguists still extend it or roll their own.

ghm2199 a day ago

Congrats on shipping this. I love handy on my Mac, my phone for STT in situations where it’s not possible/poor performance of the native Model for STT(e.g apple’s thing is not upto scruff, like mistranslating words corresponding to a domain).

Noob question: How do you think about funding from a foundation(i have no clue if you need it or not, I do hope you have a way to get paid one way or another because handy is amazing) for maintenance of this? if you did or were going to get paid by asking for maintaining such a project what might be the kind of organizations you would look for to get supported and how would you do it?

sipjca a day ago

Thanks! What an excellent question, I’m not sure I have a good answer. I kind of became an open source maintainer by accident as Handy became popular

Certainly I am very lucky that quite a few people donate to Handy, and also some people and organizations who sponsor the work I do

To be honest I just love contributing to open source and wish to continue to do so. So anyone who supports this is good to me. Organizations which believe in OSS and push it forward are typically most aligned with me

Of course you can always email me (contact@handy.computer) and we can discuss in more detail

sneak a day ago

OS-native dictation on iOS requires uploading your address book to Apple on every request, even if you don’t use iCloud. I unfortunately have to leave it disabled for this reason.

nohup2 a day ago

Are you sure? Just tested it and it works locally and offline

jjice a day ago

sneak a day ago

espetro a day ago

Oh my, this is the first time I read about it ever. Thanks for sharing it! I'll stop using the built-in dictation now.

abdullahkhalids a day ago

For anyone looking to build on top of this. I have tried a few different STT systems, and they accurately capture what I am saying. Unfortunately, they don't support the reasonable workflow

I want to open an office document, for example, and start talking. And I want the software to continuously type what I am saying at the cursor with minimal latency. The continuous part is crucial. Many software will paste whatever I said after I have stopped recording, but that is not useful.

primaprashant a day ago

Totally understandable, but I’ve found that software that transcribes everything after I finish recording actually works better for me. I’ve tried both kinds, and systems that continuously type what I’m saying distract me from completing my thought. I end up reading what’s being typed and noticing transcription mistakes instead of focusing on what I’m trying to say.

I often prefer to dictate everything in my head about a particular thing for 5–10 minutes and then go through it afterward. I find that much more useful because it doesn’t break my thought process the way continuous transcription does.

solarkraft a day ago

I can understand both modes. I mostly use transcription as input for my AI assistant and there I find it very useful to be able to check my input and just repeat myself in case something wasn’t fully captured. When using Apple’s transcription feature built into iOS and macOS, I also really like being able to edit everything right while the dictation is still active.

abdullahkhalids a day ago

Dream would be a combination of AI models that are smart like a human transcriber. One can just tell the model what mistakes it has made (e.g. "No, you wrote XYZ but I meant WXY"), and it is intelligent enough to realize when it is being instructed and when it needs to transcribe exactly what I say.

sipjca a day ago

You can fairly easily modify [Handy](https://handy.computer) to do this if you want

I’m planning on having it as a first class feature of the app too just too many other issues to work on first

mft_ a day ago

I’m really glad to hear this!

A while ago, I auditioned about 10 different STT apps on my Mac, with this realtime/streaming transcription as a goal. I failed to find that feature in an app I was happy with, but settled on Handy as the best option otherwise. So if Handy adds this, it will be perfect!

rolisz a day ago

Can you give some pointers around this? I'd gladly help with a PR for this, but if you have anything docs/ideas around this it would be helpful.

sipjca a day ago

mmmmbbbhb a day ago

You know English doesn't work like that. The word you're saying only becomes clear with the surrounding context. Eg, 'there' vs 'their'.

nilslindemann a day ago

It may be interesting to have it immediately insert the words, even if they are wrong, and when a sentence is finished, replace what has been written with the final corrected sentence.

atonse a day ago

abdullahkhalids a day ago

atonse a day ago

But this is still possible to do if you track the whole run of text. You could replace all of it each time so it LOOKS like it’s streaming but earlier words also change. I’m hoping the streaming models do this eventually.

I believe the built-in iOS dictation already does this.

kristiandupont a day ago

PhilippGille a day ago

Handy already supports streaming transcription models, and you can see the words in the small Handy pop-up while you are talking.

So in general this definitely works. Handy is just missing the feature to insert these streamed words into the app where the cursor is.

regularfry a day ago

dostick a day ago

Model should be able to understand where logical sentence ends, to stop buffering, and optionally rewrite some of the test that has already been output.

samplifier a day ago

jiehong a day ago

> The continuous part is crucial. Many software will paste whatever I said after I have stopped recording, but that is not useful.

It really depends on how one uses transcription.

For example, I really value being able to open different windows, and look at graphs, or scroll some data while I'm dictating, because it can help me with providing some support information for what I'm saying.

Some apps can even take into account things you copy or look at as part of the transcription's context to improve the results [0].

[0]: https://superwhisper.com/docs/common-issues/context#types-of...

abdullahkhalids a day ago

This is pretty cool. The dream would be that all editors (including brower tabs) have an api interface which allows arbitrary LLMs to modify the editor content. So you can indeed look at different windows as you talk, but the editor keeps getting updated live, so you can go back and see where you already at.

electronstudio a day ago

This is what I attempted with https://github.com/electronstudio/low_latency_dictation

However the accuracy of the real time models is poor, so I did a second pass with a higher accuracy model before committing the text.

mijoharas a day ago

Agreed. It's something I've found annoying about a few systems.

It looks like the rust bindings have streaming examples so hopefully there is a nice solution here.

catmanjan a day ago

You used to be able to do this with dragon naturally speaking (don’t remember if that was it’s exact name) 10 ish years ago

99catmaster a day ago

whisper.cpp has realtime capability. Been using it for 2 years at this point

LoganDark a day ago

Apple Dictation does this, or something similar, in my experience. Some apps (e.g. terminals in my experience) buffer the entire transcript but in most apps it's identical to typing as you speak. Have you tried it?

simonw a day ago

> Maintainer supported bindings in 4 Languages

Nice. Here's the Python one: https://github.com/handy-computer/transcribe.cpp/tree/main/b... - looks like it's not yet available as a binary wheel on PyPI with the dependency included (the library on PyPI right now uses ctypes to call a separately installed library) but that's planned for a future release.

sipjca a day ago

Yes, I’ve put a PR up on pypi for extra storage for CUDA but it has not been accepted yet afaik

If there’s any issues or improvements on the bindings I would love help to make the DX the best it can be

aomix a day ago

What good timing to spot this. I've been reading more and more people talk about bringing TTS into their prompting toolkit and wanted to give that a try. The idea of rambling brain dump into a doc -> edit pass -> send to the robot loop sounds appealing.

winterscott 11 hours ago

The numerical validation and WER testing are what stand out to me here. A lot of local ASR projects claim broad model support, but it is often difficult to know whether the converted models still match their reference implementations. Having one embeddable engine across Vulkan, Metal and CUDA, along with maintained language bindings, addresses a real distribution problem. How stable do you expect the C API and model format to be after v0.1? In particular, could an application eventually switch between different model families without needing model-specific preprocessing code?

bengotow a day ago

This is an incredible contribution to the community and it's just... one guy? I kept reading expecting a Series A funding announcement at the bottom.

It's a nice reminder: You can use AI to slop cannon at maximum speed, or you can use it to scale your ambitions and build something more rigorous and lasting than ever before.

I'd build Transcribe.cpp into the apps I maintain, but I feel like this functionality should (generally) be integrated into the OS or "everywhere" via an app like Handy.

sipjca a day ago

Hey, yep author and maintainer here! Certainly sponsors help and the wonderful community who donates to Handy as well! Mozilla AI was very helpful in getting this work off the ground. It was a pipe dream for me to build for Handy and they helped to sponsor me so I could make time to take this project seriously and get a v0.1.0 release out the door

I agree this should be everywhere and I hope to distribute libtranscribe some day properly so it is more a system library! It will take time to stabilize but I think we can get there

aarvin_roshin a day ago

Spot on:

> I think as we look forward to the future, more inference will start happening locally for one reason or the other. This brings the distribution story front and center. In order to have more applications running inference locally, we need to make running inference easier.

This makes these projects so much more trustworthy and easier to approach:

> Were any of the words here written using AI? Nope. They came from my mouth or my fingers.

boplicity a day ago

>This makes these projects so much more trustworthy and easier to approach:

>> Were any of the words here written using AI? Nope. They came from my mouth or my fingers.

I have to push back on this a bit, as I believe (quite strongly) that we're shaped by the tools we use; text-to-speech LLMs are still LLMs, and generally their mistakes are shaped by the expectations inherent in their training. This, in turn, shapes the words that appear on the screen. For those who regularly use them, you then learn which word sequences are likely to be accurately transcribed, and this definitively becomes part of your thinking process. Over time, the LLM becomes tangled into your thinking; the use of AI, even in this way, very much can and often does shape the resulting words.

eventualcomp a day ago

Isn't this like saying "my words are not really my own when I speak to my family, because I know my father is a non-native English speaker and hard of hearing so I try to use words which are well enunciated and are few in syllable count"?

iezepov a day ago

ChadNauseam a day ago

Parakeet, probably the most popular model for handy-style transcription, doesn't include a "text-to-speech LLMs" or any other form of LLM

leumon a day ago

Thank you! I found this to work much better then the old transcribe-rs lib. I updated my Offline Voice Input App to also use the new library and it's much faster now: https://github.com/notune/android_transcribe_app

ukuina a day ago

What's the easiest way to add speaker separation to this?

sipjca a day ago

Hey! It’s actually in progress right now, probably will come this week :)

simonw a day ago

Awesome! I found the in-progress diarization PR here: https://github.com/handy-computer/transcribe.cpp/pull/85

Looks like it's using IBM's Granite-Speech-4.1-2B-Plus https://huggingface.co/ibm-granite/granite-speech-4.1-2b-plu... and/or MOSS-Transcribe-Diarize https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize

sipjca a day ago

joshspankit 13 hours ago

Any chance you’ll be adding diarization of known speakers?

For example: from audio samples tagged with names

solarkraft a day ago

Oh, I like this! I’ve been looking into locally hosting a transcription API server and came away feeling pretty close to the problem statement. The things most frequently lacking were streaming support (which I’m so glad this has!) and the support for special words to boost during recognition (which I guess there’s some hope they might add???).

embedding-shape a day ago

> I’ve been looking into locally hosting a transcription API server

I've been hosting my own since whisper.cpp appeared on the scene, thrown up on a server with a 3090ti. Even if there is better/faster stuff out today, it just keeps on working without any issues, the weights are tiny and it's faster than I could need. This is basically what you need to get this working today:

    MODEL="/home/user/projects/ggml-org/whisper.cpp/models/ggml-large-v3-turbo.bin"
    WHISPER_SERVER_BIN="/home/user/projects/ggml-org/whisper.cpp/build/bin/whisper-server"
    "$WHISPER_SERVER_BIN" --model "$MODEL" --language en --host 127.0.0.1 --port 7812
Very simple stuff, throw it on some local homelab server and now you have a local transcription API :) Might need to play around with some of the inference parameters, but once you've locked them in, seems to work really well.

solarkraft a day ago

Does it support streaming? I find that this is the #1 thing missing from almost all implementations.

sipjca a day ago

word boosting will probably come on a much longer time horizon, but streaming is here!

I'm really hoping someone either contributes a good server example to the codebase (and is willing to help with issues) or use transcribe.cpp or the bindings to create a robust server in another language :) would be happy to link it from the main project directly as well

zaptheimpaler a day ago

Amazing, i've been looking for something like this and ended up doing transcription + diarization on a local server for now. Are you looking for contributions? Have you tried this one for diarization - https://huggingface.co/pyannote/speaker-diarization-communit... - it performed much better than Sortformer for me.

sipjca a day ago

Contributions are always welcome! There’s a WIP diarization PR rn, and after it’s merged would love to have support if it fits well into the interface. And if not would love to figure out a good interface for it

rcarmo a day ago

Yeah, diarization is the real feature these days. STT needs uniformization, but quality of diarization is what is setting personal solutions apart in this field.

sipjca a day ago

markisus a day ago

The post makes it seem like ONNX is CPU only. I've used ONNX runtime to run models on Nvidia GPUs. The runtime can even dispatch to TensorRT. I'm not sure what the performance is on Apple hardware so maybe that was the motivation for moving away from ONNX.

sipjca a day ago

TensorRT and CUDA is effectively the same speed as CPU for the speech to text models I was testing via ONNX at a huge binary bloat penalty. WGPU is hard to ship and also equivalent speed or slower. This may not be the case for LLM or other models but the runtimes did not seem well supported for what I needed to do. ONNX is incredibly well optimized for CPU, best in class even, but the other execution providers at least for STT seemed lacking.

I did this investigation before creating transcribe.cpp it would have been much more convenient and save me literal months of work. Happy to share the repo and binaries produced as well, but it was mostly throw away work to profile how to ship accelerated ONNX in Handy.

terhechte a day ago

I'm using this in one of my side projects, Emyn ( https://github.com/terhechte/Emyn ) a macOS virtual camera app for composing camera video, app windows, backgrounds, effects, notes, and captions into a polished live presentation feed.

It works very well, the integration is much easier than before, users have model choice. So happy that this exists!

sipjca 14 hours ago

Wow, it's amazing to hear that even though I released this so recently people are already using it properly! Thanks! Please let me know any issues you run into

apitman a day ago

Just wanted to say I started using Handy last week and I love it. It might single handedly cure my RSI. Well, hopefully double handedly.

jerieljan a day ago

Nice. I did transcriptions on a casual project before that went through something like this. Transcribing videos or audio files with Whisper? Very common. But having to swap it out with Qwen3 or a different family of ASR models? Oops, not as straightforward. For Qwen for example you gotta deal with the forced aligner or it won't be good as subtitles, and then gotta deal with some requirements and considerations if you want to make use of MLX on a Mac or something.

Will definitely check this out since it sounds like it eases through the pain of dealing with these.

kmfrk a day ago

Well this almost seems to be to good to be true. :)

I assume this is going to make maintaining SubtitleEdit a lot easier from now on, too: https://github.com/SubtitleEdit/subtitleedit/.

Anyone know a good Windows app that's just a window that transcribes - and translates - whatever goes through your output device, and not the microphone like most apps do?

alabhyajindal 20 hours ago

Congrats! I just tried Handy again which now uses transcribe.cpp and it works brilliantly. Love the streaming output from Parakeet Unified EN 0.6B. I remember using Handy about a year ago and it's amazing to see the improvements!

sorenjan a day ago

Why not include transcribe-cli in the release archives to make it easier to use for people that can't compile it themselves? I downloaded the Cuda version but it's only the dll files, I don't really want to have to deal with Cuda SDK, I doubt most people want to.

sipjca 14 hours ago

Right now I intend to maintain this as a library. The examples are just that, examples for programmers/agents. If someone in the community wants to step up to maintaining release binaries I will gladly have that support, it's just impossible to do as a sole maintainer

PalmPilotProMax 12 hours ago

Any plans of including the VAD step into this? When whisper.cpp added Silero it really smoothed out the UX.

sbinnee a day ago

I saw that metal is almost x10 faster than vulkan? Why so much gap?

sipjca a day ago

It very much depends on the hardware! An M4 max is being compared against a Ryzen 4750U with an integrated GPU!

The M4 max has probably 10x the compute and memory bandwidth hahaha

ctas a day ago

I'm using Handy on macOS and love it. Unfortunately, hotkeys still doesn't seem to work on Wayland, which make it unusable.

sipjca a day ago

Yeah I’m working on it, Linux is a big pain point especially Wayland

Once things are more or less ironed out on MacOS and Windows a lot of attention will be turned towards Linux

I know a lot of Linux PRs are open it just takes me so long to get around and test them. And often multiple different implementations trying to fix similar issues which is a lot of overhead sometimes

ctas a day ago

Really appreciate your work.

Is there any way people can help? From your last sentence, it sounds like another PR isn't it and the opposite might be needed. But would love to contribute with testing if helpful. I'm regularly jumping between XFCE, KDE, GNOME, Niri, etc..

sipjca a day ago

JesseHowell a day ago

Really cool that every model is actually tested for accuracy instead of just claiming it works, I think alot of 'we support everything' tools skip that step. How are you checking accuracy for models that don't have an obvious "official" version to compare against?

sipjca a day ago

Every model with open weights has some code which can be used to inference it. So we download the published weights and run against inference library they suggest, be it transformers, Nemo, etc

larnon 14 hours ago

Does the whisper model allow for entering context(which improves the accuracy greatly) as it does on Whisper.cpp?

sipjca 14 hours ago

yes

yjftsjthsd-h a day ago

So it's mostly intended to be a better replacement for whisper? Mostly? With better support for more models and maybe acceleration backends?

sipjca a day ago

More or less yes, for whisper.cpp, just trying to make local transcription more accessible to anyone building an app, etc

arikrahman a day ago

Excellent work, paired with the 500kb TTS model headlining today I can see the full stack coming together.

zuzululu a day ago

saw the demo its impressive but the audio was robotic

kelvinjps10 a day ago

Is there something but for transcribing what you watch like videos and not your microphone? Samsung has this in my phone and it's useful for language learning. (Thought is not that accurate)

lxe a day ago

What's the best local TTS model right now? I'm running parakeet on a mac which transcribes all my uh's and aahs. I'm running whisper on linux/cuda and I by far prefer that one over parakeet.

sipjca a day ago

Parakeet unified for me no longer does this and it’s also a streaming transcription model!

But the answer largely depends on you, the languages you speak, and personal preference. Whisper is still excellent and supported in transcribe.cpp

Cohere Transcribe is also excellent, but many of the new models are as well

TomGarden a day ago

I run the same, if you want try a simple filter post transcription to remove them, and while you're at it add some simple word replacements like 'cloud MD' to 'CLAUDE.MD'

jv22222 a day ago

> parakeet on a mac which transcribes all my uh's and aahs

You should be able to fix this by playing with the mic speech floor. It happens when to much ambient stuff slurps in.

It's actually gaslighting you, you don't say that many ums and ahs ;)

copypirate a day ago

Excellent work CJ

sipjca a day ago

Thanks :)

l-albertovich a day ago

Thanks CJ, you've put some pretty cool things out there!

0xnyn a day ago

handy has been invaluable in my workflow, and having a fast, local, c++-based transcription library with first-party ts bindings is incredibleee

tysm for shipping this, keep up the great work OP

vardalab 18 hours ago

One thing I find that's missing a lot or at least I haven't come across other than commercial offerings like AquaVoice is a decent injected technical vocabulary so that the initial transcript requires minimum cleanup afterwards. Because I mostly use these tools to essentially ramble at the command line with coding agents. So there's a lot of technical terms that don't translate well. Like OpenBao comes out as open bowel sometimes,lol. That necessitates significant cleanup prompt or background text available to the cleanup llm, usually in the form of screenshot or something that gets converted to text but that in turn requires good hw for speed to be almost imperceptible. For example m5 max turns cleanup into a noticeable delay while 5090 is decent.

Only way I have found that's relatively easy to inject technical vocab is to use whisper, but limited, I think to about 220 or so tokens. Whisper has sort of like a priming prompt where one can put in a bunch of technical words and it will try to recognize those. But again, that's limited to small number tokens. And that limits one use a relatively slow, by today's standards, whisper.cpp.

I benchmarked it across a bunch of different hardware that I have available, and Whisper gives decent performance as far as speed goes only on a pretty top-end GPU, such as a 5090 or 4070, like for example on Strix Halo, it's still relatively slow for longer transcriptions because I prefer just a stream of consciousness ramblings for minutes and then that being transcribed and cleaned up versus short sentences. So in that scenario something like 5090 really is good because the cleanup prompt runs fast using usually Qwen 3.6 MOE model. Whisper on 4070 itself is about 0.7 seconds for two or three minute transcription. So the total wait time for a three-minute transcription is roughly a second, or a little bit more than a second, so totally acceptable. But it does take decent hardware, and it grows to be double that on if running totally local. Well, in my case, it's all local, but it's my own hardware all over the place, but truly running on laptop, it's much faster using Parakeet, but then the cleanup is the bottleneck.

Anyway, it's just my experience messing around with this for the last year. I did start using AquaVoice, but their speed was exceptional, and tech vocab was exceptional, but they would have some annoying delays occasionally, and I didn't like paying the money and sending sensitive topics and screenshots into the cloud, and I had hardware, so my local solution is basically almost as good as commercial one. But I think they train their own model. So what I'm doing is I collect all the samples of my transcriptions, and I am slowly building my own data set that hopefully at some point when I get energy I will find some way to fine tune something.

SamPentz a day ago

Is there a way to add speaker identification easily?

sipjca a day ago

Hey! It’s actually in progress right now, probably will come this week :)

semiquaver a day ago

That would be Diarize.cpp, not Transcribe.cpp.

paweladamczuk a day ago

I've been using this one for a week for local transcription, working pretty well so far

dostick a day ago

Does this support filtering of “umm”,”err”, “ugh”, or that is nit yet possible with open source models?

sipjca a day ago

Not in the library itself, it’s pure inference. Some models have this trained out of them anyhow. Otherwise this is a post processing task which is not really inference

dostick a day ago

So it looks like something that app would be doing, or run another model over the output to smoothen out and remove these things?

sipjca a day ago

fenix1851 21 hours ago

thanks man! Been searching for solution like yours :)

hackrmn a day ago

Is transcription a form of _inference_ though? I mean I see the word being thrown around and I understand what it means (or at least I think I do) in context of LLMs doing the thing that they do -- intelligently predict the next token, but do speech-to-text models do that?

yorwba a day ago

Speech-to-text models predict the next token of text from the preceding tokens of text and the current tokens of speech.

hackrmn 15 hours ago

Thanks, I did some learning and it fell more into place.

luciana1u a day ago

three separate people in this thread independently remembered Dragon NaturallySpeaking and I think that is the funniest possible review of the state of speech recognition in 2026

bazzingadev a day ago

Hey, thanks for this.

ilaksh a day ago

I don't suppose this works in the browser?

sipjca a day ago

Out of the box no probably not, but if people are interested there’s probably ways forward

_medihack_ a day ago

Absolutely interested

kzyxx11 a day ago

Excellent work

JeremyHerrman a day ago

Another happy user of Handy here!

After seeing so many *subscription based* transcription apps all wrapping *open source models*, finding Handy was a real delight and I'm happy to see the author keep on building!

tlamponi a day ago

Has anybody experience with using this with strong dialects, like e.g. bavarian-family (German) based ones? Or other languages one too, as I'd figure basic behavior and approaches to improve detection of such is often similar in principle for dialect style variants of a language.

I mean, I naturally should try myself, and plan to do so, but slightly lower on my free time priority list and I figured someone else might have explored this already.

shade a day ago

Nice - I'm definitely going to take a look at this. I've built my own cross-platform (Mac/Win/Linux) live captioning app on top of Nemotron, and it works well but dealing with ONNX is kind of annoying. With this having Rust support (I built it on Rust/Tauri) it should be a pretty solid candidate; I'll have to see if I can find a Silero VAD implementation that doesn't depend on ONNX, or maybe I'll see if the clankers can migrate it for me.

nohup2 a day ago

Have you published your app? I would love to take a look and test it

shade a day ago

I have not formally published it, but it's open source: https://github.com/edmistond/larmindon and https://github.com/edmistond/larmindon-core - you'll want to clone them into the same root directory.

Right now it only supports languages supported by parakeet-rs and Nemotron (so... English only as far as I'm aware) and you'll need the ONNX version of Nemotron: https://huggingface.co/altunenes/parakeet-rs/tree/main/nemot...

The first run experience isn't great, you'll need to download all the files from the model, start the app, and then go to settings and configure the model directory. It runs well on Mac and Windows; I haven't tested it on Linux in a couple of months since my Linux install is out of commission currently.

diimdeep a day ago

Congrats on delivering good value to the people. I have used transcribe.cpp a few weeks ago to do near realtime offline stt on a 10 year old phone, writing simple adhoc app for my use case, it's crazy what is happening right now.

sipjca a day ago

Ha amazing, love to hear it

nojvek a day ago

I use Handy everyday. It’s a great project. Thank you for making offline ASR work great on modern machine.

wolvoleo a day ago

Looks interesting, I'll give it a try. Though I'm really happy with faster-whisper on a GPU.

zuzululu a day ago

would love to see a demo handy is fantastic although its still behind the frontier models

therealpygon a day ago

Pretty sure I saw Handy using it; if you have the latest version, you’re probably already demoing it.

sipjca a day ago

Yep the latest version has support! Virtually all of the SOTA open models are supported by Handy including the streaming ones like

Nemotron Streaming

Parakeet Unified

Voxtral Mini Realtime

If something you want is not supported, open an issue on transcribe.cpp!

loufe a day ago

author of the blogpost is the maintainer of Handy, so almost guaranteed!

zuzululu a day ago

I installed it but I don't think I see the streaming transcriptions. I do think the transcription is a bit faster. I am using the latest version.

qntmfred a day ago

ktosobcy 17 hours ago

Yet another happy user of Handy - one of the best applications out there, kudos! <3

primaprashant a day ago

Handy is an amazing cross-platform app for dictation from the author. There are other awesome open-source dictation tools as well like native macOS ones. You do not need SaaS subscription in this day and age for transcription.

I maintain this list of all the best open-source ones in this awesome-style GitHub repo. People looking for open-source dictation tools, hope you find something that works for you here:

https://github.com/primaprashant/awesome-voice-typing

rubidium a day ago

Any that also support translation? How much harder/ easier of a problem is local translation compared to transcription?

primaprashant a day ago

If you're talking about translated text, then that should be super easy. Most of these dictation tool support post-processing with LLM to remove filler words, fix punctuation, etc. I'd imagine you can change the system prompt for the post-processing step to do the translation instead, and you'd get translated text.

rubidium a day ago