How we build the best on-device meeting assistant
AnythingLLM v1.17.0 rebuilds the Meeting Assistant on NVIDIA's Nemotron 3 Diarization and Moondream's Parakeet Redux. A 3x smaller download and speaker labels you can trust, all on your Mac or Windows PC.
October 2026A meeting transcript is only half useful if you cannot tell who said what. "We will ship it Friday" means something very different coming from your engineering lead than from a customer.
As of AnythingLLM Desktop 1.17.0, the Meeting Assistant runs on two brand new open models. Nemotron 3 Diarization from NVIDIA figures out who is speaking. Parakeet Redux from Moondream, built on NVIDIA's Parakeet, writes down what they said. Both run entirely on your Mac or Windows PC, on the hardware you already own.
The result is a Meeting Assistant that is about three times lighter to download and finally gets speakers right. It is still free, and your audio still never leaves your computer.
What changed
- A much smaller download. The models the Meeting Assistant needs went from about 2.6 GB to about 0.8 GB. The speech model alone went from 2.55 GB to 0.41 GB.
- Speaker labels you can trust. Up to 8 speakers per meeting. On every recording in our test set, on every platform we ship, the new engine found the right number of people. The old one looks like random guessing by comparison.
- Less memory. Peak memory during processing dropped by roughly 1 to 1.5 GB depending on the machine.
- Put your GPU to work. On Windows, set the Meeting Assistant's Hardware Runtime to GPU and it runs on your NVIDIA, AMD, or Intel graphics card. On an NVIDIA RTX 4070, speaker identification runs about 200 times faster than real time, and an hour of audio is transcribed in under 30 seconds. It is quick on the CPU too, around 70 times real time for speakers on a modern desktop processor.
- Same price. Zero. No plan, no minute cap, no bot in your call.
Speakers were always the hard part
When we first shipped the Meeting Assistant, transcription was great and speaker identification was, honestly, not. By default we separated your voice from everyone else on the call, and full speaker identification was an opt in feature for good reason.
The problem was the model. For years, the only realistic option for identifying speakers locally was pyannote. We are genuinely grateful it existed. It was free, it had an MIT licensed version, and without it local speaker labels would not have been possible at all. But the pipeline behind it slices audio into small pieces, turns each one into a voice fingerprint, and then tries to cluster the fingerprints into people. That works on clean, well behaved audio. Real meetings are not that. Similar voices, quick back and forth, people talking over each other, and hour long calls all trip it up.
In our own testing, a two person, nine minute call came back as one speaker. An hour long meeting with three people came back as one speaker plus noise. On another machine, the same three person meeting came back as six.
Since it also added a lot of extra processing you essentially were spinning your fans for what could, sometimes, be a coin flip. It was something, but it was never the end goal.
Nemotron 3 Diarization from NVIDIA
Nemotron 3 Diarization, which NVIDIA released in September, takes a different approach. It is built on NVIDIA's Sortformer architecture. There is no fingerprinting and no clustering step. One model listens to the audio and directly answers the question "who is talking right now?" every 10 milliseconds, for up to 8 people at once, including when they talk over each other.
The clever part is how it names people. Sortformer labels speakers in the order they first speak. The first voice in the meeting is always speaker 1, the next new voice is speaker 2, and so on. A speaker cache lets the model remember who is who across silences and long stretches of audio. As NVIDIA puts it in their write up, arrival time ordering "makes the model's generic speaker labels stable and removes the need to solve a new speaker permutation for every chunk." For a meeting tool, that is exactly what you want. Speaker 2 at minute 3 is still speaker 2 at minute 53.
The numbers back it up. On DIHARD III, one of the hardest public diarization benchmarks, Nemotron 3 has a diarization error rate of 12.73%. pyannote 3.1 reports 21.7% on the same benchmark. That is about 41% fewer errors. It also ranked first of twelve systems on the independent VoiceArena Diarization Bench. And in our own tests it got the speaker count right on every recording, on Windows, on Apple Silicon, on Intel Macs, and on Snapdragon.
It is also small. About 100 million parameters, and released under the OpenMDW license, which is permissive, ready for commercial use, and ungated. No access request, no token, no strings. That matters enormously for an app like ours that ships to millions of desktops.
Parakeet Redux from Moondream
Transcription in the Meeting Assistant has run on NVIDIA's excellent Parakeet TDT 0.6B v3 since we moved off Whisper. It is fast, accurate, and handles 25 languages with punctuation and capitalization. The only downside was weight. The full precision version we shipped was a 2.55 GB download.
The team at Moondream took Parakeet v3 and made Parakeet Redux, a 1.58 bit version of the same model. Same architecture, same tokenizer, same 25 languages. The difference is that every weight in the encoder is now one of three values: -1, 0, or +1. If you know about the Bonsai Model Family from PrismML it's similar to that.
That sounds like it should wreck the model. It does not. Normally every weight is a 16 or 32 bit number. With only three possible values, each weight needs about 1.58 bits, so the model gets dramatically smaller. And multiplying by -1, 0, or +1 is just subtract, skip, or add, which is cheap for any CPU. According to Moondream's benchmarks, Redux is within about three tenths of a point of the original on English word error rate, and it actually scores better than the original on the multilingual FLEURS set and on long form talks.
In practice, we could not tell the difference in our meetings. Comparing transcripts word by word against the full model, the differences were mostly style, like "going to" versus "gonna".
How we brought them into AnythingLLM
The Meeting Assistant runs on our own transcription engine, written in Rust. Earlier this year we rewrote it from Python to cut startup time and drop a heavy PyTorch install from every user's machine. It runs models through ONNX Runtime, which gives us one engine that can use Apple Silicon, Intel and AMD CPUs, Snapdragon, and any DirectX 12 GPU on Windows through DirectML.
Both models landed in the same week, and both were in our engine within days. That speed is a testament to how good open model releases have become, and to the Rust community. The excellent parakeet-rs crate added support for Nemotron 3 almost immediately. Swapping our old diarization pipeline for Nemotron 3 took about 260 lines of our own code, and we got to delete a vendored C++ audio feature build and our custom clustering logic along the way. Parakeet Redux keeps the exact same model layout as Parakeet v3, so it needed no engine changes at all.
A few things took real work.
Running ternary weights on ONNX Runtime. ONNX Runtime does not have a native ternary format yet. So we repacked each of Redux's 264 ternary matrix multiplications into ONNX Runtime's 4 bit format, storing -1, 0, and +1 as fixed codes around a zero point. The output matches the original PyTorch model to within 0.00000025. That is why our version is 414 MB instead of Moondream's 178 MB, and it is still six times smaller than what it replaced.
The right build for every chip. On Apple Silicon and Snapdragon, a version of the encoder that uses the chip's 8 bit integer math is about 1.6 times faster. On Intel and AMD Windows PCs, the default version wins. ONNX Runtime cannot switch between them at runtime, so the Meeting Assistant downloads the right one for your machine automatically.
Who said each word. Redux only ends a sentence at punctuation, and in a lively meeting that can be over a minute of speech. If you label a whole sentence with one speaker, short interjections disappear into whoever was talking the most. So we assign a speaker to every single word. A word only changes hands if another speaker clearly overlaps it, so a quick "mm-hmm" does not hijack someone else's sentence. In a 20 minute, three person meeting, the quietest person went from 19 seconds of attributed speech to 98 seconds. Nemotron had heard them speak for 94.
Using every core, without fighting itself. Transcription and speaker identification can run side by side. On a big chip like a 12 core Snapdragon X Elite that is 20 to 27% faster. On a smaller laptop, running both at once actually slowed things down, so the engine now looks at your machine and decides.
Built on open models
None of this would be possible without the people who build these models and share them openly. NVIDIA's speech team built Parakeet and Nemotron 3 Diarization and released them on Hugging Face with permissive licenses. Moondream took Parakeet and made it dramatically lighter, then gave that away too. The Rust community made it easy to run all of it natively.
Our job is to take the best models in the world and make them feel effortless on the computer you already own. One app, on your device, private by default, free for everyone.
Try it
- Download AnythingLLM Desktop for macOS or Windows, or update to 1.17.0.
- Open Meeting Assistant from the sidebar and record a meeting, or drop in an old recording.
- The models download once, about 0.8 GB. After that, everything runs on your machine.
The Meeting Assistant is available on macOS (Apple Silicon and Intel) and Windows (x64 and ARM64). GPU acceleration on Windows uses DirectML and works with NVIDIA, AMD, and Intel graphics. Benchmarks above are from our own test recordings unless noted, and your speed will vary with your hardware. Diarization error rates are as published by NVIDIA and pyannote.
Learn more