Altranscribe
A real-time transcription and translation app for Windows and Android. It turns microphone input, system audio, or audio and video files into transcripts and shows floating bilingual captions.
- TYPE
- Personal Tool
- YEAR
- 2026
- ROLE
- Solo
- STACK
- Dart · Flutter · C++ · Kotlin
What it's for
When you’re listening to a language you don’t know well, it’s hard to catch whole sentences quickly and completely. Even when you can just about follow the gist, it’s easy to miss important details, and the longer you listen, the harder it gets to keep up. It’s even more frustrating with a less common language you’ve never studied at all, when you need to understand what’s being said anyway.
There have been all sorts of tools for this for years, but using them tends to be a letdown: recognition is inaccurate, translations are stiff, and most are tied to a pricey monthly subscription. Worse, all the audio and text have to be uploaded to the vendor’s cloud servers, with no guarantee of privacy at all.
I made Altranscribe to be done with those compromises for good. The name comes from “Alternative Transcribe”, and it’s meant as a complete alternative. High-quality transcription usually needs a GPU, so I designed it to work over the local network, with a Windows PC as the host and an Android phone as the client. That puts the graphics card sitting idle on your desk to good use. The phone can also do lightweight transcription on its own, fully offline.
As for the stack, it uses Flutter, like Canvas Quiz Exporter. I picked it mainly for its strong cross-platform support and the elegant feel of Material Design. Altranscribe is the older of the two, though, and my first real Flutter project.
Transcription engines and deployment
Altranscribe isn’t tied to any single ecosystem or cloud vendor. You can assign and combine recognition backends however you like:
- Local Whisper: built on whisper.cpp, running on the CPU or with NVIDIA GPU acceleration on Windows. It’s the best option when you want high-quality recognition that never leaves your computer.
- Local Nemotron: true streaming recognition (streaming ASR) on the CPU through sherpa-onnx, with each caption done in about a second. Its overhead is very low, so it runs smoothly on a phone too.
- Local network (Windows host): turn on sharing on the computer with one click to get a QR code, scan it with the phone to pair, and the phone borrows the computer’s processing power and models.
- Cloud APIs: works with OpenAI and Google Gemini. The project set out to run everything locally, but the cloud APIs stay in for convenience.
Translation, titles and summaries can be set up just as freely. They can run on a local Ollama or a host on your network, or go to OpenAI, Gemini, Anthropic or any third-party service with an OpenAI-compatible API. Recognition and translation are fully independent of each other, so the phone can run the lightweight Nemotron model for the original text and hand translation and summaries off to a computer on the same network. Since 0.7.0, the phone can also transcribe entirely by itself.
Floating captions and working with transcripts
-
Floating captions: show only the original, only the translation, or both, and adjust the font, size, layout and background opacity. Pause, resume and Stop & save are each one click away in the caption window, and closing the window doesn’t interrupt the transcription pipeline running in the background.
-
Archive: past transcripts get full-text search, quick copying and renaming. They export straight to TXT, Markdown or a standalone interactive web page with an audio player and timestamps you can click to jump to.
-
Term corrections: when you transcribe audio and video files, an LLM can make proper nouns, unusual names and technical terms consistent and fix recognition errors it’s highly confident about. You can also enter preferred spellings ahead of time. The raw recognition results are kept in full, along with a before-and-after of each change, so you can trace changes back to the original.
Audio preprocessing pipeline
Recognition accuracy depends largely on the audio the model receives. To keep this low-level layer consistent and efficient, all the audio processing is written in C++, and Windows and Android share the same implementation:
-
Front-end noise suppression and gain: RNNoise handles noise suppression, and adaptive automatic gain control (AGC) boosts faint and far-field audio. Each can be switched on or off as needed.
-
Isolation and decoupling: every audio source keeps its own resampling, noise suppression, gain and streaming segmentation state, isolated from the others. Segmentation runs strictly on capture time, so delays later in transcription and model inference don’t affect it.
How to use it
- 01
Download and install
On Windows, unzip the archive and run Altranscribe.exe, keeping the folder intact. To transcribe audio and video files, you also need FFmpeg on your PATH. On Android 10 or later, just install the APK.
- 02
Choose how speech is recognized
Under Settings → Speech models, pick local Whisper, local Nemotron, a Windows host on your local network, or OpenAI / Gemini. For translations and summaries, also set up an LLM under Translation & LLM.
- 03
Start transcribing
Pick the microphone, system audio or both, set the source and target languages, and start. For audio and video files, switch to Files; you can choose several at once and they're processed as a queue.
- 04
Stop & save
The transcript moves to the Library while the last sentences, translations and summary finish in the background. From there you can search and copy it, or export it as TXT, Markdown or a web page.