Glossary

Automatic speech recognitionASR

Automatic speech recognition (ASR) is the machine-learning field, and the class of models, that turns audio of human speech into text. ASR models are the engine inside every dictation and transcription product.

Modern ASR models are neural networks trained on large amounts of recorded speech. Well-known families include Whisper (OpenAI, open source), Parakeet (NVIDIA, open source), and the proprietary engines inside cloud services from Google, Apple, and others.

ASR quality is usually reported as word error rate (WER) on benchmark audio. Real-world quality also depends on your microphone, room, accent, and vocabulary, which is why per-user dictionaries and domain vocabularies matter as much as the headline model.

In Flit

Flit runs open ASR models (Parakeet, Whisper) plus Apple’s on-device recognizer, all locally on Apple silicon.

How Flit processes speech →

Your voice was always faster.

Not out yet. $29 USD once when it is, with no account, no subscription and nothing uploaded.

No spam, no sharing. One email when Flit ships, and you can leave in one click.

macOS 14 or later · Apple silicon