Building an Offline AI That Turns Meetings Into Minutes: Six Lessons
At a recent hackathon I built a system that takes the audio of a hospital meeting and produces structured minutes — decisions, action items, owners and deadlines — then e-mails them to the right distribution list. The hard constraint was privacy: hospital meetings can mention patients and staff, so nothing was allowed to leave the machine. Speech recognition, the language model, the routing and even the mail server all run locally.
The models turned out to be the easy part. Almost everything I learned was about what surrounds them. Here are the six lessons that stuck.
The pipeline
browser (upload or record)
↓
FastAPI server → job queue
→ ① speech-to-text faster-whisper large-v3 (GPU or CPU)
→ ② extraction local LLM via Ollama (qwen2.5:7b), JSON-schema output
→ ③ documents minutes as DOCX / HTML / Markdown
→ ④ routing n8n workflow: meeting type → distribution list → e-mail
(fallback: send directly over local SMTP if n8n is down)
Everything listens on 127.0.0.1. A local mail server with a web inbox stands in for the hospital’s mail system during development.
1. Speech recognition assumes one language at a time
Real meetings in a bilingual setting don’t stick to one language. People switch mid-sentence, and medical terms often stay in English. A stock Whisper model picks a single language for each 30-second window, so a switch in the middle of a window comes out translated or transliterated. Worse, it sometimes labelled our main language as a related language and transcribed it that way.
What worked was making language a per-phrase decision:
- Cut the audio at pauses with a voice-activity detector (Silero VAD), merging very short fragments so language identification stays reliable.
- Encode each chunk once, batched on the GPU, and run language identification on the encoder output, restricted to the languages we actually expect.
- Decode each chunk with its own language token plus a short domain prompt that shows mixed-language text and keeps medical terms verbatim, using a small glossary.
- When two languages are almost equally likely, decode in both and keep the more confident result.
Every transcript line then carries a language tag, which the language model uses later.
2. Silence is where speech models hallucinate
Speech models are trained to always produce text, so silence and noise get “transcribed” too: repetition loops, invented sentences, and stock phrases such as video subtitle credits. I added explicit checks — a compression-ratio test that catches repetition loops, detection of silence decoded as speech, and a list of known artefacts. A suspicious chunk is retried at a higher temperature and dropped if it still looks wrong. Without these checks, the minutes would occasionally contain action items nobody ever said.
3. Ask for structure, not a summary
A free-text summary is pleasant to read and impossible to act on. Instead, the local model fills a fixed JSON schema, enforced by constrained decoding in Ollama: decisions, action items (task, owner, the deadline as spoken, the deadline resolved to a date, priority, and the timestamp of the evidence), participants, topics and open issues. Relative deadlines like “by Friday” or “by the end of the month” are resolved against the meeting date. The evidence timestamp matters most: anyone can jump to the moment in the recording and check that an action item was really agreed.
4. Small hardware forces good design
I developed on a laptop GPU with 4 GB of memory, far below the 16 GB the system is meant for. That limitation shaped two decisions I would now make anyway. Speech recognition and the language model take turns: the Whisper model is unloaded before the LLM runs. And long meetings are processed map-reduce style — notes per slice of the transcript, then a consolidation pass — so an 8 k-token context is enough even for long meetings.
5. Save every intermediate result
Each job gets its own folder, and every stage writes its output before the next one starts: the original audio, the transcript as it streams in, the final transcript in several formats, the raw model outputs, the minutes, and the delivery result. Re-running a job skips the stages that already finished. This made debugging fast — when the minutes looked wrong I could see whether the transcript or the extraction was at fault — and it makes the system easier to trust.
6. “Offline” is a checklist, not a switch
Pulling models from the internet is the obvious leak, but there are many quiet ones. Model libraries are forced into offline mode at runtime, every service binds to localhost by default, the workflow engine’s telemetry, version checks and template downloads are switched off, the mail server’s update checks and reverse-DNS lookups are disabled, and the web interface loads no fonts or scripts from a CDN. Only the one-time setup step needs internet.
How I would evaluate it next
A demo that works on a few recordings is not the same as a system you can trust, and that is where statistics comes in. The next step would be to measure it properly:
- Transcription quality per language — word error rate split by language tag, since the mixed-language parts are where errors concentrate.
- Extraction quality — precision, recall and F1 for action items against minutes written by a person, with bootstrap confidence intervals, because a handful of meetings gives very wide error bars.
- Comparing two models fairly — when trying a different speech or language model, score both on the same meetings and use a paired bootstrap or McNemar test rather than eyeballing two averages.
- Agreement with people — how often a human and the model agree on what counts as a decision, measured with Cohen’s kappa.
- How many meetings to test — sized in advance with an eval sample-size calculation.
What this taught me about learning AI
Going in, I expected the challenge to be the models. In practice, a capable speech model and a capable language model were a download away. The real work was everything around them: deciding which language each phrase is in, catching made-up text, forcing the output into a shape people can act on, fitting it all onto modest hardware, and keeping the data where it belongs. That is also why I keep building tools for evaluating AI systems: knowing whether a system actually works is the part that is still easy to skip.