[05]entry 5 of 29
Ultron
this entry is written twice, in full
- 1.in brief
real-time AI overlay for interviews & meetings
- 2.in general
A desktop AI copilot for live conversations, invisible to the screen share.
- 3.in particular
Dual-stream audio, 6 STT providers, multi-LLM routing, and local RAG over your meeting history.
- 1.in brief
a live AI helper for interviews and meetings
- 2.in general
A desktop assistant for conversations as they happen, which does not appear when you share your screen.
- 3.in particular
It hears both sides, has six transcription services and several AI models behind it, and can search everything that was said in past meetings.
context
An Electron1 desktop app that sits as a stealth overlay, invisible to screen-capture and sharing via a macOS window-layer trick. It captures both sides of a conversation and lets you summon AI assistance with a hotkey, mid-meeting.
the problem
Most "AI meeting" tools only hear your mic. To actually be useful the AI needs full conversational context (both sides of the audio) plus low latency, provider resilience, and a UI that never blocks while tokens stream.
what i built
- 1.
Dual-stream capture: microphone + system loopback transcribed concurrently for full context
- 2.
Six STT providers including on-device ONNX Whisper9 (no cloud); multi-provider LLM routing with a fallback chain and circuit breakers
- 3.
Meeting RAG: live transcript indexed into sqlite-vec7 for semantic search over past sessions, with structured summaries on session end
- 4.
Screenshot vision analysis and a phone-mirror that streams transcript + answers to any device on local Wi-Fi
- 5.
All API keys stored with Electron1 safeStorage (OS-level encryption), never in plaintext
outcome
The hard parts are exactly where you would expect them: real-time audio on dedicated threads bridged to an async runtime, token-by-token markdown streaming over IPC, and a ~3k-line settings hub driving every provider, keybinding and mode.
Six transcription providers behind a fallback chain and a circuit breaker, because whichever one you picked will go down mid-sentence.
context
A desktop app that floats above everything else and stays invisible to screen sharing and recording. It listens to both halves of a conversation and lets you summon AI help with a keyboard shortcut, in the middle of a meeting.
the problem
Most meeting AI tools only hear your own microphone, which is half a conversation and therefore not much use. To be genuinely useful it has to hear both sides, answer quickly, keep working when a service goes down, and never freeze the window while it is thinking.
what i built
- 1.
It records your microphone and the computer’s own sound at the same time, so it has the whole conversation rather than your half of it.
- 2.
Six different speech-to-text services, one of which runs entirely on your own machine with nothing leaving it, and several AI models arranged so that when one fails the next takes over without you noticing.
- 3.
Everything said is indexed as it is said, so past meetings can be searched by what was meant rather than by exact words, with a written summary when a session ends.
- 4.
It can look at a screenshot and explain what is on it, and mirror the transcript and the answers to your phone over the local network.
- 5.
Every key you give it is stored using the operating system’s own encryption rather than as readable text in a file.
outcome
The hard parts are exactly where you would expect them: catching live audio without stuttering, writing an answer out word by word as it arrives rather than waiting for the whole thing, and a settings screen large enough to drive every service, shortcut and mode in the app.
Six transcription services behind an automatic fallback, because whichever one you picked will go down mid-sentence.