One tap records a meeting. About a minute later a speaker-attributed summary and a translation into the listener's language arrive by email, on iPhone, Android, Wear OS and the web. Korean and Chinese recordings pass through a domain-adapted recognition stage built from recent ASR research.
Scroll sideways on a narrow screen to see the whole diagram.
One tap records on iPhone, Android, Wear OS or the web, even with the screen locked. Uploads queue in the background with idempotency keys, so a retry never processes a meeting twice. Optional participant names travel with the audio.
Files are identified by content, not extension, then converted to 16 kHz mono and trimmed of long silences. There is deliberately no denoising: recent studies show it makes modern ASR less accurate.
The use case's vocabulary and the participant names bias Soniox, which also separates speakers. Korean and Chinese text then gets dictionary sound-alike fixes and optional LLM span edits, each checked by the phonetic verifier.
gpt-oss-120b on Groq writes key points, decisions, action items and open questions as strict JSON. Speaker labels tie tasks to people, and an owner is kept only if the transcript or the typed names support it. Notes are translated into the chosen language.
A bilingual email goes only to the address the sign-in provider verified, and notes sync to every device. Audio is deleted right after processing; transcripts leave the server after 30 days.
A harness runs every engine through every correction chain on 14 synthesized meetings and scores character error rate, facts, owners and hallucinations. Its findings chose the default engine and caught two bugs before release.