September 4, 2026
The Baseline here is a whisper model. Except for a
i dont want to go into details here. The problem is more about live transcription, since we face the following problem: If we send audio-chunks that are too long, it's no longer live. If they are shorter, they get inacurate, and the transcription suffers. The solution is straight forward: Start with shorter chunks, and then further on, send the results so far as context.
This idea is refined in this paper (1). There is an implementation posted on github. The main point there is that it runs locally, which was a requirement for me anyways because of privacy concerns.
For my project, i built a bridge server that implements the basic whisper api, so this can serve as a drop-in replacement:
001def do_POST(self):002 if self.path.startswith("/inference"):003 self._handle_inference()004 elif self.path.startswith("/reset"):005 self.server.session.reset()006 logger.info("Session reset.")007 self._send_json({"status": "reset"})008 else:009 self.send_error(404)010011def _handle_inference(self):012 try:013 length = int(self.headers.get("Content-Length", 0))014 body = self.rfile.read(length)015 content_type = self.headers.get("Content-Type", "")016 file_bytes = extract_multipart_file(body, content_type)017 if file_bytes is None:018 self._send_json({"error": "no 'file' field in multipart request"}, status=400)019 return020 audio = decode_wav_to_16k_mono(file_bytes)021 text = self.server.session.process_chunk(audio)022 logger.debug("chunk (%d samples) -> %r", len(audio), text)023 self._send_json({"text": text})024 except Exception as e:025 logger.exception("inference failed")026 self._send_json({"error": str(e)}, status=500)
Line 21 just calls the python API of the paper's implementation, with a session wrapped around it:
001class Session:002 """Wraps one persistent SimulStreaming online-ASR processor."""003004 def __init__(self, asr, online):005 self.asr = asr006 self.online = online007 self.lock = threading.Lock()008009 def process_chunk(self, audio: np.ndarray) -> str:010 with self.lock:011 self.online.insert_audio_chunk(audio)012 out = self.online.process_iter()013 return out.get("text", "") if out else ""014015 def reset(self):016 with self.lock:017 self.online.init()
Importantly, i added a reset api: self.server.session.reset(): If the inital transcription is in the wrong language, that same language will be used for the rest of the session. Then, the service will translate. (This is a known issue with whisper!, not because SimulStreaming also implements a translation service, which i did not install.)
To correct this, the reset endpoint will clear the context sometimes.
Macháček, Dominik, and Peter Polák. 2025. “Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025.” Pp. 389–98 in Proceedings of the 22nd International Conference on Spoken Language Translation (IWSLT 2025), edited by E. Salesky, M. Federico, and A. Anastasopoulos. Vienna, Austria (in-person and online): Association for Computational Linguistics. ↩