All work

Legal audio

A fully local pipeline that turns court audio into speaker-attributed notes, where every line links back to the moment it was said. Nothing leaves the machine.

When
May–Jul 2026
Role
Research intern (remote)
Where
Applied AI Laboratory, HEC Lausanne
Stack
Python, Whisper, pyannote, local LLMs
The app in demo mode: note modes on the left, verbatim notes with speaker and timestamp citations in the middle, and the transcript on the right with the cited line highlighted.

Why it exists

Recordings of legal proceedings are often privileged, so sending them to a cloud API is not an option. And a note about what a court said is only useful if it is easy to check against the record.

Most local meeting-notes tools feed a Whisper transcript to a language model and inherit every hallucination the model makes. For a document someone may rely on as an account of a hearing, that is the whole problem.

What I built

  • A local pipeline. Whisper large-v3-turbo for speech recognition, pyannote for speaker diarization, and a 4-bit quantized language model, all on a consumer 8 GB GPU.
  • Notes that cannot be made up. In model-selected extractive mode the LLM reads a numbered transcript and may only return integer sentence ids, enforced by a grammar constraint at decode time. The note text is then copied from the transcript by id, so it is byte-for-byte the record.
  • Confidentiality you can verify. A run marked privileged refuses any non-loopback endpoint before inference starts, and every run writes a manifest of hashes and endpoints that can be re-checked later.
  • A one-click check. In the app, every note is a button. Click it and the transcript scrolls to the source sentence and the audio seeks to that moment.

Results

word error rate on 62 minutes of gold Supreme Court audio (Oyez)
6.9%
diarization error rate on the same audio
7.9%
verbatim rate in the extractive mode, verified
100%
how much more often legally important words are misheard than average words
3.0×

The evaluation harness covers WER, DER, a checklist-style LLM judge in the spirit of CheckEval, and planted prompt-injection tests against that judge.

I also built a speaker-label bias probe. It re-runs note selection on transcripts whose words are identical but whose speaker labels are anonymised or shuffled, and compares the result against a resampling noise floor. Relabelling alone shifted one speaker’s share of the notes by 40.4 points, a failure that WER, DER and verbatim rate cannot see.

What the numbers don’t say

Supreme Court oral argument is a floor, not a typical case: the recordings are unusually clean and the speakers unusually clear. Depositions and trial-court audio will be harder. The code is private.