During my research internship at the Applied AI Laboratory, HEC Lausanne, I built a fully local pipeline that turns court audio into notes. In one mode a local language model reads a numbered transcript and only picks which sentences go into the notes; it cannot change a word. So every note is verbatim, and every accuracy metric I had said the pipeline was working. None of them asked whether the notes were fair to the people speaking.
The probe
Keep the transcript byte-for-byte identical, change only the speaker labels, run the selection again, and compare. Three conditions:
| Condition | What changes |
|---|---|
| Control | Nothing. The identical transcript is selected again, to see how much the selection moves on its own. |
| Anonymised | Every label becomes SPEAKER_XX, so the model can no longer tell the speakers apart. |
| Permuted | Labels are shifted by one, so the same labels sit on different people. |
For each condition I measured every speaker’s share of the selected lines and how far it moved from the original run. A shift only counts as evidence if it is at least 5 points and at least twice whatever the control moved.
The first run was confounded
My first anonymised run used a plain “SPEAKER” label. It was shorter than the real labels, so more lines fit in each window the model reads, and the transcript was split at different points: windows of 155, 175, 187, 176 and 61 lines against the original 151, 178, 176, 173 and 84. That run changed two things at once, so its 10.7-point shift could not be blamed on the labels.
The fix was a label of exactly the same width, SPEAKER_XX, and a rule: any condition whose windows don’t match the original is reported but never counted.
The result
| Condition | Overlap with original selection | Largest shift in one speaker’s share | Counts as evidence |
|---|---|---|---|
| Control | 100% | 0 points | This is the floor |
| Anonymised | 18% | 40.4 points | Yes |
| Permuted | 41% | 3.6 points | No |
Removing who-said-what replaced most of the selection and moved one speaker’s share of the notes by 40.4 points. The words were identical.
Word error rate, diarization error, verbatim rate and the judge all score these two sets of notes the same. Only the probe saw the difference.
What it does and doesn’t show
- It shows that what gets selected depends on the labels, not just the words. No accuracy metric in the pipeline checks for that.
- It doesn’t show that the model favours particular people. Swapping who is who moved shares by only 3.6 points, below the bar. The more likely reading is that the model uses the labels to follow the structure of the hearing, who is asking and who is answering, and selects differently when that structure disappears.
- It is one case: a 62-minute Supreme Court argument with 10 speakers. That makes it a finding to chase, not a result to generalise.
Why this matters for diarization
If removing the labels can move a speaker’s share of the notes by 40 points, then the labels a diarizer produces are not just a detail of the transcript. Diarization error on this audio was 7.9%, and overlapping speech, where two people talk at once, is one of the hardest cases for a diarizer. A turn given to the wrong speaker could change what ends up in the notes.
Measuring how diarization errors, especially in overlapping speech, carry through into what a summary selects is the next thing I want to study.