All notes

Same words, different notes

Changing only the speaker labels on a court transcript moved one speaker’s share of my pipeline’s notes by 40 points. What that does and doesn’t show.

During my research internship at the Applied AI Laboratory, HEC Lausanne, I built a fully local pipeline that turns court audio into notes. In one mode a local language model reads a numbered transcript and only picks which sentences go into the notes; it cannot change a word. So every note is verbatim, and every accuracy metric I had said the pipeline was working. None of them asked whether the notes were fair to the people speaking.

The probe

Keep the transcript byte-for-byte identical, change only the speaker labels, run the selection again, and compare. Three conditions:

ConditionWhat changes
ControlNothing. The identical transcript is selected again, to see how much the selection moves on its own.
AnonymisedEvery label becomes SPEAKER_XX, so the model can no longer tell the speakers apart.
PermutedLabels are shifted by one, so the same labels sit on different people.

For each condition I measured every speaker’s share of the selected lines and how far it moved from the original run. A shift only counts as evidence if it is at least 5 points and at least twice whatever the control moved.

The first run was confounded

My first anonymised run used a plain “SPEAKER” label. It was shorter than the real labels, so more lines fit in each window the model reads, and the transcript was split at different points: windows of 155, 175, 187, 176 and 61 lines against the original 151, 178, 176, 173 and 84. That run changed two things at once, so its 10.7-point shift could not be blamed on the labels.

The fix was a label of exactly the same width, SPEAKER_XX, and a rule: any condition whose windows don’t match the original is reported but never counted.

The result

ConditionOverlap with original selectionLargest shift in one speaker’s shareCounts as evidence
Control100%0 pointsThis is the floor
Anonymised18%40.4 pointsYes
Permuted41%3.6 pointsNo

Removing who-said-what replaced most of the selection and moved one speaker’s share of the notes by 40.4 points. The words were identical.

Word error rate, diarization error, verbatim rate and the judge all score these two sets of notes the same. Only the probe saw the difference.

What it does and doesn’t show

  • It shows that what gets selected depends on the labels, not just the words. No accuracy metric in the pipeline checks for that.
  • It doesn’t show that the model favours particular people. Swapping who is who moved shares by only 3.6 points, below the bar. The more likely reading is that the model uses the labels to follow the structure of the hearing, who is asking and who is answering, and selects differently when that structure disappears.
  • It is one case: a 62-minute Supreme Court argument with 10 speakers. That makes it a finding to chase, not a result to generalise.

Why this matters for diarization

If removing the labels can move a speaker’s share of the notes by 40 points, then the labels a diarizer produces are not just a detail of the transcript. Diarization error on this audio was 7.9%, and overlapping speech, where two people talk at once, is one of the hardest cases for a diarizer. A turn given to the wrong speaker could change what ends up in the notes.

Measuring how diarization errors, especially in overlapping speech, carry through into what a summary selects is the next thing I want to study.