gremlin.group / blog/ how-phonic-knows-who-is-speaking

PhonicAI

How Phonic Works Out Who Is Speaking

A transcript of a meeting is far more useful when it shows who said what. Here is how speaker identification works in plain English, how Phonic does it, and where it still gets things wrong.

A transcript of a meeting is only half useful if it is one long block of text. Knowing that someone agreed to send the proposal by Friday matters less than knowing who agreed to it.

Working out who is speaking is a separate problem from working out what is being said. It has its own name, its own models and its own failure modes. This post explains how it works in general, how Phonic handles it, and where it falls short.

Two different jobs

Transcription turns speech into words. Speaker identification, usually called diarisation, answers a different question: which parts of the recording belong to the same voice?

Diarisation does not know who anyone is. It listens for the qualities that make one voice sound different from another, such as pitch, tone and the way someone speaks, and groups stretches of audio that sound alike. The output is a set of labels: this section is one speaker, that section is another, and this later section is the first speaker again.

Putting a real name on each label is a further step. The audio alone cannot say that the second voice belongs to Sarah from finance. Something else has to supply that information.

What clustering means

One common way to do the grouping is clustering. The recording is cut into short pieces, each piece is turned into a kind of numerical fingerprint of the voice, and pieces with similar fingerprints are put in the same group. Each group becomes a speaker.

It is a bit like sorting a pile of handwritten notes by handwriting. Most of the time it works. Two people with very different handwriting are easy to separate. Two people whose handwriting looks alike end up in the same pile.

Voices behave the same way. Two speakers of similar age, accent and pitch can produce fingerprints that are close enough to be treated as one person. Short interjections, people talking over each other, poor microphones and background noise all make it harder. Every system that groups voices this way runs into this.

How Phonic does it

Speaker identification is part of Phonic Pro, the £2.99 in-app purchase. The free version of Phonic still records and transcribes. Pro adds the speakers, along with system audio recording, Meeting mode and the Insights tab. The Phonic Pro article lists everything it includes.

Speakers are identified on the device, in the same way the transcription is. Each speaker starts with a placeholder: Person A, Person B, and so on.

Phonic then tries to put names on them. A speaker is named automatically when they introduce themselves or when someone addresses them by name. If the recording belongs to a calendar meeting, Phonic uses the meeting’s attendees as well. That helps because the people in a meeting are usually the people on the invite.

When the automatic naming gets it wrong or misses someone, you can rename speakers yourself in each recording. If there is a calendar meeting, its attendees are offered as suggestions, so you are choosing from the people who were invited rather than starting from scratch. The speakers article covers the details.

Using the calendar needs Calendar permission. The Phonic privacy policy explains what each permission is used for.

The models

Diarisation needs its own models, separate from the speech model used for transcription. Phonic uses models from FluidAudio. They are about 120 MB, downloaded the first time you use speaker identification and reused after that.

If the first time is a long meeting on a weak connection, that download is an unwelcome wait. You can fetch the models in advance from Settings, under Speaker Models. The models article has the steps. Doing it on Wi-Fi before an important meeting is the sensible approach.

Under an hour, and over

Phonic uses two approaches depending on the length of the recording.

Recordings up to an hour use a model called Sortformer. Longer recordings use clustering alone, the technique described above. Clustering can merge similar voices, so on a long recording two people who sound alike may come out as a single speaker.

In practice, a typical meeting of under an hour is handled by Sortformer. A two-hour workshop is handled by clustering, and its speaker labels deserve a closer look before you rely on them.

Where it still struggles

Even on shorter recordings, some situations make diarisation harder for any system.

Two people who sound alike are the hardest case. When people talk over each other, the audio holds two voices at once and is hard to assign cleanly. Someone who only says “yes” or “agreed” gives the system very little to go on. And a single phone in the middle of a large room picks up some voices far better than others.

Naming has its own gap. If nobody introduces themselves or uses anyone’s name, and there is no calendar meeting, speakers stay as Person A and Person B until you rename them.

None of this makes speaker identification useless. It means the labels are a strong first draft rather than a verified record. For most meetings, that is enough. For anything where attribution matters, such as who agreed to what, it is worth checking.

Why it is worth having anyway

Imperfect speaker labels still make a transcript far more useful. You can scan for what a particular person said. The Insights tab, also part of Pro, shows action items with owners and talk time per speaker.

Renaming a speaker once in a recording is much quicker than reading through a transcript with no speakers at all and working out who said each line from memory.

The practical takeaway

Diarisation groups voices. It does not recognise people. Phonic adds names from what is said in the meeting and from the calendar, and lets you correct anything it gets wrong.

Download the speaker models before your first long meeting. Check the labels when attribution matters, especially on recordings over an hour, where clustering can merge similar voices.

More about Phonic, including what is free and what comes with Pro, is on the Phonic page.

M Written by Michael, Co-Founder Cloud architect and co-founder of Gremlin Group. Spends most of his time designing AWS infrastructure and writing about cloud architecture, cost optimisation, and DevOps.

Got a process that eats your week?

Tell us about it. If AI is the right fix we will build it, and if a spreadsheet would do, we will say so.