Before getting into the tools themselves, I want to say something about what this review is for. Transcription AI is genuinely one of the most defensible uses of AI in journalism — it automates a mechanical task that has nothing to do with the quality of the reporting. But "defensible" does not mean "without issues," and the issues with transcription AI are specific enough that they are worth going through carefully rather than assuming that because the use case is legitimate, the tool can be trusted without scrutiny.
Context
Why journalists consider transcription tools
The starting point is time. An hour of interview audio, transcribed by a human, takes somewhere between two and three hours of careful listening and typing. For journalists who interview regularly — several times a week, sometimes more — that is an enormous proportion of available working time spent on a task that adds no reporting value whatsoever. The transcript is not the story. It is the raw material the story comes from. Every hour spent transcribing is an hour not spent on the analysis, the writing, or the additional reporting that a story needs.
The second pressure is accuracy. Human transcription is not infallible. Writers who transcribe their own recordings know how easy it is to fill in a gap you could not quite hear with what you expected the person to say — a subtle error that feels like recollection but is actually reconstruction. AI transcription has its own accuracy problems, which I will get to, but the comparison is not with a perfect human transcript. It is with the kind of transcript a journalist actually produces under deadline pressure.
So the pull toward these tools is real and reasonable. The question is what you are actually getting.
What works
What these tools do reasonably well for journalism
The speed is real. Both Otter.ai and Whisper produce a transcript of an hour-long interview in a matter of minutes. This is not a marginal improvement over human transcription. It is a qualitative change in what the post-interview workflow looks like. A journalist who previously set aside half a day for transcription can now have a searchable, reviewable text within the time it takes to make a coffee.
Speaker identification works adequately in straightforward conditions. Otter.ai, in particular, attempts to separate and label different speakers in a conversation. When there are two clearly distinct voices in a clean recording, this works reasonably well. It is not perfect — it will sometimes merge two speakers or split one speaker into two — but it provides a usable scaffold that is faster to correct than to produce from scratch.
The searchable transcript is genuinely useful. Being able to search a transcript for a specific word or phrase — rather than scrubbing through audio — changes how a journalist moves through interview material. If a source mentioned a specific name or date and you need to find exactly what they said around it, a searchable transcript recovers that in seconds. This is a real workflow improvement that has nothing to do with AI quality and everything to do with having the text in a digital form.
Whisper's accuracy on clear audio is high. In good conditions — two people speaking clearly in a quiet room, recorded on a decent microphone — Whisper (which underlies many commercial transcription tools, including some competitors to Otter.ai) produces transcripts with relatively low word error rates on standard English speech. For a journalist's standard office or sit-down interview, it is reliable enough to be the starting point for quote verification rather than a document requiring wholesale correction.
Required reading
Where these tools cause real problems in journalism work
This is the section most tool reviews skip or minimise. I am not going to do that.
Accuracy degrades sharply in real interview conditions
The conditions under which transcription AI performs well — clear audio, quiet environment, standard accent, two speakers — are often not the conditions in which journalism happens. A street interview. A noisy conference room. A source with a strong regional accent or a quiet voice. A phone recording. A group conversation. In these conditions, accuracy drops significantly. The transcript still saves time over manual transcription, but the error rate is high enough that the journalist must read every line against the audio rather than spot-checking. The efficiency gain shrinks accordingly.
This matters more than it might seem. The errors that transcription AI makes in difficult conditions are not evenly distributed. They tend to cluster around proper nouns — names, place names, technical terms — which are exactly the details that matter most for factual accuracy in journalism. A source's name. The name of an organisation. A figure they cited. The AI will produce something phonetically similar that reads as plausible but is wrong.
The confident wrong word problem
AI transcription does not express uncertainty. It does not leave a blank where it could not hear something. It fills in. And what it fills in sounds like what was said. A human transcriber who cannot hear a word leaves a gap — [inaudible] — or puts something in brackets to flag uncertainty. AI transcription produces a complete-looking sentence that may contain a word the speaker never said.
This is the specific failure mode that creates corrections. A journalist uses the transcript to pull a quote. The quote contains a word the AI substituted. The journalist does not go back to the audio because the transcript looks complete. The substituted word goes to print.
Otter.ai's cloud processing raises source confidentiality questions
Otter.ai processes audio in the cloud. This means that when you upload a recording, the audio passes through Otter's servers. For most interviews, this is an acceptable tradeoff. For interviews with confidential sources — people who have agreed to speak on the condition that their identity and the content of the conversation is protected — uploading the recording to a third-party cloud service is a confidentiality risk that is not hypothetical. The question is not whether Otter.ai is trustworthy. It is whether uploading the recording to any third-party service is consistent with the confidentiality obligations the journalist accepted when the source agreed to speak.
Whisper can be run locally — on your own computer, without any audio leaving your device — which addresses this concern. The setup requires some technical confidence and is not the same experience as the Otter.ai web interface. But for confidential source material, local processing is the appropriate standard.
Speaker labels are suggestions, not facts
Otter's speaker identification labels are not reliable enough to be treated as ground truth. In any transcript that will be used for quotes, every attributed statement needs to be verified against the audio to confirm that the label is correct. This sounds like a minor issue until a journalist pulls a quote attributed to Speaker A that was actually said by Speaker B — which, in a political or business interview context, can be the difference between accurate reporting and a serious error.
When it helps
Situations where these tools make sense
I want to be specific here, because "use it for transcription" is not helpful advice on its own.
Standard sit-down interviews with a single source, recorded clearly, in English. This is the scenario where accuracy is high enough and the efficiency gain is large enough that the tool is straightforwardly worth using. The journalist still verifies every quote against the audio before publication, but the transcript is a reliable starting point.
Long recordings where the journalist needs to search for specific content. A three-hour background briefing. A lengthy expert conversation being used for context rather than direct quotes. The transcript's value is navigational — finding the relevant section quickly — rather than providing publishable quotes directly.
Recordings that are not going to be quoted directly. Background interviews. Notes calls. Conversations that are informing the reporting without being cited. The accuracy standard for non-quoted content is lower because the specific wording does not go to print.
High-volume routine interviews where the subject matter is familiar. A beat reporter who interviews the same kinds of sources repeatedly — on the same technical subject, using the same vocabulary — will find that transcription AI performs better on their recordings than on unfamiliar subjects, because the vocabulary it needs to handle is more predictable.
Clear limits
Situations where journalists should avoid or be very careful
Confidential source recordings should not go through cloud transcription services. This is a firm line for me. The source agreed to speak under confidentiality conditions. Uploading their recording to a third-party server changes the risk profile of that conversation without the source's knowledge or consent. Use Whisper locally or transcribe manually.
Poor-quality audio recordings should not be trusted as primary transcripts. If you cannot hear clearly what was said, neither can the AI. The transcript will look complete and plausible but contain substitutions you cannot identify without listening carefully. For poor-quality recordings, manual transcription — slow and frustrating as it is — gives you the information you need: the gaps and uncertainties that the AI will hide from you.
Direct quotes should never be pulled from a transcript without audio verification. This applies to both tools without exception. The transcript is a finding aid, not a quotation source. Before any statement attributed to a named source appears in published work, the journalist should have listened to that statement in the recording and confirmed that the transcript accurately captures what was said. This is not a best practice. It is the minimum standard.
Recordings in languages other than English are unreliable without specific testing. Whisper has multilingual capability, but accuracy varies significantly by language and decreases further with accents, code-switching, and technical vocabulary in that language. Do not assume that because the tool handles English well, it handles other languages at the same standard.
Professional practice
How to use these tools without losing professional judgment
The core principle is straightforward, even if following it consistently requires discipline: the transcript is a working document, not a record. It helps you find your way through the material. It is not the source of what a person said. The recording is.
In practice, this means building audio verification into the workflow rather than treating it as an optional final step. Some journalists I have spoken with do this by maintaining a rule: before a quote is typed into the story draft, they listen to it in the recording. Not skim the transcript. Listen. This takes perhaps thirty seconds per quote and catches substitution errors before they become corrections.
For Otter.ai specifically: review the speaker labels at the start of the transcript and correct them before you use the document. If you cannot identify which label belongs to which speaker from the first few exchanges, go back to the audio and establish it before you proceed. Correcting speaker attribution at the beginning takes two minutes. Discovering a misattributed quote after publication takes much longer to repair.
For sensitive material, ask yourself before uploading: does this recording contain anything that would embarrass me, embarrass my source, or create legal or confidentiality complications if it were accessed by a third party? If yes, process it locally with Whisper or transcribe it yourself.
Corrections
Common misconceptions journalists have about transcription AI
"The AI got it wrong so it must have been a bad recording." Not necessarily. Transcription AI fails on good-quality recordings when the vocabulary is unusual, the accent is unfamiliar to the model, or the speaker talks quickly or quietly. The quality of the audio is one factor. It is not the only one.
"I can just use the transcript for quotes — it's fast enough and close enough." This is the most dangerous misconception. The AI produces text that looks authoritative. It is not. Every quote that goes to print needs audio verification. There is no situation where "close enough" is acceptable for a direct quote attributed to a named person.
"Otter.ai is HIPAA compliant so it's fine for sensitive content." HIPAA compliance addresses a specific set of healthcare data protections. It does not address journalistic source confidentiality obligations, which are a different standard with different requirements. Compliance with one regulatory framework does not substitute for independent assessment of confidentiality obligations under professional ethics standards.
"Whisper is free so it's probably worse than Otter.ai." Whisper (the underlying model) is the basis for a significant portion of commercial transcription tools, including some competitor products to Otter.ai. The accuracy of a tool depends on how Whisper is implemented and what audio processing is applied, not on whether you pay for a branded interface. Running Whisper locally may produce better or comparable results to cloud services for many use cases, with the added benefit of keeping audio on your own device.
"AI transcription is replacing the need to listen to interviews." This is worth addressing directly because the temptation exists. The transcript gives you the words. It does not give you the tone, the hesitation, the moment a source's voice changed when they said something, the laugh that contextualises a statement differently than the bare text suggests. Journalists who work from transcripts only miss things. Listening to interviews — even if partially, even if at increased speed — is a reporting skill that the transcript does not substitute for.
Closing
Final advice for journalists using transcription AI
These tools are worth using. The time savings are real, the workflow improvement is genuine, and the use case — automating the mechanical conversion of audio to text — is about as straightforward a candidate for AI assistance as journalism offers. None of that changes the fact that the tools have specific, predictable failure modes that create professional risk if they are not understood and managed.
The journalist who uses Otter.ai or Whisper well is the one who treats the transcript as a navigational aid — something that helps them move through the material faster — while maintaining the listening and verification habits that protect them from the errors the AI will inevitably make. That combination is efficient and defensible. The transcript alone, without those habits, is a liability.
If you are new to these tools, start with a recording where accuracy is not critical — a background briefing, a note-taking conversation, something you will not be quoting directly — and spend the time to understand where the tool fails on your specific type of content before you rely on it for material that will go to print.
That is about as simple as the advice gets: know the failure mode, build the verification step, don't skip it under deadline pressure. The tool will save you time. The verification step is what prevents it from costing you more than the time you saved.