How to transcribe a research interview

Weeve author portrait

Dylan de Heer

How to transcribe a research interview

Transcribing a research interview is four decisions and one long correction pass. The decisions are whether you need a full transcript at all, how much of the speech you are going to preserve, who or what produces the first draft, and what you will write in your methods section afterwards. The correction pass is the part nobody warns you about, and it is where most of the hours go.

The mechanics are not difficult. What makes this hard is that transcription is an analytical act dressed up as an administrative one. Julia Bailey makes the point in Family Practice: turning a recording into written form is an interpretive process, and therefore the first step in analysing the data rather than the chore before it. (Bailey, J., 2008, Family Practice 25(2), 127 to 131.)

This covers the decisions in order, with a worked excerpt, a methods-section template you can adapt, and an honest account of where automatic transcription falls over on real interview audio.

First, decide whether you need a full transcript

The most useful question about transcribing research interviews is whether you need to do it at all, and it is not a fringe one. Halcomb and Davidson asked it in Applied Nursing Research in 2006 and proposed a structured alternative to verbatim transcripts. Researchers reach the same place independently all the time. A team on r/academia describes dropping transcription for a spreadsheet: one row per participant, columns for the themes they were coding against, and short quoted extracts pulled from the audio at the point they were needed.

That is not laziness, it is a legitimate method choice, and it is the right one more often than the literature admits. Full verbatim transcription is expensive. Bailey puts it at a minimum of three hours of work per hour of talk, rising towards ten where fine detail is being marked. Correcting a machine draft is faster but not fast: Eftekhari, transcribing 41 interviews with speech recognition, reports 1.5 to 3.5 hours of accuracy checking per interview. For a study with 20 interviews of 45 minutes each, that difference is weeks.

Full transcription earns its cost when:

  • Your analysis is on the language itself. Discourse analysis, conversation analysis, narrative work, anything where how something was said carries the finding.

  • You need to demonstrate an audit trail. Some ethics committees and most doctoral examiners want to see the data you coded from.

  • More than one person is coding. Shared coding needs a shared artefact, and audio is not one.

  • You will be quoting extensively in a publication. Getting a quote wrong in print is worse than the transcription hours.

Summary notes with timestamped extracts are enough when you are doing straightforward thematic analysis on a small sample, you are the only coder, and the recordings will stay accessible to you throughout. Decide this before you start rather than 12 interviews in.

Naturalistic or denaturalised

If you are transcribing, this is the decision that shapes every page. The terms come from the methods literature and they matter more than the software. One warning before the definitions: the labels are used in opposite directions by different authors. This piece follows Oliver, Serovich and Mason (2005), where naturalistic means maximally detailed. Bucholtz (2000) uses naturalised for the opposite, text tidied towards written conventions. Say which sense you mean and you are safe either way.

Naturalistic transcription keeps the speech as it happened. Every "um", every false start, every three-second pause, laughter, overlapping talk, the word repeated twice while the participant thinks. It is slower to produce and harder to read, and it is the only option if your analysis involves how people say things rather than only what they said.

Denaturalised transcription removes the noise of speech and keeps the meaning. Filled pauses go, stammers are cleaned up, grammar is lightly repaired where the speaker clearly meant something else. It reads like prose. It is what most thematic and content analysis actually needs, and it is what most people mean when they say "clean verbatim".

Neither is more rigorous than the other. What is not defensible is drifting between them without noticing, which is what happens when you clean up one participant because they were hard to read and leave another untouched.


A middle position most qualitative researchers land on: denaturalised as the default, with pauses, laughter, and emphasis marked only where they change the reading of a quote you intend to use. Say so in your methods section and it is fine.

The five steps

1. Record with the transcription in mind. An external microphone placed between you and the participant will save more correction time than any software choice. Room noise, air conditioning, and a café table are what break automatic transcription, not accents. Record a 30-second test and listen to it on headphones before the interview proper starts.

2. Produce a first draft. Either by hand, or with software, or by paying a service. See the section below on which of these fits, and where the audio file goes in each case, which matters if your consent form promised the recording would not be shared.

3. Correct against the audio, in full. This is not optional and it is not skimming. Play the recording at 0.75 speed with the draft open and fix as you go. On a machine draft you are looking for misheard technical terms, wrong speaker attribution at turn boundaries, and confidently-wrong sentences that read fluently and say the opposite of what the participant said. That last category is the dangerous one, because nothing in the text flags it.

4. Anonymise. Replace names, employers, place names, and anything else identifying with consistent pseudonyms or bracketed labels. Keep the key in a separate file, stored separately, per whatever your ethics approval says. Do this on the transcript, and do it before the transcript goes anywhere near a shared drive or a coding tool.

5. Format and label. Speaker labels, timestamps at a consistent interval, a header block with the participant pseudonym, date, duration, and interview number. Boring, and the thing you will be grateful for when you are trying to find a quote 14 months later.

A worked excerpt

Here is the same 20 seconds of audio in both styles, with the notation spelled out.

Naturalistic:

INT:   And how did that, um, how did that land with the team?
P07:   [laughs] Yeah. Yeah, it was, (2.0) I mean it was fine? It was
       fine but I think people were, they were a bit- nobody said
       anything at the time. [pause] Nobody said anything for weeks.
INT:                                                    [Right.]
INT:   And how did that, um, how did that land with the team?
P07:   [laughs] Yeah. Yeah, it was, (2.0) I mean it was fine? It was
       fine but I think people were, they were a bit- nobody said
       anything at the time. [pause] Nobody said anything for weeks.
INT:                                                    [Right.]
INT:   And how did that, um, how did that land with the team?
P07:   [laughs] Yeah. Yeah, it was, (2.0) I mean it was fine? It was
       fine but I think people were, they were a bit- nobody said
       anything at the time. [pause] Nobody said anything for weeks.
INT:                                                    [Right.]

Denaturalised:

INT:  How did that land with the team?
P07:  It was fine, but I think people were a bit... nobody said
      anything at the time. Nobody said anything for weeks

INT:  How did that land with the team?
P07:  It was fine, but I think people were a bit... nobody said
      anything at the time. Nobody said anything for weeks

INT:  How did that land with the team?
P07:  It was fine, but I think people were a bit... nobody said
      anything at the time. Nobody said anything for weeks

The conventions used above are a working hybrid, not a single system. Timed pauses and cut-off words follow the Jefferson conventions used in conversation analysis. The bracketed non-verbals and the timestamped inaudible marker come from ordinary transcription practice; strict Jefferson uses double parentheses for non-verbals, ((laughs)), and reserves square brackets for overlap. If your analysis is conversation-analytic, use full Jefferson. If it is not, pick one set and keep it. One trap either way: an ellipsis inside a quoted extract conventionally means you have cut material, so many researchers reserve ... for omissions and mark trailing off some other way.

  • (2.0) a timed pause in seconds, [pause] an untimed one

  • [laughs], [sighs] non-verbal sounds in square brackets

  • word- a cut-off or self-interrupted word

  • [Right.] on an indented line, speech overlapping the line above

  • ... trailing off, in denaturalised transcripts only

  • [inaudible 00:14:32] with a timestamp, never a guess

Whichever set you use, define it in a short legend at the top of your transcripts and use it consistently. Examiners care far more about consistency than about which convention you picked.

What to write in your methods section

The one rule the methods literature is unanimous on is that you have to state how the transcription was produced. A 2019 review in RAUSP Management Journal is titled after the failure: papers say the interviews were transcribed and give the reader nothing else. This is now more important than it used to be, because "transcribed by the researcher" and "transcribed by a speech recognition model and corrected by the researcher" are genuinely different data provenance claims, and a reviewer is entitled to know which one they are reading.

Something in this shape covers it:

Interviews were audio-recorded and transcribed denaturalised, with pauses and non-verbal sounds marked only where relevant to the interpretation of extracts used. First-pass transcripts were generated using [tool and model], processed [locally on the researcher's device / on the provider's servers under the data processing agreement at reference X], then checked in full against the original audio by the researcher and corrected. Identifying details were replaced with pseudonyms at the transcription stage.

Adapt the wording, keep the four facts: the style, the tool, where the audio was processed, and who did the correction.

Where automatic transcription actually fails

Speech recognition on clean single-speaker audio is now very good. Research interviews are frequently not clean single-speaker audio, and the failures cluster:

Overlapping speech. Two people talking at once is where every system degrades most, and research interviews are full of it because interviewers make encouraging noises. Expect to rebuild these passages by ear.

Speaker attribution at turn boundaries. Diarisation gets the speakers right in the middle of a turn and wrong at the edges, so the first few words of a response often get attached to the interviewer. This is quiet and easy to miss and it will misattribute a quote.

Domain vocabulary. Your field's terminology, drug names, instrument names, local acronyms, and organisation names all get replaced with plausible common words. Build a find-and-replace list after the first transcript and run it on the rest.

Non-native speakers and strong regional accents. Word error rates rise, and more importantly they rise unevenly across your participants, which introduces a bias into your data that is invisible at the point of coding. The hallucination work below found the same unevenness, which means the bias is not only in what the model mishears but in what it invents.

Confident fabrication. Some models will smooth a garbled passage into a fluent sentence rather than flag it. A human transcriber writes [inaudible]. This is measured, not folklore: Koenecke and colleagues found roughly 1% of Whisper transcriptions contained whole hallucinated phrases that were never spoken, and that those hallucinations clustered in speakers with speech impairments. It is the single strongest argument for the full correction pass, and it is worth knowing about the model Harvard recommends and the model class most tools now run.

Eftekhari's 2024 methods paper in the European Journal of Cardiovascular Nursing is worth reading on this, and it lands roughly where practice has: fast first draft, mandatory human check.


Where the audio goes

This is a methods question and an ethics question rather than a software preference, and it is worth separating from accuracy.

Most transcription services upload your recording to their servers and process it there. For an interview with a public official about published policy, that is unremarkable. For interviews with patients, children, undocumented workers, or employees speaking about their employer, it is a disclosure to a third party that your participant information sheet very likely did not mention.

Harvard's library guidance for qualitative researchers lists Whisper, OpenAI's open speech model, as the option for sensitive recordings, specifically because it can be run on your own machine. Its guide is blunt about the distinction: the Whisper API, where the audio goes to OpenAI and comes back, is "NOT appropriate for sensitive data", so the model has to be downloaded and run locally. Harvard points Mac users at Aiko or MacWhisper for that, and Windows users at aTrain, all of which wrap the model in an interface. That route is honest, it is cheap or free, and it is worth trying before you buy anything.

Weeve sits in the same local-processing category, with recording, transcription and summarisation in one Mac app rather than a transcription step you then take somewhere else. The audio file does not leave the device, so the recording, the transcript and the summary all stay where your ethics approval says they should.

Being straight about the limits, because they matter for research use. Weeve is Mac only and needs Apple Silicon on macOS 14 or later. It is not a qualitative coding tool, so MAXQDA, NVivo, ATLAS.ti or Dovetail still do your analysis and you export the transcript to them. It does not offer a human-verified accuracy tier, so if you need certified verbatim for legal or archival deposit, a service like Rev or Verbit is the correct answer. And Weeve does not currently hold SOC 2 or ISO 27001 certification. Some ethics committees ask for one by name, and if yours does, that is a hard stop. The free Starter plan also covers ten recordings a month, which will not carry a study of twenty interviews on its own.

It will not remove the correction pass. Nothing does.

FAQ

How long does it take to transcribe a one-hour interview? Bailey's estimate for hand transcription is three to ten hours of work per hour of audio, the upper end being naturalistic transcription with detailed notation. Correcting an automatic draft against the audio is quicker, though Eftekhari's study of 41 interviews still needed 1.5 to 3.5 hours per interview. Budget the checking time, not the typing time.

Should I transcribe my interviews myself or pay someone? Transcribing your own interviews is genuinely useful analytically, because the repeated close listening surfaces themes that reading a finished transcript does not. It is also the first thing to outsource when the sample size makes it impossible. If you outsource, check what your ethics approval says about third-party processing and whether a confidentiality agreement with the transcriber is required.

What is the difference between verbatim and clean verbatim? Verbatim, sometimes called naturalistic, keeps everything: filled pauses, repetitions, false starts, non-verbal sounds. Clean verbatim, or denaturalised, removes those and keeps the meaning intact. Clean verbatim is the more common default for thematic analysis.

How do I format an interview transcript in APA style? APA does not prescribe a transcript layout. What it governs is how you cite and present extracts in the text: quotations from participants are formatted like other quotations, with a participant pseudonym rather than a citation, and long extracts set as block quotes. Institutions usually have their own transcript template, so check that first.

Can I use AI to transcribe research interviews? Yes, and it is now standard practice, with two conditions. Check the whole draft against the audio rather than sampling it, and state in your methods section which tool produced the draft and where the audio was processed. The processing location is the part that interacts with your consent form.

Do I need to keep the original audio after transcribing? Usually yes, until the retention point specified in your ethics approval, because the audio is the primary data and the transcript is an interpretation of it. Your approval will also say when it has to be destroyed, and that instruction is as binding as the one to keep it.

A transcript is not a recording of an interview. It is the first argument you make about what the interview meant.

If your participants were promised the recording stays with you, Weeve's free Starter plan keeps the audio and the transcript on your Mac, ten recordings a month. If you are not on a Mac, the browser transcriber turns a recording you already have into text inside the tab, up to an hour at a time.