Research

Interview transcription software: where does the audio go?

Weeve author portrait

Dylan de Heer

Interview transcription software: where does the audio go?

Most interview transcription software works by uploading your recording to a server. That is fine for a product demo and a problem for an interview, because an interview is usually the most sensitive recording anyone makes, and the person who agreed to it did so on terms you set.

This covers what actually happens to an interview file in each kind of tool, how transcription works when nothing is uploaded, and where a local tool is the wrong choice.

The question the tools do not answer

Search for interview transcription software and near the top you will find something that is not a vendor. It is a thread of researchers asking each other what to use, and the most upvoted reply begins by telling the asker to check with their university, because there are ethical concerns about where the audio goes.

That is the real question, and almost no product page answers it. They answer accuracy, price, export formats and turnaround. Where the file lives is left to the privacy policy.

For a lot of interview work, it is the only question that matters. A research participant covered by an ethics submission. A source speaking on condition of anonymity. A candidate being honest about why they left their last job. A client describing a dispute that is still live. In each case the value of the conversation depends on the person believing it stays where they left it.

What happens to the file, by tool type

Cloud transcription services. You upload, the audio sits on the vendor's servers, and a transcript comes back. Retention is set by a policy you did not write. Some vendors train on customer audio unless you find the setting and switch it off. This is most of the market, and it is genuinely convenient.

Human transcription services. A machine produces a first draft and a person checks it line by line. Accuracy is the highest available, the file still goes to the vendor's servers, and a human being outside your institution has heard the recording. For court use or published quotes this is often the correct trade.

Qualitative analysis suites. Tools built for research workflows, with coding and thematic analysis attached. They transcribe too, but that transcription runs in the cloud, so the audio still goes to a server.

On-device transcription. The model runs on your own machine and the audio never leaves it. Slower to set up, no collaboration features, and nothing to disclose on an ethics form beyond the machine itself.

None of these is wrong. They are different answers to different constraints, and the useful move is knowing which constraint you actually have.


How transcription works when nothing is uploaded

The practical objection to local transcription used to be that it was not good enough. Small models on a laptop produced text you had to rewrite. That changed when open speech models got good and Apple Silicon made it realistic to run one on a Mac you already own.

Weeve works this way. The speech model downloads once on first run, then transcription happens on the Mac. You can import an interview you already recorded, on a phone, a dictaphone or from a video call, or capture one live. Speaker labels are applied on the device. Once the models are installed the transcription itself runs with the network off, which is the simplest way to prove to yourself that the audio is not being sent anywhere.

Being straight about the limits, because they matter for this work:

  • Mac only, and it needs Apple Silicon with macOS 14 or later.

  • Nothing on screen is captured. Weeve records the Mac's audio, not the display, so a screen share or shared slides are not in the transcript. Video files you import are transcribed from their audio track.

  • Accuracy drops on heavy accents and where people talk over each other. Two-person interviews are close to the best case for a speech model; a focus group is not.

  • Speaker labels degrade as voices multiply. A two or three person interview reads cleanly. On a busy focus group, expect to correct the labels by hand.

  • The free Starter plan covers 10 recordings a month, which is a real constraint if you are running a study with thirty participants.

  • No SOC 2, ISO 27001 or HIPAA attestation. If your institution requires one, that is a hard stop regardless of the architecture, and it is worth raising before the ethics submission rather than after.

Where a local tool is the wrong choice

Three cases, and they are common enough to state plainly.

You need human-verified accuracy. For a deposition, court use, or quotes going into print under your name, a human transcription service is the right answer. Rev and Verbit exist for exactly this. No on-device model is a substitute, and neither is a cloud one.

You need qualitative coding, not just a transcript. MAXQDA, NVivo and ATLAS.ti are built for thematic analysis. They all transcribe as well, in the cloud, but the coding is what you are paying for and that is a different job. Weeve produces the transcript. It does not code it. Plenty of researchers use a local tool for transcription and then bring the text into one of those.

You need a shared workspace. If a team of five needs to annotate the same transcript, a cloud tool is the honest answer. There is no shared workspace here.


Choosing, in one question

Ask what you promised the person you interviewed.

If you promised nothing in particular, and the material is not sensitive, the cloud tools are faster to start and better at collaboration. Use them.

If you promised confidentiality, named a specific data-handling arrangement on an ethics form, or protected a source, then where the audio is processed is not a preference. It is the commitment, and it should decide the tool.

Picking the tool is the smaller half of the job. The transcript itself still has to be produced, corrected and written up, and if this is academic work there are conventions attached to all three: how to transcribe a research interview covers the method.

FAQ

What is the best interview transcription software? It depends on what you promised the interviewee. For collaboration and speed, the established cloud services are strong. For interviews covered by a confidentiality commitment or an ethics submission, a tool that processes on your own machine removes the question entirely. For anything going into print or court, use a human transcription service.

Can I transcribe an interview without uploading it? Yes. On-device transcription runs the speech model on your own computer, so the audio never reaches a server. On a Mac with Apple Silicon this is now fast enough for routine use, and the transcription itself runs with the network switched off once the model has downloaded.

Is AI transcription accurate enough for research interviews? For clear two-person audio, generally yes, and researchers routinely correct the output rather than typing from scratch. It is materially worse on heavy accents, crosstalk and poor recordings. Budget time for a correction pass, and do not quote directly from an uncorrected automatic transcript.

Do transcription services train their models on my interviews? Some do unless you opt out, and the setting is not always obvious. Read the specific clause rather than the marketing page, and if the recording is covered by an ethics approval, check what that approval actually permits before uploading anything.

How do I transcribe an interview I recorded on my phone? Transfer the audio file to your computer and import it. Any tool worth using accepts common formats directly and reads the audio track without a conversion step. If you would rather not install anything for a single file, a browser transcriber will handle it.

What do I tell an ethics committee about transcription? Name the tool, say where the audio is processed and stored, state the retention period, and say whether the vendor uses customer data for training. If processing happens on your own machine, say that, and say what is backed up, because "it stays on my laptop" and "it is not in a backup" are different claims.

An interview is a promise before it is a recording. The tool should keep it.

If that promise is the constraint you are working under, Weeve's free Starter plan transcribes on a Mac with the audio staying on the device, 10 recordings a month, with speaker labels and export. It needs Apple Silicon and macOS 14 or later. If you are not on a Mac, or you want to judge the output before installing anything, the browser transcriber handles a single file of up to an hour in the tab and the recording never leaves the page.