How Much Does It Cost to Hire a Freelance Transcriptionist?
Human transcription costs $0.75–$3.00 per audio minute for most work. A one-hour interview with clean audio and two speakers comes to $45–$150 — about $98 in the middle — and the same hour costs $90–$210 as a legal transcript or $120–$300 as a medical one. The calculator below is set to that hour of clean interview audio, so change the recording type, length, or turnaround to price your own file.
Price Your Transcription Job
The default. Clean studio or Zoom audio, one or two speakers
Runtime of the file, not how long it takes to type
Fillers, false starts, and stutters removed — what most clients want
No surcharge
No surcharge
No rush fee
Expect to pay, per audio minute
$0.75 — $2.50
60 audio minutes = $45–$150 for the file
Per audio minute
$0.75–$2.50
Midpoint $1.63
This file
$45–$150
60 audio minutes
Per audio hour
$45–$150
The same rate × 60
How this quote is built
Behind that price is roughly 4 hours of work — 4 hours per hour of audio is the standard pace for this kind of file. At the midpoint quote that's about $25/hr of the transcriptionist’s own time, which is the sanity check to run on a quote that looks cheap: below roughly $20/hr you are usually buying an unedited machine transcript or a first-week freelancer.
The same file, priced as every kind of transcript
| Recording type | Per audio minute | 60 audio min | Per audio hour |
|---|---|---|---|
| General | $0.75–$2.50 | $45–$150 | $45–$150 |
| Qualitative research | $1.25–$2.75 | $75–$165 | $75–$165 |
| Academic / lecture | $1.00–$2.50 | $60–$150 | $60–$150 |
| Legal | $1.50–$3.50 | $90–$210 | $90–$210 |
| Medical | $2.00–$5.00 | $120–$300 | $120–$300 |
| Captioning / subtitling | $2.50–$7.50 | $150–$450 | $150–$450 |
Every row carries your current style, speaker, audio, and turnaround settings — so this is what the subject matter alone is worth. Legal and medical cost more for the same minute of audio because the terminology has to be verified, not guessed.
Per Audio Minute, Per Audio Hour, Per Real-Time Hour
Transcription is quoted in three units, and only two of them are about your file. Knowing which one you are being quoted in is the difference between comparable bids and a spreadsheet of numbers that can't be lined up.
| Unit | Typical range | Used for | What to watch |
|---|---|---|---|
| Per audio minute | $0.75 – $3.00specialist work runs above it | The standard unit for recorded files, and the one on most rate sheets | Whether the quoted rate is the base or the all-in figure. Verbatim, four speakers, and noisy audio turn $1.50 into $2.63 |
| Per audio hour | $45 – $180 | The same number times 60. Agencies and services quote this way because it sounds cleaner | Partial hours. Ask whether a 70-minute file bills as 70 minutes or as two hours |
| Per real-time hour | $60 – $150/hr | Live captioning and CART, where there is no file to measure | On recorded audio this unit bills you for their typing speed. Only accept it for live work |
What the per-audio-minute rate is worth then depends on what is on the recording. These are the base rates for clean audio, one or two speakers, clean verbatim, and a standard 3–5 business day turnaround — the same bands the calculator starts from.
| Recording type | Per audio minute | Per audio hour | What the rate buys |
|---|---|---|---|
| General — podcast, interview, meeting | $0.75 – $2.50 | $45 – $150 | Clean-verbatim transcript, speakers labeled, delivered as a document |
| Qualitative research | $1.25 – $2.75 | $75 – $165 | Focus groups and in-depth interviews, usually to a verbatim standard |
| Academic / lecture | $1.00 – $2.50 | $60 – $150 | Citations, proper names, and field terminology checked rather than guessed |
| Legal — deposition, hearing | $1.50 – $3.50 | $90 – $210 | Strict verbatim, speaker identification, and line numbering included |
| Medical | $2.00 – $5.00 | $120 – $300 | HIPAA-aware handling and drug, dosage, and diagnosis terminology |
| Captioning / subtitling | $2.50 – $7.50 | $150 – $450 | A timed caption file — timecodes and line lengths, not just the words |
US market ranges for freelancers hired direct, on clean audio with a standard turnaround. The same rates from the transcriptionist's side, with the full surcharge stack, are on the freelance transcription rates page.
The gap between the bottom and the top of the general row is the whole story of this market. At $0.75 an audio minute you are buying an hour of someone's typing; an hour of audio is roughly four hours of work, so that quote pays about $11 an hour before tax and equipment. At the $98 midpoint of the same file, the transcriptionist is earning around $25 an hour — a normal freelance rate for skilled work, and the level at which you can expect names spelled correctly and inaudible passages flagged rather than invented.
That arithmetic is the most useful thing on this page when you are comparing bids. Divide any per-audio-hour quote by four and ask whether the resulting hourly rate buys the accuracy your transcript needs. Legal and medical audio divides by five or six instead, because the terminology has to be verified rather than typed — which is why their base rates start where general transcription tops out.
What Actually Moves the Price
Transcription is priced on listening time, not runtime — and everything that makes a file slower to listen to is a surcharge. These are the standard ones, applied to the base rate above.
| Surcharge | Typical add | Why it costs more | How to avoid paying it |
|---|---|---|---|
| Poor audio | +25 – 50% | Phone audio, room echo, and crosstalk turn one pass into three over the same passage | Record each speaker on their own track, or hand out lapel mics. This is the cheapest fix on the list |
| 3+ speakers | +25% | Every turn has to be attributed, and people talk over each other | Have speakers say their name once at the start; supply a participant list |
| 5+ speakers | +40% | Roundtables and focus groups — voices stop being individually recognizable | Multitrack the recording if you can; it can drop this to the 3+ band |
| True verbatim | +25 – 40% | Every filler, stutter, and false start is transcribed rather than tidied — 30–50% slower | Ask for clean verbatim unless how it was said is part of the record |
| Accents / technical jargon | +20 – 30% | Unfamiliar terms and pronunciations mean lookups and verification, not typing | Send a glossary of names, products, and acronyms with the file |
| Timestamps | +10 – 20% | 10% at every paragraph, 20% at every speaker change — the second marks every turn | Ask for the interval you will use. Paragraph-level is enough to find a quote |
| Rush turnaround | +25 – 150% | Next-day +25–50%, same-day +50–100%, 6-hour or overnight +100–150% | Book known deadlines a week out, or split a batch between two transcriptionists |
Surcharges add rather than compound: a verbatim focus group on noisy audio is base +25–40% +40% +20–25%, not those figures multiplied together.
Two of these are worth a moment before you record rather than after. Audio quality is the one you control and the one that costs the most — a conference-room speakerphone recording of six people can price at double the same conversation captured on separate mics, and no amount of negotiating gets that back. Verbatim style is the one clients get wrong in the expensive direction: true verbatim sounds like the thorough option, but it costs 25–40% more and produces a transcript nobody enjoys reading. Ask for it when the ums are evidence, and clean verbatim the rest of the time.
This is also why quotes for one file land so far apart. A transcriptionist who has listened to a sample is pricing what they heard; one who hasn't is pricing studio audio and will come back with a revised number, or absorb it once and decline your next job. Send a 60-second sample of the worst part of the recording — not the introduction, which is always the cleanest minute — and the quotes you get back become comparable.
Where the Very Cheap Quotes Come From
Automated transcription is priced in cents per audio minute rather than dollars, so on cost alone nothing competes with it. It is a different product, not a cheaper version of the same one: it does not reliably handle crosstalk, proper names, jargon, or accented speech, and — the part that matters — it cannot tell you which passages it got wrong. For searchable notes from an internal meeting, that is a fine trade. For anything that will be quoted, published, filed, or coded as research data, someone still has to check it against the audio, and that person is either a transcriptionist or you.
A machine draft plus a human clean-up pass is the honest middle, and it does cost less than transcription from scratch — but by less than most clients expect, because verifying a draft still means listening to the whole file. It works best on exactly the audio the machine handles well: one or two clear speakers, no jargon. On the files where a raw AI transcript is worst, cleaning it up can take longer than typing from nothing, which is why some transcriptionists quote the same rate either way.
A quote far below the bands above usually means one of three things: an unedited machine transcript being resold, a first-week freelancer who hasn't yet timed themselves against a real file, or an assumption of clean studio audio that your recording does not meet. Ask which — and ask what happens to inaudible passages. The answer you want is that they get flagged with a timestamp. The answer that costs you later is a plausible-looking sentence nobody can trace back to the tape.
Price Your Own File
Three jobs clients hire transcriptionists for by name, each loaded into the calculator at the top of this page with the inputs already set.
90-minute focus group
Research · true verbatim · 5+ speakers · mixed audio
Returns $2.31 – $5.64 per audio minute — $208 – $508 for the session, midpoint $357. Three surcharges stacked on one file, which is what a room full of people costs.
Load the focus group →45-minute deposition
Legal · 3–4 speakers · clear audio · next-day
Returns $101 – $276, midpoint $189. Verbatim adds nothing here — a legal transcript is verbatim by definition, so it's already in the base rate.
Load the deposition →Podcast episode, same day
General · 60 minutes · clean verbatim · 12-hour turnaround
Returns $68 – $300 against $45 – $150 on a standard deadline. The rush fee is the single most avoidable line on this page.
Load the rush job →Change any input to model your own file — or swap model=per-hour in the URL to see the same job priced per audio hour.
Recommended gear
Transcribing it in-house instead?
If the answer to the rates above is that you'll type it yourself, the cost moves from an invoice to your own afternoon — and an hour of audio is about four hours of that. These are the six things that decide whether it's four hours or seven. The pedal and the transcription headset are the pair that does the real work: playback control moves to your foot, so both hands stay on the keyboard and you stop losing a minute to the mouse every time you need to hear a sentence twice. The software side isn't buyable here — Express Scribe, oTranscribe, and the AI tools are downloads or subscriptions — so what's worth owning is the hardware. Affiliate links — buying through them helps fund this calculator.
Infinity IN-USB-2 foot pedal
VEC · USB transcription foot control
The most widely used transcription pedal, and the one playback software is most likely to recognize without extra setup. Play, rewind, and fast-forward move to your foot — the single change that cuts most from a transcription ratio.
Spectra SP-USB transcription headset
Under-the-chin USB headset with volume control
Built for hours of playback rather than music: light enough to forget, with the volume dial on the cable so you can lift a mumbled passage without leaving the document. The natural pair for the pedal.
Sony WH-1000XM5
Noise-cancelling headphones
For the poor-audio files that carry the +25–50% surcharge when you send them out. Cancellation this good is the difference between replaying a passage twice and replaying it six times.
Ergonomic split keyboard
Kinesis Advantage2 / Keychron Q11
Four hours of continuous typing per hour of audio is the ratio, and wrists are what gives out first. A split keyboard pays for itself the first time it heads off a flare-up mid-deadline.
Dell U2723QE 27″ 4K USB-C
Second monitor
Player and timestamps on one screen, the transcript full-height on the other. Transcribing in a window stacked above a media player is how the fifth hour becomes the sixth.
Time Timer Original 8″
60-minute visual timer
Run it beside one file and you'll know your real transcription ratio instead of the optimistic one. That number is what decides whether typing the next batch yourself is cheaper than hiring it out.
As an Amazon Associate this site earns from qualifying purchases. Links are sponsored.
Related Calculators & Guides
Freelance Transcription Rates
The same rates from the transcriptionist's side, with the full surcharge stack
Cost to Hire a Freelance Writer
What it costs to turn that transcript into an article, per word or per project
Cost to Hire a Translator
What the finished transcript costs in a second language, by pair and per word
Podcast Editor Rates
Per-episode editing, with transcripts priced as an add-on to the edit
Voice-Over Artist Rates
The other side of audio work — per finished minute rather than per audio minute
Virtual Assistant Rates
Hourly and retainer pricing when transcription is one task among many
Frequently Asked Questions
How much does it cost to hire a freelance transcriptionist?
Most human transcription costs $0.75–$3.00 per audio minute. General audio — podcasts, interviews, meetings with one or two clear speakers — runs $0.75–$2.50 per audio minute, which is $45–$150 per audio hour. Qualitative research sits at $1.25–$2.75, academic and lecture audio at $1.00–$2.50, legal at $1.50–$3.50, medical at $2.00–$5.00, and captioning or subtitling at $2.50–$7.50 because timecoding is a separate job on top of the transcript. Verbatim style, extra speakers, poor audio, and rush turnaround are surcharges on those base rates.
How much does it cost to transcribe one hour of audio?
An hour of clean, two-speaker audio costs $45–$150 from a freelance transcriptionist, with the midpoint around $98. The same hour costs $90–$210 as a legal transcript, $120–$300 as medical, and $150–$450 as captions with timecodes. Add roughly 25% if there are three or more speakers, 40% for five or more, 25–50% for poor audio, 25–40% for true verbatim, and 25–100% for rush turnaround.
Why do transcription quotes for the same file vary so much?
Because the quote is a function of listening time, not file length, and the things that add listening time are invisible in a file name. Five surcharges account for nearly all of the spread: true verbatim over clean verbatim (+25–40%), three or more speakers (+25%) or five or more (+40%), poor audio such as phone recordings, room noise, or crosstalk (+25–50%), heavy accents or technical jargon (+20–30%), and rush turnaround (+25–100%). A quote given without hearing the audio is a guess, which is why the cheapest quote is usually the one that assumed studio conditions.
Is AI transcription cheaper than hiring a transcriptionist?
Automated transcription is priced in cents per audio minute rather than dollars, so on cost alone it wins by an order of magnitude. It is a different product: it does not reliably handle crosstalk, names, jargon, or accented speech, and it has no way to flag what it got wrong. That matters where the transcript is evidence, a quotation, or a deliverable — a court filing, a published interview, a research coding frame — and it matters much less for searchable notes from an internal meeting. The middle path is a machine draft with a human clean-up pass, which does cost less than transcription from scratch but not by as much as clients expect, because verifying a draft against the audio still means listening to all of it.
Should I ask for verbatim or clean verbatim?
Clean verbatim for almost everything: fillers, false starts, and stutters are removed, and the transcript reads the way the speaker meant it. Ask for true verbatim — every um, pause, and repetition kept — only when how something was said is part of the record: legal proceedings, behavioral and linguistic research, and compliance recordings. True verbatim takes 30–50% longer to produce and carries a 25–40% surcharge, so requesting it by default on podcast or marketing audio buys a harder-to-read transcript at a higher price. Legal rates already include it, which is part of why they start higher.
How fast can a transcriptionist turn a file around, and what does rush cost?
Standard turnaround is 3–5 business days for general transcription, because an hour of audio is about four hours of work and a freelancer is rarely holding that block open for you. Next-day delivery adds 25–50%, same-day or 12-hour adds 50–100%, and 6-hour or overnight adds 100–150%. Rush pricing is the surcharge worth avoiding: booking a known deadline a week early costs nothing, and on a multi-file project splitting it between two transcriptionists usually beats paying one of them a rush fee.
What should I send with the file to get an accurate transcription quote?
The recording itself, or a 60-second sample of the worst part of it — not the introduction, which is always the cleanest minute. Then five details: the runtime, how many speakers there are and whether you need them named, verbatim or clean verbatim, whether you want timestamps and how often, and the deadline. Add a list of proper nouns, product names, and technical terms; it takes five minutes to write and removes most of the errors a transcriptionist would otherwise have to guess at.
Do transcriptionists charge extra for timestamps and speaker names?
Speaker labels are standard on multi-speaker files and are already reflected in the multi-speaker surcharge. Timestamps are usually a separate line: about 10% for a timestamp at every paragraph and about 20% for one at every speaker change, since the second means marking every turn in the conversation. Ask for the interval you will actually use — timestamps every 30 seconds sound thorough and cost real money to produce, while paragraph-level marks are enough to find a quote in the audio.