Published
A 60-second video script is about 130 spoken words, not the 150 the words-per-minute rule suggests. The difference is silence: title cards, end cards and on-screen numbers take seconds in which nobody talks. Write to the word budget, in two columns, and the document becomes something a vendor can quote from.
Almost every page ranking for this question treats scripting as a creative act. Define your goal, know your audience, open with a hook, write like you talk, use two columns, end with a call to action. All of it is true and none of it is checkable. You can follow every line of that advice and still hand over a script that runs 80 seconds when the slot is 60, which is discovered late, costs a re-record, and is one of the most common reasons a video comes back wrong.
This is the arithmetic version. A script is a specification with a runtime, a word count and a delivery format, and all three can be verified before anyone is booked.
How this was checked. For this query in the United States on 10 August 2026, Google returned an AI Overview, a video pack, a forum block and six organic results — a screen-recording vendor, a university library guide, two AI-video tools, a copywriting blog and an e-learning site. Every one of them gives the same six-step outline; not one publishes a word budget that accounts for silent time, and the People Also Ask box carries two separate questions about script length that the indexed answers handle with a single multiplication. Speaking rates below come from published guidance, named in the table; the two radio conventions are trade figures from a US agency and a US recording studio rather than a single citable study. The word budgets are a model built from those rates: every figure follows from one stated formula and one stated overhead assumption, so the arithmetic can be checked line by line.
Word budgets for 15, 30, 60 and 90 seconds, and for two to five minutes
Start from the finished duration and work backwards. A video is not a block of speech; it is a block of speech with fixed silent overhead at both ends and a beat of silence every time something on screen has to be read rather than talked over.
The overhead used here is deliberately modest: 2 seconds of opening title, 3 seconds of end card, and one 1-second beat per 20 seconds of runtime for an on-screen number, product reveal or interface state. A 15-second cut cannot afford that, so it compresses to a 1-second open and a 2-second end card with no reveal beat at all.
| Finished length | Naive count at 150 wpm | Silent overhead | Speaking time | Word budget | Shortfall |
|---|---|---|---|---|---|
| 15 sec | 37.5 | 3 sec | 12 sec | 30 | 20% |
| 30 sec | 75 | 6 sec | 24 sec | 60 | 20% |
| 60 sec | 150 | 8 sec | 52 sec | 130 | 13% |
| 90 sec | 225 | 9 sec | 81 sec | 202 | 10% |
| 2 min | 300 | 11 sec | 109 sec | 272 | 9% |
| 3 min | 450 | 14 sec | 166 sec | 415 | 8% |
| 5 min | 750 | 20 sec | 280 sec | 700 | 7% |
Every row is the same three operations: subtract the overhead from the finished length, multiply the remaining seconds by 2.5 words per second, round down. The 60-second row is worth memorising, because it is the length most briefs default to and the one where the naive number is wrong by a clean 20 words.
The shortest videos are where the words-per-minute rule lies most
Read the shortfall column downwards. It is not a constant discount — it runs 20% at 15 and 30 seconds and decays to 7% at five minutes. The reason is structural rather than stylistic: the overhead is largely fixed, so it consumes a shrinking fraction of a growing runtime. Three seconds of top and tail is a fifth of a 15-second cut; the full five seconds is one sixtieth of a five-minute one.

This has a practical consequence that runs against instinct. Short-form scripts are usually treated as the easy ones, so they get the least planning — but they are the format where a naive word count does the most damage, and the format with the least room to absorb it. Going 20 words long in a five-minute piece costs eight seconds nobody notices. Going eight words long in a 15-second cut means something gets cut off, or the read is sped up until it stops sounding like a person.
The corollary is that the shorter the video, the more a written script is worth against the effort of writing one. Nobody improvises a five-minute explainer. Plenty of teams improvise a 15-second one, which is exactly the case where the seconds are least forgiving.
Choosing a planning rate: 150 words a minute, and when to move off it
The 150 figure is not arbitrary. It is the average conversation rate for English speakers in the United States as measured by the National Center for Voice and Speech, and it survives as a planning number because delivery contexts cluster near it.
| Context | Rate | What the number actually measures |
|---|---|---|
| Everyday US English conversation | ~150 wpm | The NCVS average, the source of the rule of thumb |
| Audiobook and long-form narration | 150–160 wpm | Audible plans at 9,300 words per finished hour, which is 155 wpm |
| Live presentation to a room | 100–150 wpm | Slower: the room, the slides and the pauses absorb time |
| Radio spot, conservative convention | 120–170 wpm | One US agency publishes 30–40 words for 15 seconds, 75–85 for 30, and 150–170 for 60 |
| Radio spot, fast convention | 180 wpm | One US studio publishes a flat “3 words per second” — 45, 90 and 180 words for the same three lengths |
| Rehearsed conference talk | 173 wpm average | A five-talk sample, range 154–201 — a small and heavily rehearsed set |
Two things in that table matter more than the headline number.
The first is that the published radio counts are ceilings for video, not targets, and the reason is mostly arithmetic rather than delivery. The 30-second row of the budget table says 60 words while the broadcast guidance says 75 to 90 — but that budget row is already computed at 150 wpm, which is exactly the conservative agency’s own low end. The rates agree. What differs is that the broadcast figure assumes all 30 seconds are speech and six of them are not. A smaller part of the gap is genuine pace: radio has no picture competing for those seconds, so the read can sit at the top of the band in a way a video read cannot.
The second is that the conservative convention is not a flat rate, and that is the most useful thing in the table. Its own numbers work out to 150–170 wpm at 60 seconds but only 120–160 wpm at 15 seconds: it quietly slows down as the spot gets shorter. The fast convention does not, and the gap between the two is consequently 50% at 15 seconds and 20% at both 30 and 60. Somebody writing radio for a living arrived at the same conclusion this model does from the opposite direction — short slots need proportionally more headroom — which is a reasonable sign the effect is real and not an artefact of the assumptions here.
Move off 150 deliberately, not by feel. Go to 130 if the audience is non-native, the material is technical, or the read is over dense on-screen data. Go to 165 for a presenter-led social cut where energy matters more than absorption. Then rebuild the budget, because every row above changes.
Video left, audio right: the two-column format in full
The two-column audio-visual script is the standard for corporate, explainer and commercial work for one reason: three different people read it, and each of them only needs one column. The director and the editor read the left. The voice artist reads the right. Nobody has to parse a paragraph to find their instruction.

The conventions that make it work are unglamorous and worth following exactly:
- Left column narrower than the right, roughly 40/60. Visual notes are phrases; audio is sentences, and sentences need the room.
- One row per shot, not per sentence. A row is a unit the editor can move. If a row contains two visual ideas, it is two rows.
- Number rows in tens — 10, 20, 30. Inserting a shot then becomes 15 rather than a renumber that breaks every note anyone has already written against the old numbers.
- On-screen text goes in the left column in quotation marks, spelled exactly as it must appear. Legal wording, product names and pricing are set here or they get set wrong by someone with no authority to set them.
- Sound that is not voice is still audio. Music changes, silences and effects belong in the right column, because they occupy time.
- Carry a words column and a seconds column at the right edge. This is the part almost every template omits, and it is what turns the document from a description into a specification.
Anything that is only atmosphere belongs in a note under the table, not in a cell. Cells are instructions.
What a production company needs in the document before it can quote
A script that reads beautifully and cannot be priced is not finished. Quoting requires knowing what has to be created, and the script is where that becomes visible or stays hidden. Nine things make the difference between a document a vendor can price in an hour and one that generates a week of questions.
| Field | Why a quote depends on it |
|---|---|
| Version and date | Prevents two people costing two different documents |
| Target duration and tolerance | “60 seconds ±2” is a spec; “about a minute” is not |
| Deliverables and aspect ratios | One master or one master plus three crops is a different job |
| Shot numbers | The unit everything else is counted and scheduled in |
| Video cell per row | What the camera or the animator is actually pointed at |
| Audio cell per row | Voice, dialogue, music and effects, each occupying time |
| On-screen text, verbatim | Copy that has to be typeset, checked and possibly localised |
| Word and second count per row | Lets the vendor verify the runtime instead of trusting it |
| Asset origin per visual | Existing, to be shot, to be animated, or to be licensed |
That last field is the one that moves the number most. A visual marked “to be animated” and a visual marked “already own it” cost differently by an order of magnitude, and whether a piece is drawn, filmed or built in 3D is a decision the script forces without meaning to — the choice between 2D and 3D is mostly settled by what the visual column asks for. Writing “exploded view of the mechanism” is a modelling budget. Writing “hand opens the box” is an afternoon.
The same logic runs the other way for anything filmed. Every row of the visual column that describes something other than a person talking is an item on a shot list, and the supporting footage laid over the narration has its own counting problem once the day gets planned. Write the visual column as though someone will have to point a camera at each cell, because someone will.
If the project has a written design brief behind it, the script inherits the tone and the constraints from there rather than reopening them. Where no brief exists, the script quietly becomes one, which is workable for a single video and a problem by the third.
A 60-second script, counted line by line
Below is a complete 60-second script for a fictional inventory product, built against the budget in the first table: 8 seconds of silent overhead, 52 seconds of speech, 130 words. The counts are shown so the arithmetic is checkable rather than asserted.
| # | Video | Audio | Words | Sec |
|---|---|---|---|---|
| 10 | Title card, logo on flat colour | silent | 0 | 2.0 |
| 20 | Wide: Dana at a desk, three screens, warehouse behind | VO: “Every Monday, Dana rebuilds the same stock report by hand. Four spreadsheets, ninety minutes of work, and by Tuesday the numbers are already wrong.” | 24 | 9.6 |
| 30 | Screen recording: dashboard assembling, sources connecting | VO: “Kestrel pulls those same four sources into one live view. It refreshes every fifteen minutes, nobody has to touch it, and the number on the screen is the number in the warehouse.” | 32 | 12.8 |
| 40 | On-screen text: “Refreshes every 15 minutes” | silent | 0 | 1.0 |
| 50 | Close: hands, phone alert, shelf label | VO: “When a line runs low, the alert reaches the person who can order more, not the person who files the report. That is the difference.” | 25 | 10.0 |
| 60 | On-screen text: “90 minutes back, per person, per week” | silent | 0 | 1.0 |
| 70 | Wide: two staff on the floor, tablet between them | VO: “Teams running it get back about ninety minutes each week, and stop arguing about whose spreadsheet is right. Setup takes one afternoon, and no one has to change how they already work.” | 32 | 12.8 |
| 80 | Logo lockup over the floor shot, held | silent | 0 | 1.0 |
| 90 | Talking head: operations lead, mid-shot | VO: “We stopped holding the Monday meeting altogether. The report was the meeting, and now it writes itself.” | 17 | 6.8 |
| 100 | End card: logo, URL, one line of CTA text | silent | 0 | 3.0 |
| Total | 130 | 60.0 |
Three properties of that table are the point of the exercise. The words column sums to exactly the 130-word budget. The seconds column sums to exactly 60.0. And every spoken row’s duration is its word count divided by 2.5, so any line can be lengthened or shortened and the consequence is immediately visible in the total rather than discovered in a recording booth.
Note also what the silent rows are doing. The two on-screen figures — the refresh interval and the ninety minutes — are given a full second alone, because a number that is said and shown at the same moment is usually neither read nor heard. The third beat, row 80, is the held logo lockup: the one second in the script bought for rhythm rather than for reading.
The example also contains a fault worth seeing. Row 30 says “every fifteen minutes” in the voice and row 40 shows the same figure on screen, so one number is paid for twice. That is exactly what step three of the cut list below exists to catch, and it is the first place to look if this script came back four seconds long.
From script to storyboard to a timed animatic
The script is the first of three documents, and each one tests something the previous one could not.

The script fixes words and running time. The storyboard fixes framing: one frame per shot, taken straight from the numbered rows of the visual column, answering what size and what angle. Because the rows are already numbered, the storyboard inherits the numbering and the two documents stay married through every revision.
The animatic is where the timing stops being theoretical. It is the storyboard frames cut to the actual recorded voice track, so it plays at the real duration rather than the calculated one. This is the first moment the project can be watched, and it is the last cheap place to discover the script is long.
That asymmetry is the whole argument for building one. Cutting 15 words at the animatic stage means editing a text document and re-recording a scratch track. Cutting the same 15 words after animation has started means the voice is re-recorded, every shot that sat under those words is re-timed, and anything already rendered is rendered again. The work is identical; only the cost of doing it has changed.
Two habits make the gate useful rather than ceremonial. Record a scratch voice track — anyone’s voice, a phone is fine — as soon as the script is signed off, because the real duration of a read is a fact and the calculated one is a forecast. And time the animatic against the target with a tolerance, not a hope: if the spec said 60 seconds ±2 and the animatic runs 64, the script goes back before a booth is booked.
Seven checks that say the script is shootable
Run these before the document leaves the building. Each one fails loudly and cheaply now, and quietly and expensively later.
- The spoken word count is at or under the budget for the target duration, and the number is written on the document.
- Every audio cell has a video cell beside it. An orphan line of voiceover means a shot nobody has planned.
- Every on-screen text string is written verbatim, including punctuation, capitalisation and any legal line.
- Every visual is marked with its origin — existing, to be shot, to be animated, or to be licensed. No cell is unmarked.
- Shot numbers run in tens in the first draft, later inserts take the numbers in between, and no number is reused anywhere in the revision history.
- The read has been timed out loud once, by a person, against a stopwatch, and the result is recorded on the document.
- Nothing in the visual column needs a person, a location or a permission that is not already secured. A script that requires a factory floor nobody has asked about is a schedule risk written in prose.
Check six catches more problems than the other six combined, because it is the only one that tests the document against reality rather than against itself. It also takes ninety seconds.
What to cut first when the read comes back long
Reads come back long. The order in which words are removed decides whether the video survives it, and the instinct — trimming evenly throughout — is the worst available option, because it shortens everything and improves nothing.
- Cut the setup, not the payoff. The first 15% of a first draft is almost always throat-clearing that the visual has already established.
- Cut adjectives before clauses. Adjectives cost words and carry little; clauses carry the argument.
- Move numbers from the voice to the screen. A figure costs a second on screen and two or more in the voice once it is wrapped in a sentence, and it lands harder in type. Watch the net: the on-screen beat is not free.
- Merge two shots making the same point. Two proofs of one claim is one proof and a redundancy.
- Delete the second example. Nobody has ever needed the second example in under two minutes.
- Only then, speed the read — and treat it as a cost, not a saving. Comprehension pays for it.
- Last resort: change the duration and re-price it. Sometimes the brief was wrong, and 75 seconds is the honest answer.
Steps one to five remove words without removing meaning. Step six removes clarity. Step seven removes the constraint. Working in that order means the expensive options stay available instead of being spent first.
Scripting well is mostly refusing to find out late. The word budget, the two columns and the timed animatic exist to move every discoverable problem to the cheapest hour in the schedule, which is the one before anyone is booked. If a script is being written against a live production date, our video production team works from exactly this document, and larger campaigns that need the script to survive across several formats at once are handled through media production.
10 / Reader questions
Frequently asked questions
01How many words is a 60 second video script?
About 130 words, not 150. A minute at the usual 150-words-per-minute planning rate looks like 150 words, but a 2-second title card, a 3-second end card and three 1-second on-screen reveals take 8 seconds in which nobody speaks. That leaves 52 seconds of speech, which is 130 words.
02How long should a 5 minute video script be?
About 700 words of spoken audio. The figure usually quoted is 750, which is five minutes multiplied by 150 words per minute with nothing subtracted. Remove a 2-second open, a 3-second end card and one 1-second on-screen beat per 20 seconds of runtime and 20 seconds of the five minutes are silent.
03How long is a 3 minute video script?
About 415 words. Three minutes at 150 words per minute is 450 words on paper, but 14 seconds of a three-minute video are typically silent: 2 seconds of opening title, 3 seconds of end card, and nine 1-second beats where an on-screen number or product has to be read rather than talked over.
04What is a two-column video script?
A page split so that everything the viewer sees is in the left column and everything the viewer hears is in the right, with one row per shot. It is the standard format for corporate, explainer and commercial work because a director, an editor and a voice artist can each read only their own column.
05What is the difference between a script and a storyboard?
The script fixes the words and the running time; the storyboard fixes the framing. A script says a hand picks up the device and what the voice says while it happens. A storyboard says whether that is a top-down shot or a close side angle. Timing is settled first, because it decides how much there is to frame.
06Should a video script be written in full sentences?
Yes for anything spoken, because full sentences are the only way to count words and therefore the only way to know the runtime. Bullet points hide length: four bullets can read in nine seconds or twenty-five depending on who says them, and neither the voice artist nor the editor can plan against that.