Skip to main content
ToolsBay

TEXT UTILITIES

Word Count and SEO: What the 1,447-Word Study Actually Measured

7 min read · ToolsBay editorial · Published · Updated

Just want to do it now?

Count words, sentences and paragraphs, plus estimated reading time.

Open Word Counter

Every few months someone republishes the figure: the average first-page Google result is 1,447 words. It is a real number from a real study. It is also the answer to a question nobody asked, and it has been used for a decade to justify padding.

An earlier version of this post repeated that number approvingly and then went further, claiming Google uses "LSI keywords" to confirm a page's topical authority. That claim was false. This is the replacement, and it starts by taking the number apart.

Where 1,447 came from

Backlinko analysed millions of search results and reported that the average word count of a page-one result was 1,447. The same publisher's earlier study, over a smaller sample, put it at 1,890. Nothing about Google changed by four hundred words between those two studies. The sample changed.

That is the first problem with the figure: it is an average over a corpus somebody chose. Change the mix of queries and the average moves. Include navigational and answer-shaped queries — "npm cache clear", "http 418", "what time is it in tokyo" — and it collapses, because the winning page for those is a paragraph.

The second problem is bigger. The study measured pages that already rank. There is no control group. Nobody assembled the set of 1,447-word pages that rank nowhere, and that set is enormous. Pages that reach page one for competitive queries tend to come from publishers who also have links, a brand, an editorial process and a site that renders fast. Those publishers also write long. Length rides along with the causes; it is not one of them.

Google's search advocates have said plainly, more than once, that word count is not a ranking factor. Both things are true at once — the correlation exists in the data, and the lever does not exist in the algorithm. Adding 800 words to a thin page does not move it, because the 800 words were never what the ranking pages had.

"LSI keywords" are not a thing

This one deserves killing on its own, because it is the load-bearing myth under most length advice.

Latent semantic analysis is real. It comes out of information-retrieval research in the late 1980s. You build a term-document matrix, run a singular value decomposition on it and keep the strongest dimensions, which lets "car" and "automobile" land near each other without anybody writing a synonym list. It works on a fixed collection of documents. Adding documents means redoing the decomposition over the collection.

That is the problem. The technique assumes a corpus that sits still. The web does not sit still, and a decomposition over hundreds of billions of documents is not something you re-run because a page changed. Google's own people have said there is no such thing as an LSI keyword. The phrase survives because it sounds like machinery, and because it gives "write longer" a technical-sounding reason.

What is actually true is duller and more useful: a page that genuinely answers a question tends to contain the words that question drags with it. Write a real answer about Nginx and you will mention certificates, ports and logs without being told to. The vocabulary is a symptom of having covered the thing. Sprinkling the vocabulary onto a page that has not covered it produces a page that has not covered it.

A word count is not one number

Here is where a tools site can be more useful than an opinion. Take this paragraph:

A JWT payload is Base64, not encryption — anyone holding the token can
read it. The signature is what proves it is genuine, and checking that
needs the secret. A well-known consequence: a token pasted into a
server-side decoder is a token you have handed to a stranger... and you
cannot take it back.

Split it on whitespace and you get 54 words. Split it into letter-and-digit tokens instead, the way a frequency counter does, and you get 53. The entire difference is the standalone em dash: with spaces either side it is a whitespace-delimited token, so one rule counts it as a word and the other never sees it. well-known survives as one word under both. Set the English function words aside and you are down to 28 — just over half of where you started, from the same paragraph.

The word counter here uses the whitespace rule and says so, because a marker or an editor counts well-known as one word rather than two. The character counter reports the same text as 300 characters and 302 UTF-8 bytes, and those two bytes of difference are entirely the em dash — one character, three bytes. Put emoji or accented names in a draft and that gap stops being a rounding error.

Scale this up and the "word count of a page" gets worse. A crawler reads rendered text, so your header, nav, footer, cookie notice and related-links block all count toward it. A tutorial with a 60-line configuration file in it carries several hundred tokens that no reader reads as prose. Two articles can both report 1,500 words and contain 1,500 and 900 words of actual argument.

Even reading time is a convention rather than a measurement. The header on this post divides by 200 words a minute. The word counter divides by 238, a figure from a 2019 review of reading-rate research on English non-fiction. Neither is wrong. They are different agreed constants, which is what most content metrics turn out to be when you look underneath them.

Keyword density measures your denominator

Density is worse than word count, because it is a ratio of two numbers that are both unstable.

That same JWT paragraph uses token three times. Counting every word, that is a density of 5.66%. Setting function words aside — a legitimate way to count, and a toggle in the frequent words extractor — it is 10.71%. Same text, same three uses, two densities that differ by nearly a factor of two. A number that depends that heavily on how you tokenise is not a number a search engine could be scoring you against.

There is also a trap in the arithmetic that the length advice never mentions. Density is count over total. Write more words about anything and your keyword density falls without you touching the keyword. "Reduce your keyword density" and "write longer content" are the same instruction in different clothes, and following both at once is how a 400-word answer becomes a 1,500-word article with 400 words in it.

Google's spam policies name keyword stuffing as a violation. They do not name a percentage, and no percentage has ever been published, because a published threshold would be trivially gamed. Run the frequency ranking on your own draft for a duller reason: it shows you the word you are leaning on. If simply appears eleven times, that is a writing problem, and it was a writing problem before anyone mentioned search.

The limits that are real

Some counts do bind, and every one of them is a character limit rather than a word limit.

Google truncates the displayed title link at around 60 characters and the description snippet at around 160. "Around" is doing work there: the cut is by rendered width, not by character, so a title in capitals runs out of room sooner than one in lowercase. The page ranks the same either way — what you lose is the part of your own sentence the searcher never sees. Front-load the distinguishing words and the truncation costs you nothing. The meta tag generator previews both against a mock result row for that reason.

Worth knowing alongside that: Google rewrites title links when it judges its own version more useful, and frequently replaces your description with a passage lifted from the page. You are writing a strong suggestion, not a specification.

What to do instead of counting

The question that matters is whether a reader who arrived with a question leaves without needing another tab. That is not a length target, and it cannot be measured in advance, which is exactly why the industry substituted a number that can.

Two habits beat a word target. Write the answer, then cut every sentence that exists only to reach a number — a section that survives that cut has earned its length. And check counts at the end rather than the start, against limits that actually exist: the title, the description, the form field you are pasting into.

This post replaced one of 332 words. The replacement is longer, but the length is not what makes it better. What makes it better is that the old one said Google uses LSI keywords to prove topical authority, and it does not, and now the page says so.

Tools covered in this guide