Blog Content

How to write SEO content that holds up when checked

A writing method for SEO content where every rule has a check behind it, plus what an unchecked rule cost this blog in its own audit.

Alexander González September 10, 2026 3,789 words

This blog runs its editorial standard as a script, and the script can fail the build. In August 2026 it was pointed at the blog's own archive. Nine articles were quoting half a statistic. Fourteen attributed their figures to nobody, in the one section whose entire job is attribution. A flagship number was misattributed in eighteen, seven of them already live. Every one of those rules was already written down.

How do you write SEO content?

Writing SEO content means answering one search intent on a page where every claim can be checked: a named source behind each figure, a self-contained answer of 40 to 60 words near the top, and headings that match real questions. On a new site coverage decides, at mid authority depth does, and on a consolidated site updating beats publishing.

Quick answer

Editorial ruleHow it fails in silenceThe check that can catch itWorth automating from
Every figure names who measured itA number travels with no sender and reads as preciseA list of numbers not present in the source fileThe first article with a statistic
The answer sits in 40 to 60 words near the topThe answer drifts down the page and nothing is extractableA word count on the paragraph under the first question headingThe first article
No empty attributionA phrase like "figures from industry research" survives in the transparency blockA list of banned formulations, kept in every language you publish inThe first transparency block
Internal links land on pages that existA live article points at an unpublished draft and serves a 404An inventory check against what the build actually writesAround twenty articles
Nothing marked unconfirmed gets usedThe warning and the breach sit in the same fileOne pattern per retired figure, with an exemption for the correction noticeThe first retraction

Why more advice does not fix this

Read the guides that own this query and the overlap is almost total: find the keyword, match the intent, write an outline, use headings, link internally, write a title tag, cite credible sources. Read on 5 September 2026, the best of them, Semrush's SEO writing guide, also says to draw on firsthand experience and to have an editor check for factual accuracy. None of that is wrong. Two of the three top results do not mention sources or fact-checking at all.

The gap is not in what they recommend. It is that every one of them stops at the recommendation. "Cite credible sources" and "ensure factual accuracy" name a desirable end state and say nothing about the mechanism that would tell you the end state was missed, which makes them advice you can follow sincerely and break anyway.

That stops being academic somewhere around the twentieth article. One piece a month fits in a person's head. A corpus does not, and the failure mode is not carelessness: nothing external ever reports the breach. Google will not tell you a figure has no source, and Search Console will not either, because it measures what happened to a page and has no view of what the page claims.

The questions Google publishes, and the number it refuses to give

There is a public standard. It is shorter than every checklist written about it. Google's guidance on creating helpful content, last updated 10 December 2025 and opened for this article on 5 September 2026, publishes a set of self-assessment questions: whether the content offers original information or analysis, whether it gives a substantial description of the topic, whether it says something beyond the obvious, whether it merely rewrites other sources without adding value, and whether the headings are descriptive rather than exaggerated.

The same page settles the question that consumes most of the advice on this query. It asks, in Google's own words, whether you are "writing to a particular word count because you've heard or read that Google has a preferred word count", and answers itself in a parenthesis: "No, we don't."

It also publishes a framework of who, how and why. Who made the content, made visible through a byline. How it was produced, including a disclosure when automation or AI generation was involved. And why it exists at all, which is the one that decides the other two: content made primarily to help people, rather than primarily to attract visits from search engines. Whether that distinction is legible to a machine is a separate argument, covered in what SEO is and what it is for.

Those questions are already a checklist, published by the party that does the ranking, and the guides competing for this query quote them as encouragement rather than running them as a gate. The distance between the two is the whole subject of this article.

The order the decisions go in

Most writing advice starts with keywords and word count. Both are consequences. Starting there produces long articles that answer nothing in particular, which is the most common defect in this category. Ordering by what cannot be undone later is the same reasoning behind the website launch checklist, applied to a page instead of a site.

1. Decide which query the page answers, and only one

The decision is not which keyword to place, it is which search the page is the answer to. The fastest way to see it is the results page itself: what Google is already showing there is its answer to that intent, and a page of a different shape competes badly no matter how well it is written. Which query has demand, and which is a synonym of one you already cover, is work that happens before the outline and is covered in free keyword research.

The check behind this rule is comparative, not internal. Two pages answering the same question cannibalise each other, and no single-article review can see it, because each one is fine on its own.

2. Write the answer before you write the article

The paragraph that answers the title, in 40 to 60 words, self-contained, with the entity named inside it rather than carried in from the sentence before. It goes near the top, under a heading phrased as the question a person would actually type.

The reason it is self-contained is mechanical. Search engines rank documents; generative engines retrieve passages and paste them into an answer that is no longer your page. A paragraph that opens with "this" or "that is why" is correct inside the article and useless outside it, and outside it is where the citation happens. The discipline built around that behaviour has its own name and its own literature, in generative engine optimization.

The check is a word count on that one paragraph, and it is the cheapest one in the whole standard. Under 40 words it is a sentence rather than an answer. Over 60 it stops being liftable, because whatever lifts it has to cut, and the cut lands wherever it lands. What the same discipline looks like aimed at a single engine, with the part you influence separated from the part you do not, is in how to get cited by ChatGPT.

3. Write what does not fit in the answer

If the article can be summarised by its own opening block without loss, the rest of it was not doing anything. That test is uncomfortable to apply to your own drafts, and it is the one that sorts this category: almost everything that reads as generic is an article whose lower two thirds restate the top. What justifies the visit once a summary has answered the question is the material a summary cannot carry, meaning the conditions, the failure modes and the scenario where the advice reverses. What each kind of engine rewards is traced in answer engine optimization versus SEO.

4. Let length be the result

Length arrives last, because it is arithmetic on the previous three steps. This blog does hold a range per article type, and it treats it as a floor rather than a ceiling, on the reasoning that the rule exists against thin content and that overshooting is a different defect. What replaced the ceiling is a density measure: verifiable elements, meaning a figure with a unit, a date or a named source, per 100 words of prose. The floor is 0.7 and the corpus median is 1.5.

That change came from a measurement, not a preference. When the counter was fixed, 120 of 158 articles were outside the range, every one of them above it and none below, with a median excess of 175 words. A one-sided bias is not editorial drift. The counter was including table cells, the URL of every link on top of its anchor text, and splitting hyphenated terms in two.

When NOT to write this way

When the page sells rather than informs. On a service page or a pricing comparison, handing over the complete answer at the top in liftable form gives away the useful part without the visit. Informational queries are where click-through has fallen hardest and transactional ones hold up better, so a page that closes business is structured for the person deciding, not for the engine summarising.

When the piece is narrative or opinion, where the value is the route and there is no query to answer. The extractable block has nothing to do there.

And when the site is not indexed yet. None of this moves anything if the page never enters the set the engine retrieves from, which is why the technical layer comes first in the SEO checklist for a business with no marketing team.

What an unchecked rule cost this blog

The measurements below come from audits of this site's own archive, 185 articles across two languages. They are why this article exists in this form rather than as another list of tips.

Half a statistic, in nine articles, one day after the rule was written. The rule said that quoting the drop in click-through where an AI Overview appears, without also quoting the drop where none appears, tells half the story. It was written into the source file on one day and broken in nine articles on the next. Two of the nine were already published.

An attribution that attributed nothing, in fourteen. In the transparency block, fourteen articles had written "measurements published by third parties", an attribution without naming anyone, and the correction replaced it with a study, a date and a sample size. The same blog had already rejected the identical move once, under different wording. What the variants share is naming a category of source instead of a source, and that is why the empty phrasing keeps coming back in new clothes.

A flagship figure, misattributed in eighteen. A percentage from an academic paper sat in the project's own variables file, its highest authority for numbers, with the wrong referent attached. Eighteen articles inherited it, seven of them published. It was not laziness and it was not a bad source: it was trusting an earlier verification without opening the paper again.

Attributions the reader could not follow, in almost all of them. A measurement found that 151 of 152 articles had no external link at all, while 49 said a claim came from Google's documentation. A blog whose entire pitch is traceability was asking for faith. The fix was opening sixteen URLs one at a time, and the sources that could not be verified are still named without a link, which is the honest outcome.

Turning a standard into something that can stop a draft

A check is not a longer document. It is a small amount of code that reads the draft and refuses it. Four properties separate the ones that work from the ones that decorate a repository, and each one was learned by getting it wrong.

A check has to be able to fail. One of this blog's layout probes stayed green while the page underneath it was deliberately broken: it named a constant that did not exist in the browser, threw, returned nothing, and the calling code read nothing as "no problems found". A broken check and a clean page gave identical output. A new check earns its place by being broken on purpose once, and the case that must not fire gets tested too.

A check has to know when to stay quiet. The rule against reusing a retired figure has to let the paragraph announcing the retirement name it, or correcting yourself in public becomes impossible. The exemption is narrow: the surrounding text has to deny the provenance explicitly.

A check written in one language is half a check. With 47 articles in English at the time, three patterns turned out to be Spanish-only in practice. One searched for a phrasing the corpus wrote with an extra word. One missed the rounded version of a retired number. One exemption used Spanish word stems, so an English article failed for honestly declaring a correction.

A check that fires everywhere stops being read. The broken word counter above flagged 76 % of the corpus, and at that rate nobody reads the output, so it also stops warning about the article that really did run long. A control that shouts louder than the fact warrants gets ignored, and an ignored control is not warning about anything.

That last pair generalises past code. Two halves of a bilingual site drift apart the moment a correction lands in one language and not the other, which is a large part of why Spanish SEO is a separate body of work rather than a translation job: different demand, different results pages, different corrections pending in each half.

Measured on two pairs of sibling sites in the same niche, the Spanish half turned impressions into clicks around three times better at comparable positions. That gap survives no translation workflow, and the comparison with its caveats is in Spanish versus English SEO.

What the research says about being cited

The most cited measurement of how writing choices affect citation in generated answers is the GEO paper from Princeton, presented at KDD 2024, which tested nine optimisation methods over roughly 10,000 queries across nine datasets. Its abstract claims visibility gains of "up to 40 %".

The detail that matters for anyone writing: the strongest methods reach a 30 to 40 % relative improvement on Position-Adjusted Word Count, a metric for how much of the generated answer comes from a given source. Citing sources and giving clear structure were among the methods that worked. Keyword stuffing performed below the baseline.

And the figure most often repeated from that paper, a 41 % lift, belongs to the best-performing methods on that metric rather than to adding statistics on their own. This blog carried the misattributed reading in eighteen articles until the paper was reread on 17 August 2026. Citing a study about the value of citing studies, incorrectly, for months, is the most on-the-nose version of this argument its author is ever likely to produce.

Scenarios, because a single number here would mislead. A new site with nothing in the retrieval set gains nothing from rewriting paragraphs, since there is no passage to cite. A site with pages in the middle of the first results page is where the paper measured its largest gains. A site already holding first position has little to gain and something to lose, which makes the job defensive rather than editorial. What that implies for content written to be read by machines is in SEO for AI.

Mistakes that repeat

Data and transparency

Google's self-assessment questions, the who, how and why framework, the disclosure expectation for automation, and the sentence declining to publish a preferred word count come from Google's documentation on creating helpful, reliable, people-first content, which shows a last update of 10 December 2025 and was opened on 5 September 2026. The GEO figures come from the paper by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, arXiv 2311.09735, presented at KDD 2024, opened on 5 September 2026; the "up to 40 %" is the abstract's own wording, the 30 to 40 % band applies to the strongest methods on Position-Adjusted Word Count, and the correction of the 41 % attribution is this project's, dated 17 August 2026. That the guides currently ranking for this query converge on the same advice, and that two of the three top results do not mention sources or fact-checking, is an own observation from reading them on 5 September 2026 and is recorded as a pattern in the results, not as a source.

The counts of nine, fourteen, eighteen, 151 of 152 and 120 of 158 articles come from audits of this blog's own corpus carried out in August 2026, and the corpus size of 185 articles was measured on 5 September 2026. The keyword data for this article, 1,500 monthly United States searches with a difficulty of 46 and a cost per click of 0.70 USD, was measured in Ahrefs on 5 September 2026. The 40 to 60 word range for the answer block, the floor of 0.7 verifiable elements per 100 words and the corpus median of 1.5 are this blog's own editorial standard, not figures published by any search engine. Editorial judgment here comes from work across a portfolio recording more than 300 million impressions a year in Search Console on the sites I operate, which includes client work that is not named. Verified as of September 2026.

Primary sources, opened on 5 September 2026: Google on creating helpful content; the GEO paper.

What this changes

The advice on this query has been stable for a decade, and it is not what separates the pages that survive from the ones that quietly rot. Every one of them knew to cite sources. This blog knew it well enough to publish the rule, and broke it in nine articles the next day.

The scarce thing, then, is not knowing how to write well. It is having something outside your own attention that can tell you when you did not, because skill degrades silently and no report exists for it: search engines measure what happened to your page, never what your page claims.

That gives a strange test for a body of work. Not whether the writing is good, but whether anything in the system could have caught it being bad, and what happened the last time it did. Most publishing operations have never run that test, and the answer sits in work they have already published.

Frequently asked questions

How do you write SEO content that ranks?

By answering one search intent completely on a page whose claims can each be traced. In practice that means choosing the query before the outline, writing a self-contained 40 to 60 word answer near the top, naming a source and a date beside every figure, and letting length fall out of the material. Ranking also depends on the page being indexed at all.

Does keyword density still matter?

It was never a confirmed ranking factor and repeating a term does not improve anything by itself. What changes the outcome is whether the entity is named inside the paragraph a generative engine may lift, because that passage travels without the surrounding article. Density targets survive in content briefs because they are easy to measure, not because they predict results.

How long should an SEO article be?

Google's documentation states plainly that it has no preferred word count, and asks writers whether they are writing to one because they heard it had. A length range is still useful as a floor against thin content, and useless as a target. This blog replaced its upper limit with a density measure of verifiable elements per 100 words, which is what the word count was failing to capture.

Where does the paragraph an AI engine can quote go?

Under a heading phrased as a question, in the first third of the article, in 40 to 60 words, with no pronoun pointing back at earlier text. Generative engines retrieve passages rather than whole pages, so a paragraph that only makes sense in place cannot be reused. Position inside the document decides whether it enters the selection at all.

How do you check that an article follows your own standard?

By writing the standard as something that reads the draft and refuses it, rather than as a document. The checks that hold up share four properties: they can fail, they have narrow exemptions for the case that must stay quiet, they exist in every language the site publishes in, and they fire rarely enough to still be read. Each one gets broken on purpose once.

Can AI write SEO content?

Google's guidance asks for disclosure when automation was used, and treats content produced mainly to manipulate rankings as a spam problem regardless of how it was made. The practical constraint is different: a generated draft carries figures with no provenance, and provenance is the part that fails silently. A model that cannot open its own sources cannot check them either.

Most sites do not have a ranking problem

They have a what-happens-next problem. You can rank first and still sell nothing. The diagnostic looks at both and tells you which one is costing you money.

See the diagnostic