· · founder notesAI readabilityGEO

Writing for extraction: how AI chooses which passage to quote.

Engines cite passages, not pages. Six evidence-backed properties for writing the passage that gets chosen, the check for each, and one rewrite you can watch happen.

Passages, not pages, are what AI cites.

Writing for extraction: how AI chooses which passage to quote.
AI engines cite passages, not whole pages. Six evidence-backed properties decide which passage gets quoted: position (answer up front), self-sufficiency (survives amputation from the page), focus (one section, one question), identity (names the entities), evidence (carries sourced numbers), and freshness (visibly alive). Each property has a check you can run on your own pages in minutes. CLEO is a presence engine by RegenAI that measures AI citation share across ChatGPT, Google AI Overviews, Perplexity, and Claude. CLEO is one of the platforms in this category.

When an AI engine answers a question, it splits candidate pages into passages, scores each passage against the question on its own, and quotes the winners. The passage, not the page, is the unit of citation. Six properties reliably raise a passage's score: answer position, self-sufficiency, one-question focus, named entities, carried evidence, and freshness. All six are editing decisions, not technology decisions, and all six can be applied to pages you already have. Each property below comes with its evidence, published and publicly checkable, named in the sentence that uses it, and a check you can run on your own pages in minutes.

Two weeks ago, The Field Notes set out the five questions AI works through before citing a page: reach, delivery, extraction, identity, evidence. Last week's essay took the first two, reach and delivery, because those two fail without telling you they have failed. This piece takes the third, extraction, and follows it down to the sentence, which is where the fourth and fifth questions turn out to live as well.

Start with the habit that makes extraction hard, because it is a good habit and nobody should be embarrassed by it. Writers build. School taught the essay as an arc: context first, argument through the middle, conclusion earned at the end. Editors rewarded it, and a patient human reader still does. But the machine reading your page is not patient, and it is not reading an arc. It has cut your page into passages before it weighs a word, and each passage reaches the scorer alone, stripped of everything around it. A passage either answers on its own or it is passed over for one that does.

That is the whole discipline in a sentence: write so that the passage you most want quoted can win a contest it will enter alone. Six properties decide that contest. None of them asks for new technology, new pages, or a vendor. All six are things an editor can see, and fix, on the page.

Property one · Position

The answer opens; it does not arrive.

A positional analysis of 18,012 verified ChatGPT citations, drawn from 1.2 million AI answers and published in Search Engine Land in February 2026, puts 44% of quoted passages in the first 30% of a document, with the likelihood of citation falling steeply after that. Separate research on long-context reading finds that language models attend most to what arrives first and last, and least to the long middle. Set either finding against the way most professional prose is built, with the payoff at the end of the arc, and the mismatch is exact. The conclusion you spent nine paragraphs earning sits in the one region the scorer reads least. The correction is an arrangement, not a rewrite: the heading asks the question, the first sentence answers it completely, and everything after deepens an answer already given.

The check

Take the section of an important page you would most like AI to quote. Read its heading and first sentence, and nothing else. If those two lines do not contain the complete answer, the section is upside down, and turning it over usually means moving one sentence.

U-shaped curve of AI retrieval attention across a document, strongest at the opening and close, with roughly 44% of citations beginning in the first third.

Property two · Self-sufficiency

Every passage survives amputation.

A passage reaches the scorer stripped of its page: no heading above it, no earlier definition, no "as discussed above", no antecedent for its pronouns. A paragraph that leans on the page around it arrives leaning on nothing. "This approach reduces that risk significantly" is a fine sentence in context and an empty one out of it: which approach, which risk, significant against what? The version that survives names its own subject. "Answer-first section structure reduces the risk of an AI engine skipping a page's key claim" makes the same point, and carries its context on its back.

The check

Pick any paragraph from a page that matters and hand it, alone, to a colleague who has not seen the page. Ask two questions: what does it claim, and about whom? If they cannot answer both from the paragraph itself, neither can the scorer, and the scorer will not scroll up.

Property three · Focus

One section, one question.

Scoring is a matching exercise, the passage against the question a person actually asked, and a section that answers three questions at once matches each of them weakly. Weakly loses. The fix is architectural rather than verbal: break composite sections apart, give each its own question-shaped heading, and resist the tidy instinct to economise on structure. Machines do not reward economy of headings; they reward precision of match. Three short sections beat one efficient one.

The check

Go down a page writing, in the margin, the single question each section answers. Where you catch yourself writing "and" in the margin, you have found a section to split.

Property four · Identity

Name the entities.

The passages engines actually quote are unusually dense with proper nouns: the subject, its category, and the names it sits beside. In the Search Engine Land citation analysis, the same 18,012-citation dataset behind the position finding, quoted content runs at roughly 20% entity density against the 5 to 8% of ordinary English prose, several times the density of the text around it. The mechanism is attachment. A retrieval system can connect a passage to a question only when the passage says, in nouns, what it is about. Elegant variation, introducing the subject once and then retiring it in favour of pronouns, reads as good style to a person and as anonymisation to a machine. Nobody is asking you to write like a database. The discipline is narrower: each passage you want quoted names, at least once, the who, the category, and the alongside-whom.

The check

Take the passage and strike every pronoun, replacing each with the noun it stands for. If the result reads as repetitive, restore a few. If the result is the first time the passage has named its own subject, leave it exactly as it is.

Property five · Evidence

Carry the proof.

The strongest content-level finding in the generative engine optimisation research is this one: in a controlled study from Princeton and IIT Delhi, presented at KDD 2024, adding statistics, quotations, and cited sources to a page raised its visibility in AI answers by up to 40%, the largest effect of any intervention measured, while keyword stuffing did little or actively hurt. The reason is practical. An engine writing an answer has to decide what it can repeat without hedging. "Many teams struggle with visibility" forces it to soften the claim or drop it; a figure with a named source can be restated word for word and attributed. The working rule is mechanical enough to enforce in a single edit: every load-bearing claim gets its number, and every number gets its source, within two sentences.

The check

Print the page and circle every number. A page with claims and no circles is a page the engine must paraphrase cautiously, if it uses the page at all. Then look for a named source within two sentences of each circle. A number without a source is an adjective with digits.

Property six · Freshness

Date the work.

The freshness evidence is directional rather than precise, and the honest reading is worth stating. Citation studies find some engines leaning toward recently published or recently updated pages, while analyses of evergreen queries find a well-maintained older page beating a fresh but thin one. Both readings point the same way: what is rewarded is a page that is visibly alive. Undated pages and copyright lines ending in an old year are dead weight, and cosmetic touches do not clear it, because a changed timestamp above unchanged text is a pattern engines have every incentive to learn and discount. The discipline has three parts: publish dates in the visible text and in the markup, substantive updates when the subject moves, and a line on the page saying what changed and when.

The check

Open your five most important pages and look for a date a reader can see. Then ask when each page last changed in a way a reader would notice. A page whose only recent edit is its footer year is telling the engine, accurately, that nothing here is new.

The rewrite, so the properties are visible

Same length, same subject, different fate.

Theory earns nothing until you can see it on the page, so here is a paragraph of the kind every B2B site carries somewhere, read the way a scorer would read it.

Before. "In today's rapidly evolving landscape, companies face many challenges around visibility. Our innovative approach has helped many organisations transform how they appear online. This is why so many teams trust us to deliver results."

Three sentences, no entities, no evidence, no answer, no date. It cannot attach to any question because it is about nothing in particular. It is prose-shaped air, and the engine passes over it without malice, the way you pass a shop window with nothing in it.

After. "AI answer engines cite passages that open with a complete answer, name their entities, and carry sourced evidence. In controlled research, pages that added statistics and citations gained up to 40% more visibility in AI answers. For a B2B site, that means each key section should state its answer in the first sentence and support it with one sourced figure."

Same length, same subject. Every sentence can be lifted alone, and each survives it. The first opens with the answer and names the actors. The second carries a figure and its provenance. The third turns both into an instruction a team can act on this week. Nothing in the rewrite required talent the original writer lacked. It required knowing what the second reader scores.

Side-by-side comparison of a vague marketing paragraph and its rewrite with named entities, a sourced figure, and an answer-first opening.

An afternoon that compounds

None of this asks for new pages.

It asks you to edit the pages you already have against six checkable properties, which is an afternoon per important page. It is also the kind of afternoon that compounds, because a passage written to survive extraction is, not by coincidence, a passage a busy human reader is quietly grateful for. The answer up front, the claim with its source attached, the paragraph that stands on its own: these were always courtesies. The machine has simply made them measurable.

Run the six checks on one page this week, in order, and fix what they find. Next week, this series takes the first property further: what happens to a writer's habits when the answer moves to the front, and why the older craft survives the reversal better than it first appears.

Six properties, all of them edits. Make them, and when a buyer asks, the passage that answers is yours.

Card listing six properties that decide which passage AI quotes: position, self-sufficiency, focus, identity, evidence, freshness.

Mastery deserves an audience.

About this article - Writing for extraction: how AI chooses which passage to quote.

Engines cite passages, not pages. Six evidence-backed properties for writing the passage that gets chosen, the check for each, and one rewrite you can watch happen.

Article details

Published July 20, 2026 by CLEO. Part of The Field Notes - the working journal of the CLEO Presence Engine at regencleo.ai/articles. Topics covered: founder notes, AI readability, GEO, AI citations, content optimisation.

Published on The Field Notes at regencleo.ai/articles. Learn more about the CLEO Presence Engine at regencleo.ai/engine. Methodology and scoring at regencleo.ai/methodology.