Yōtei

Insights

How to get cited by ChatGPT

What the studies actually show moves citations, what does nothing, and the order to do it in. Written from doing it on this site.

Put the answer in the first two sentences under every heading, because Kevin Indig’s analysis of 1.2 million ChatGPT answers found 44.2 percent of citations come from the first 30 percent of a page. Use numbers and direct quotes, which lift visibility by roughly 31 and 41 percent in benchmark tests. Take a position, since a threshold gets quoted and "it depends" cannot be. Then get mentioned on pages that already get cited, because over a quarter of US citations come from the Wikipedia and Reddit layer and no on-site change competes with that. Schema helps machines classify you. An llms.txt file does close to nothing measurable.

What actually gets a page cited?

Answering the question in the first two sentences under each heading. Indig’s study of 1.2 million ChatGPT answers found 44.2 percent of citations came from the first 30 percent of the content, a pattern he calls the ski ramp. A model lifts a passage, not a page, so the passage has to be self-contained.

The same analysis found heavily cited text ran about 20.6 percent entity density, three to four times ordinary English. Entities are the concrete nouns: named tools, prices, platforms, thresholds. That is the mechanical version of the same advice, which is that vague writing is not quotable writing. If a sentence would survive being pasted into an answer with no surrounding context, it can be cited. If it needs the paragraph above it, it cannot.

Do statistics and quotes really help?

Yes, and it is the cheapest change on this list. Benchmark testing puts the lift at about 31 percent for adding statistics and about 41 percent for adding direct quotations.

The reason is the same as the ski ramp: a number is self-contained and an adjective is not. "Cheap" cannot be cited. "Twenty two of the seventy one tools are under fifty dollars a month" can. This also forces honesty, because once you commit to numbers you have to go and check them, and the checking is what makes the page worth citing in the first place.

Why does taking a position matter?

Because a threshold is quotable and a hedge is not. "Do not start UGC before five thousand a month" can be lifted into an answer whole. "It depends on your goals and budget" gives a model nothing to work with.

The cost of a stance is that you have to keep it consistent. If two of your pages give different numbers for the same question, a model reading both has no reason to trust either, and you have spent authority to gain nothing. Pick the number once and use it everywhere, including the FAQ, the schema and the llms.txt file.

What matters more than anything on your own site?

Being named on pages that already get cited. Over a quarter of US citations come from the Wikipedia and Reddit layer, and genuine participation there beats any on-site change you can make.

For a tool or a directory this means the roundups, the comparison posts and the "best X for Y" lists your buyers already read. Those pages are already in the model’s reading; yours is not yet. Getting a line in one can show up in AI answers within days, where a new page of your own takes months. Reddit works the same way and only if you are actually answering the question. A link drop reads as a link drop to both the subreddit and the model.

What about schema and llms.txt?

Schema is worth doing and llms.txt is close to worthless. SE Ranking analysed nearly 300,000 domains and found zero correlation between having an llms.txt file and how often AI engines cited a site. OpenAI has made no commitment to it.

Keep the file anyway. It costs an afternoon, it is trivially maintained from your existing data, and the downside is nothing. Just do not count it as work. Schema is different: FAQPage, HowTo, ItemList and Article give machines an unambiguous read on what a page is, which speeds up classification even though no engine promises a ranking benefit for it. Do both, expect the citations to come from the writing and the mentions.

Not yet

  • Paying for an AEO tracker before you have pages worth tracking. There are free CLI options that check citations with your own API keys.
  • Publishing volume. Ten pages that each answer one question beat fifty that circle a topic, because a model cites a passage and volume does not make passages better.
  • Chasing schema types nobody consumes. FAQPage, HowTo, ItemList and Article cover almost everything.

Does keeping a page updated matter?

More than it did for search. Recency is a real signal inside AI answers, particularly for anything time-sensitive like pricing, and a page that is visibly current beats an identical page that looks stale.

The practical version: put a visible last-updated date on anything with numbers in it, and actually re-check them on a schedule. Prices are the worst offender and the easiest thing to get caught on. A price table that is right is one of the most citable things you can publish, and one that is a year stale is worse than not having published it.

Good to know

Is AEO different from SEO?

The plumbing overlaps and the writing does not. Search rewards a page that covers a topic; AI answers reward a passage that answers a question, which means front-loading, numbers and a stated position rather than comprehensiveness.

Do I need an llms.txt file?

It does no harm and there is no evidence it helps. A study across nearly 300,000 domains found no correlation with citations. Generate it from data you already have, spend the rest of the afternoon on the writing.

How long before I see citations?

Getting named in a roundup that already gets cited can show up within days. A new page of your own usually takes months. That order of operations is the whole strategy.

When should I start on this?

Around your first paying customers. The foundations are an afternoon with a coding agent and they compound for months, so starting late costs you months you cannot get back. Writing content for search waits until you know the words customers use.

How do I know if it is working?

Pick eight to ten prompts a buyer would actually type and check them by hand once a month in a tool that prints its sources. Record whether you appear. The number that matters is how many of the ten you are in, not your position in any one.

This expands the First stretch camp on the climb, which lays out the stages either side of it. Tools by job are on the route index, the free ones are here, and the other answers are in insights. Updated 30 July 2026.