AI search
The pages AI actually cites (and the ones it ignores)
We traced every link in 160 AI answers. Four page types won the citations, while llms.txt and schema markup didn't separate leaders from laggards.
When an AI system answers a buyer’s question, it builds the answer from a small set of pages and links to some of them. Our research into one category, price monitoring software in the US, shows which pages those are. The answer has little to do with domain size or technical markup.
We put forty questions that buyers actually ask to four AI systems (ChatGPT, Google AI Overviews, Google AI Mode and Copilot), collected 160 answers and checked every link they cited. The full method is in our white paper. This post covers one question: which pages get cited, and which don’t.
Which pages do AI answers cite?
Among the individual pages cited most often, four types stand out: a community discussion, a public pricing page, a named comparison with a competitor and a listing in an app directory. Each one answers a specific question a buyer asks.
| Page type | Example | Answers with a link |
|---|---|---|
| Community discussion | An r/SaaS thread comparing nine price monitoring tools | 9 |
| Public pricing page | prisync.com/compare-plans | 7 |
| Named comparison with a competitor | price2spy.com/price2spy-vs-prisync | 5 |
| App directory listing | apps.shopify.com/prisync | 5 |
Number of answers, out of 160, that link to the example page. USA, September 2026.
Five of the twenty most-cited pages are one vendor comparing itself by name with another. Two are pricing pages, and seven of the forty queries were about cost. Cost questions are also where answers link to a vendor’s own site most often: 79% of them do, against 65% of answers where the buyer asks for a short list of tools. To quote a price, the system needs a primary source. A demo request form doesn’t give it one.
Named comparisons work for a similar reason. When a buyer asks the system to compare two named products, only 25% of answers name three or more services. If you aren’t one of the two, you are rarely added. A page comparing you with your closest competitor is how you get into that conversation before it starts.
A third of the sources aren’t on your site
The services in the study don’t control about a third of the sources behind the answers. Reddit is linked in 11 answers, as many as price2spy.com and not far behind prisync.com with 14. YouTube follows with 10 and the G2 catalogue with 9.
A single r/SaaS thread appears in nine answers on its own. Size doesn’t buy a place either: Priceva, the site with the most US search traffic in the sample, is cited three times, all three from the same page.
In practice, the platforms where your category is discussed shape your visibility whether you take part or not.
Domain authority is not a pass
A small site can be cited more than a large one. thepricegeek.com has a domain rating of 8 and 52 visits a month, and it’s linked in seven answers. That is more often than the entire web presence of sixteen of the twenty services we measured.
To compare sites of different sizes fairly, we divided citations by traffic. thepricegeek.com earns 135 citations per 1,000 visits. prisync.com, one of the category leaders, earns 2.8, and priceva.com earns 0.6. Across the sites in that comparison, the order is almost exactly the reverse of domain authority.
The small sites aren’t cited for their links or their markup. They are cited because they published something specific and verifiable that the category’s questions need. That makes the barrier to entry lower than in classic search, which is also why a Google ranking is not an AI shortlist.
Do llms.txt and schema markup help?
In this study, no technical signal separated the leaders from the rest. We crawled all twenty domains on 21 September 2026 and compared the ten services named most often in answers with the ten named least often.
Technical signals: top ten vs bottom ten
Share of sites where the signal is present · services ranked by mention share in AI answers
- Top ten
- Bottom ten
Source: direct crawl of all 20 domains, 21 September 2026
The gaps are small and point both ways. llms.txt is somewhat more common in the top ten (60% against 40%), FAQPage markup barely differs (40% against 30%), and structured data of any kind is more common in the bottom ten (100% against 80%). Nineteen of the twenty sites don’t mention AI crawlers in robots.txt at all; the twentieth allows all of them.
Individual cases tell the same story. Price2Spy is second by mention share without a single piece of structured data. PriceShape is third with neither llms.txt nor a robots.txt file. 42Signals publishes llms.txt and isn’t named in a single answer.
None of this shows that these files are harmful, and it isn’t a reason to remove them. It’s one category observed at one point in time. What it does show is narrower: they don’t explain the gap between first and last place, so they don’t belong at the top of a work plan. Markup changes how a page is read. It doesn’t change whether the page has anything worth citing.
A checklist for getting cited
The cited pages share one thing: information that exists nowhere else, such as a named judgement about two products, a real price or a test someone ran themselves. Four actions follow from that, each with a metric you can track.
- Publish real prices on an open page. Cost questions are where answers link to vendor sites most. Track the share of cost answers that link to you.
- Publish named comparisons with your two or three closest competitors. This is a decision for product marketing and legal, not only the search team. Track your mention share in two-way comparisons.
- Publish something nobody else has: a measurement, a test, numbers from your own data. Track citations per 1,000 visits.
- Be present where your category is discussed. Know which threads, videos, catalogues and app directories the answers draw on, and be there before the buyer asks. Track cited pages outside your domain.
Structured data and llms.txt take a few days of work and showed no measurable effect in this study. Do them if you like, but don’t plan growth on them.
One caution: this is a single category, US price monitoring software, measured once, in September 2026. It shows what the cited pages have in common; it doesn’t prove that copying them will move your numbers. A repeat measurement on the same query list is the first chance to tell change from chance.
The full data is in our white paper. If you need pricing pages, comparisons and original research written to be cited, our content service builds them, and our AI search (GEO) service measures whether AI systems pick them up.