The six dimensions below are the ones behind Avid's AI readiness check for publishers. The front end and back end framework is the standard we brief titles against on a live piece.
A title can be the best-reported publication in its category and still be unquotable. Not badly written, not badly ranked: unreadable to the software writing the answer. One line in a file your commercial team has never opened is enough to do it, and nobody finds out, because there is no error message and nothing visibly lost. The answer gets written from somebody else's pages instead.
An AI answer is assembled out of pages, and most of them do not belong to the brand being asked about. Foundation Marketing and AirOps, May 2026, found that 90% of the sources AI cites when answering are ones the brand does not control. The market the sample covers is not stated. Read the study. From your side of the desk, that is the brief: the pages engines quote are publisher pages. What decides whether any of that reaches you is mechanical, and most of it has nothing to do with the writing.
Access is a precondition, not a score
Our readiness check asks eight questions across six dimensions, in this order: access, readability, content standard, evidence, commercial, partnership fit. That order is not a ranking of importance. It is the sequence an answer moves through, and each dimension only means something once the one above it holds. Which is why access does not average in with the rest. It caps the result.
A crawler that cannot read the page makes everything below access unearnable, so the check does not let a title in that position score its way up on the other ten questions. Score one or less of four on access and the reported band stops at the second of the four, whatever the total says.
What each one is actually asking:
- Access. Can the crawlers that answer questions read your pages, and could you open one commissioned article to them as a term of a deal without changing your site-wide position?
- Readability. Is the full article text in the HTML on first load, at one stable address, with Article and Organization structured data carrying author and dates?
- Content standard. Does the writing carry something quotable: a number, a comparison, a named product, a direct answer, an expert byline?
- Evidence. Do you know which of your pages get cited today and on which topics, and can you separate AI referrals from direct in your analytics? Those visits arrive unlabelled, not missing.
- Commercial. Is any of this a product with a name on the rate card, and does editorial still write the piece?
- Partnership fit. Can you go from brief to live inside two weeks, and will you carry a claim other titles are carrying that month, in your own words?
The first two are technical and cheap. The last two are slower and commercial, and they are where a readable title becomes a bookable one.
A bigger domain does not get you cited
Once a page can be read, size stops helping. Surfer, Joshua Hardwick, 14 July 2026, found effectively no correlation between a domain's authority score and whether AI cites it, at Spearman -0.073 to +0.026. The market is not stated. Read it.
The masthead down the road is not cited because it is the masthead. It is cited because a specific claim sat on a page a crawler could read, in a shape a model could lift. For a mid-sized specialist title that is the whole opening: a standard you can meet this month rather than a position you have to buy over years.
The other half is repetition. One article on one site is a data point a model can weight or ignore. The same claim carried by several unaffiliated titles inside one window reads as the settled view of a category. It is the reason a campaign starts at four publishers rather than one, and the reason the slowest title sets the date.
What a quotable page looks like
Split it in two, because at a publisher the halves are owned by different people. Get one and miss the other and the page still does not work.
The front end: what the content says
This is what the brief asks the writer for, settled on a live piece rather than handed over as a rule sheet.
- Self-contained claims. Each claim a plain sentence that stands alone, out of context, inside somebody else's answer.
- Question-led headings. Reframed as the questions buyers actually ask.
- Answer first. The direct answer above the narrative, not after it.
- Evidence attached. Figures, sources and dates on the claim itself, not gathered at the end.
- Entity clarity. The product or company named alongside the claim, never implied.
- Comparative formatting. Where it fits, tables a model can quote in parts.
The back end: how the content is published
Your technical team's work, raised at booking rather than at delivery.
- Crawler permissions. Access for the answering bots, not only for the search engines.
- Access settings. No noindex, and no paywall or metering on the campaign URL. A noindex tag is not a Google-only switch: unless it is scoped to one named crawler, it is addressed to every robot that reads the page.
- Rendering method. Body copy in the HTML rather than loaded by script afterwards, because several crawlers do not run JavaScript.
- Structured data. Article, author and FAQ markup on the page.
- Link architecture. A clean slug, correct sitemap placement, internal links from related editorial.
- Freshness signals. Update dates that reflect a genuine change, and a URL that stays live afterwards. Citations build over months.
The back end is where a campaign quietly fails, and the branded content template is where we find it rather than the editorial one. In one assessment this year, a title open to every AI crawler on its editorial and cited regularly in our tracking was running its commercial work on microsites carrying a plain noindex tag. The tag was the only difference between the two, and it had been added on purpose, by a team who believed it spoke to Google alone.
Two kinds of crawler, and only one of them can cite you
Most robots.txt files have never made this distinction. Take it to your technical team with a decision in mind: two different jobs, two different sets of bots.
Bots that collect text to train a future model. GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot and Bytespider.
Bots that build and fetch the pages an answer is written from. OAI-SearchBot, Claude-SearchBot and PerplexityBot index ahead of the question. ChatGPT-User, Claude-User and Perplexity-User fetch a page at the moment somebody asks one. Both jobs feed answers, which is why they belong on the same side of the line.
Only the second group can put your page inside an answer. Blocking the training crawlers costs you nothing in today's answers. It is a rights and licensing decision, not a visibility one, and plenty of publishers block those crawlers deliberately because they intend to sell that text rather than give it away. That is entirely compatible with being cited. What is not is blocking the answering crawlers, which is what a site-wide block usually does by accident. Worth knowing that the user-initiated agents are the loosest part of this: a fetch made because a person asked for it is not treated as ordinary crawling, so robots.txt is a reliable lever over the indexing crawlers and a weaker one over the rest.
Google-Extended sits in neither list, and it is the line most often misread. It is not a bot. Google's own documentation says it "doesn't have a separate HTTP request user agent string" and that "Crawling is done with existing Google user agent strings": it is a robots.txt token, not an agent that visits you. What it governs is whether content Google has already crawled may be used to train future Gemini models and for grounding in Gemini apps and Vertex AI. So blocking it is not quite free either. The training half is a rights decision. The grounding half is a live answer.
What it is not is the AI Overviews control. The same page says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". On AI inside Search, Google's guidance says "AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access". The crawler list, and the AI features guidance. There is no AI Overviews switch that sits apart from Search: access is the Googlebot directive, and short of that you have the snippet controls, nosnippet, data-nosnippet and max-snippet, which limit what Search may show from a page rather than removing the page. A publisher blocking Google-Extended to stay out of AI Overviews has given up Gemini training and grounding and changed nothing in Search.
OpenAI labels its own crawlers the same way: OAI-SearchBot for search, ChatGPT-User for user-initiated visits, GPTBot for training, plus a separate agent that checks pages submitted as ads and cites nobody. The lists change, so review them twice a year. New agents appear, and a block you never chose can creep back in with a template change.
Where to start
Access and rendering first, then one commissioned piece written to the front end standard on a template that meets the back end one.
If you would rather know where your title sits before you move anything, ask us. The AI readiness check is eight questions, no sign-up, and your answers stay in your browser unless you send them. Or talk to our partnerships team and we will tell you what we see from the buying side: which of your templates can be cited today, and which of your topics advertisers are asking us about.