AI Search Tests

Can AI Search Cite a New Website?

A repeatable field test of whether AI search engines discover and cite a brand-new website, with the prompts, timeline and citation results documented.

Paper-craft magnifying glass over a search bar, charts and a rising graph, illustrating an AI search citation test.

Current finding: Early signals point to retrieval and mentions arriving before a direct citation, and to passage clarity mattering more than domain age. No final conclusion has been published.

This study asks one narrow question. When a brand-new website publishes original, verifiable material, how long until it starts showing up in AI search answers, and what does it take to get there? We are running it in the open and publishing the protocol, the change log, and the limits while the test is live.

Why a New Site’s First AI Citation Is Worth Testing

Classic SEO taught everyone the same lesson. A new domain waits months, sometimes years, to earn the links and authority that move it up Google. AI citation appears to work on different rules. Answer engines pick sources by whether a specific passage answers the question and holds up against other sources, not by how old or well-linked a domain is. Reported research on how these engines choose even suggests they lean slightly away from the largest domains, and that most cited pages do not sit in Google’s top results at all.

If that holds in practice, a small, new site could earn a citation far sooner than it could earn a ranking. Picture a new software company with a strong product and no backlinks. On Google it faces a wall, since the pages that rank have years of authority behind them. In an AI answer, the same company might get named the week it publishes a clear, honest comparison backed by its own numbers, because the engine needs that specific answer and does not care how old the domain is. We wanted to know whether that story is a real path or a hopeful one, and what a new site would actually have to do to make it happen.

How the Experiment Is Set Up

We fixed the method before publishing anything, so the setup could not bend toward a hoped-for result. First we wrote a set of prompt families, the natural questions a person would ask an assistant about our topic, and recorded the baseline answers across four engines: ChatGPT, Perplexity, Google’s AI Mode, and Claude. That baseline captures who gets cited before our pages exist.

A prompt family groups the different ways people ask the same thing. For one topic that might be “how do AI engines pick sources,” “why does ChatGPT cite some sites and not others,” and “does ranking on Google get you cited by AI.” We record the full answer and every named source for each, so when our page later appears we can point to the exact prompt, date, and engine where it showed up. Saving that baseline is what lets us claim a change was ours rather than a shift the engine made on its own.

We chose four engines on purpose, because they work differently and a new site can win one while staying invisible in another. Perplexity and Google’s AI Mode search the live web for most queries, so they can surface a fresh page quickly. ChatGPT and Claude lean more on what they learned during training, then reach for live results in some cases, which tends to slow how fast a brand-new source shows up. Checking all four keeps us from mistaking one engine’s behavior for a general rule.

Then we published the source pages with the basics an engine needs to find and read them. Each page carries a descriptive title, crawlable internal links, and a clear, self-contained answer near the top of every section. We allowed the AI crawlers in the robots file and confirmed the pages were indexed in both Google and Bing, since an engine cannot cite what it cannot reach. From launch, we repeat the same prompts on a fixed weekly schedule and log what changes.

What the Test Varies

A test means little if everything moves at once, so we hold the important things steady. The underlying claims stay the same, and the publishing identity stays the same, so we are not shifting the message and then crediting the format for the change.

Around that fixed core, we vary three things. The first is passage format, comparing a direct answer stated first against the same fact wrapped in buildup. One version opens a section with “AI engines cite sources they can verify across several places,” while the other reaches that point only after two paragraphs of context. The second is source specificity, comparing original data we gathered against a restatement of what others already published. The third is internal linking, comparing an orphaned page against one supported by clear contextual links from related pages.

We picked these three variables because the research on how engines choose sources points at all three: a passage the model can lift cleanly, a claim it can verify, and a page it can reach through links. If any one of them turns out to move citations on its own, that is worth knowing, and if they only work together, that is worth knowing too.

We change one variable at a time wherever we can, and where we cannot, we note the overlap so we do not overclaim. This is slower than shipping every idea at once, and it is the only way to say with any confidence that a lift came from the direct-answer format rather than from the original data or the extra links.

How We Measure Retrieval, Mentions, and Citations

The single most useful decision was to record three outcomes separately rather than lump them into one idea of visibility. Retrieval means an engine fetched and considered the page, which we infer from live-search engines that show their working. A mention means the answer names our brand in the text without crediting a specific page. A citation means the answer points to the exact page as the source of a claim.

A mention tells us the model has heard of us. A citation tells us it trusted one page enough to stand behind it. Tracking them apart shows the path a new source actually travels, instead of collapsing it into a single yes or no. It also tells us where a page is stuck. A page that gets retrieved but never mentioned has a different problem than one that gets mentioned but never cited, and the fix for each is not the same.

What We Are Seeing So Far

The experiment is still running, so treat everything here as a preliminary signal, not a conclusion. A few patterns have shown up often enough to note while we keep collecting data.

Retrieval-first engines that search the live web for every query pick up a new, crawlable page sooner than engines that lean on training data. Mentions have tended to appear before direct citations, which fits the idea that broad recognition builds before a model will point at one page. Pages that state the answer plainly at the top have shown up more than pages that bury the point, and an original figure has drawn attention that a restated one has not.

The gate that mattered most in the early weeks was plumbing, not prose. Until a page was indexed in Bing, the engine that leans on that index would not surface it at all, no matter how clear the writing. Once indexing landed, the content signals started to matter. That ordering fits the theory: reachability decides whether you are eligible, and clarity decides whether you win.

The engines have also disagreed with each other more than we expected. A page that one engine named as a source went unmentioned by another in the same week for the same question, which is a reminder that there is no single ranking behind AI answers to optimize toward. We treat that divergence as a finding in itself: visibility has to be tracked per engine, and a win in one place says little about the others. None of this is settled. We will publish a firmer read only after enough repeated checks exist to separate a real pattern from the normal week-to-week variation these engines produce.

What This Test Cannot Tell You

Honesty about limits is part of the method. Answer engines personalize and sample their responses, so two people asking the same question can see different sources, and the same prompt can shift from one week to the next. The study also runs on a single site in one topic area, which means a pattern we see may not transfer to a different field or a larger domain. Results from one engine or one prompt family do not automatically hold for another.

There is also a moving-target problem no protocol can remove. These engines update their models and their retrieval systems on their own schedule, so a behavior we record this month can change next month for reasons that have nothing to do with our pages. We cannot see inside the ranking system either, so we infer retrieval from the sources an engine chooses to show rather than from any log it hands us. Because of that, we treat a single strong week as a hint, not proof, and we weight patterns that repeat across many checks and more than one engine far more than any one striking result.

We also keep a clear line between an observation and a claim. An observation is something we saw on a given date and engine. A claim is something we are willing to say holds in general, and it has to clear the bar below before we make it. Holding that line is what separates a real experiment from a marketing story dressed up as one, and it is the reason our early notes read as cautious rather than confident.

How to Run This Test on Your Own Site

You do not need our setup to try the same thing. Write down the questions your customers would ask an assistant, record the answers across engines today as a baseline, then publish or rework one page to answer one question directly and keep it crawlable and indexed in Google and Bing. Rerun the prompts weekly and log retrieval, mentions, and citations apart. Our AI visibility audit skill lays out that routine step by step.

Keep the first test small. One page and one question give you a clean read, and they make it obvious which change moved the needle. Once you trust the routine, you can widen it to more pages and more prompt families without losing track of what caused what.

What Would Count as a Result

No final conclusion is available yet. We set the bar for a real finding before the data could tempt us to lower it. A result means a documented timeline from publish to retrieval to mention to citation, repeated across several weekly checks and visible on more than one engine, clearly above the background noise of answer variation. Anything less stays in the observations column.

We also want to answer the practical version of the question, not just the yes-or-no one. If a new page can be cited, how long did it take, which engine named it first, and which of the three variations came along for the ride? Those details are what would make the finding useful to someone launching a site next month, rather than a headline that a new site can get cited with no sense of the cost or the wait.

The protocol and the change log stay public while the study runs. For the mechanism this experiment probes, see our explainer on how AI search engines choose which sources to cite, and for the content signals that seem to earn a citation, see what AI search cites.

Site search

Find research and resources

Type at least two characters to search.