Why your site does not appear in AI answers
An answer engine is not looking for the best page: it is looking for the paragraph it can copy without getting it wrong.
More and more people settle a question by asking ChatGPT, Perplexity or the summary Google puts above the results. Those answers cite pages. And many companies with a decently ranked site discover that they never appear in them.
It is not bad luck or a secret algorithm. It is that the work is a different one, and it helps to separate three things that usually end up mixed on the same invoice.
SEO, AEO and GEO are not the same thing
| What it aims for | What you gain |
|---|---|
| SEO | |
| Appearing in the list of results | A click through to your site |
| AEO | |
| Being the answer the search engine shows at the top | Visibility, sometimes without a click |
| GEO | |
| Being cited inside a generated answer | A mention with a link, and credit as an authority |
The practical difference is this: a classic search engine decides which page to show; an answer engine decides which sentence to copy. Optimizing for the first means making the page good. Optimizing for the second means making sure there is, inside the page, a paragraph that can be extracted without getting it wrong.
What has to happen for an AI to cite your page
With different mechanics, every answer engine needs the same things. And it fails in the order they appear.
1 · That they can get in
The crawlers of answer engines are different from those of search engines, and many sites block them without knowing it —sometimes because of an inherited setting, sometimes because a firewall treats them as odd traffic.
If you want to be cited, those crawlers have to be explicitly allowed in the robots.txt. And it is worth deciding deliberately, because there are two different families: the ones that search live in order to answer a question —the ones that give visibility— and the ones that collect content to train models. You can allow one group and not the other.
2 · That they can read without running anything
This is the most common cause and the most expensive to fix late.
Many modern sites deliver an almost empty document and build the content in the browser with JavaScript. The large search engines are able to run it; several answer engine crawlers are not, or not reliably. If your content only exists after code has run, as far as they are concerned your page is blank.
You can check it in ten seconds without tools: open the page source —not the inspector, the source— and look for a paragraph of your text. If it is not there, it is not there for anyone who does not run JavaScript either.
3 · That the statement stands on its own
An answer engine extracts fragments. A paragraph that begins with “this” or “as we said earlier” cannot be extracted: out of context it means nothing, and the system prefers a source that does not force it to guess.
Hence a writing rule that sounds silly and works: have every important statement repeat its subject. Not “it costs about 50 dollars”, but “an annual license for X costs about 50 dollars”. The second can be cited; the first cannot.
4 · That the question is written on the page
People ask an AI the way they would speak to a person: “how much does it cost to put a WhatsApp bot in my restaurant?”. A page that titles that section “Our messaging solutions” is not competing for that question.
Putting the literal question as a heading, with the direct answer below it in two or three sentences, is the simplest technique and the one that pays off most. The rest of the section can explain, qualify and give context: none of that gets in the way once the answer is at the top.
5 · That the entities are defined
The systems that generate answers work with entities: companies, products, places, concepts. If your company is named one way on the site, another on the business profile, another on LinkedIn and another in a directory, those mentions do not consolidate into a single entity. An identical name, address and phone number everywhere is worth more than any fine-tuning of keywords.
A detail that gets overlooked: many companies have two legitimate addresses —the registered office and the place where the work actually happens— and use them interchangeably. To a machine they are two entities. You have to choose one for the directory, the business profile and the site, and always use that one.
Structured data helps with the same thing: it tells a machine what each thing is instead of leaving it to work it out.
6 · That there is something to cite
An AI cites because it needs to back up a statement. The pages that get cited are the ones with clear definitions, honest comparisons, figures with a source and a view of their own. A page that only says its company is a leader is no use to anyone for backing anything up.
About the llms.txt file
You will hear that you have to put an llms.txt file at the root of the site so that AIs understand it better. It is worth saying this precisely.
llms.txt is a proposed convention, not an adopted standard: the project that maintains it describes itself as “a proposal to standardize” the use of that file, and no standards body backs it. Putting it there costs little and does no harm; selling it as the lever that will get you citations is an exaggeration.
There is an argument you will hear in its favor that is worth knowing how to take apart: “but OpenAI, Anthropic and Google publish their own llms.txt”. That is true, and it does not prove what it appears to. Publishing a file for others to read is not the same as committing to read anyone else's, which is what matters to you. Before spending on this, ask about the second thing and not the first.
The lever is still having content that is readable without running code and statements that stand on their own.
What we did on this site
It serves as a checkable example, because you can verify it yourself right now:
- All the HTML is generated at build time. This article is complete in the page source. The English and French versions of the site are also genuinely static pages, not translations that happen in the browser.
- Answer engine crawlers are allowed by name in the
robots.txt, and kept apart from the training ones so that each group can be decided on separately. - Every page declares its structured data, and the frequently asked questions in the structured data are generated from the visible text, so that they cannot say something different from what a person reads. If an article carries no questions, the block is not declared: an empty block is worse than none.
- One language per address, with the alternatives declared to each other — and only when the translation really exists. An article that is only in Spanish does not announce versions that do not exist.
About the scope of this article. What is here is technical judgment and verifiable decisions about how a site is built to be readable by an answer engine. Every technical claim we make about this site can be checked by opening its code.
In what order to fix it
- Check that your content is in the page source. If it is not, everything else is decoration.
- Review the
robots.txtand decide deliberately which crawlers get in. - Rewrite the headings as real questions and put the direct answer underneath.
- Unify the name and the contact details everywhere, including the address.
- Add structured data that describes what the page actually shows.
- Publish something worth citing. It is the slowest part and the only one that cannot be replaced by anything else.
The first five are technical work of known scope: they are reviewed, fixed, and then checked to confirm they were fixed. It is exactly what the studio division does when it reviews a site, and what comes out of it is a list of what is wrong, ordered by what weighs most, with what each item costs alongside.
The sixth is not solved by any tool, and it is the one that decides whether your site ends up cited or not. Publishing something worth citing is sustained editorial work: choosing the questions your customer actually asks, answering them better than whoever is already there, and doing it consistently enough for an engine to take it seriously. We do that too, and this blog is what comes out of applying the method to ourselves: you can judge it by reading.
If you want to know what state your site is in before deciding anything, the review of the five technical points is a thirty-minute meeting, no cost, no obligation to hire us.
Frequently asked questions
How do I get ChatGPT to cite my website?
Three conditions are needed. That answer engine crawlers can get in, which is declared in the robots.txt. That the content is in the HTML without needing to run JavaScript, because several of those crawlers do not run it reliably. And that there are paragraphs that stand up out of context, with the subject repeated, so that they can be extracted without ambiguity. No tool replaces having something worth citing.
What is the difference between SEO, AEO and GEO?
SEO aims to appear in a search engine's list of results and win a click. AEO aims to be the featured answer the search engine shows at the top, even if there is no click. GEO aims to be cited inside an answer generated by an AI system. The first rewards the page being good; the other two reward there being, inside the page, a sentence that can be extracted without getting it wrong.
Is there any point in putting an llms.txt file?
It is a proposed convention, not an adopted standard: the project that maintains it describes itself as a proposal, and no standards body backs it. That OpenAI, Anthropic or Google publish theirs does not mean they read yours, which is what would matter to you. Putting it there costs little and does no harm, but it does not replace what does work: content readable without running JavaScript and statements that stand on their own.
Can my JavaScript site be cited by an AI?
Only if the content arrives already written in the HTML. Open the page source and look for a paragraph of your text: if it does not appear there, the page is blank for any crawler that does not run JavaScript, and several of the ones that feed the answer engines do not run it reliably.
Sources and notes
- The technical decisions described in the section “what we did on this site” can be verified directly on navhera.com.
- llmstxt.org, the project that maintains the
llms.txtspecification and describes it as a proposal for standardization. No standards body has adopted it. - The names of the crawlers this site allows or keeps apart are in its own
robots.txt, commented one by one. The operators of those crawlers publish their names and their behavior in their documentation; it is worth consulting it, because they change. - Article marked as semi-evergreen: answer engines change their behavior frequently. Review in six months.
Written by the Navhera team and reviewed before publishing. If you spot an error, write to us and we will correct it with a note.