Blog · Technical
Can an AI engine read your store? Three checks
Before asking whether AI engines cite your products, find out whether they can read them. Three checks in twenty minutes: robots.txt, llms.txt and structured data. On our own site one of the three was in bad shape.
Before asking whether AI engines cite your products, it is worth answering a duller question: can they read them? Three checks, twenty minutes, and on our own site one of the three was in bad shape.
First: can the crawlers get in?
The robots.txt file, at yoursite.com/robots.txt, tells automated programs what they may read. For a few years now answer engines have had their own, with their own names, and some sites block them without knowing: sometimes from a decision made years ago, sometimes because a plugin added restrictive rules on its own.
The check is trivial: open that file and look for blocks. If it says User-agent: * followed by Allow: /, you are open to everyone. If you find lines naming specific crawlers with Disallow, you are excluding somebody, and it is worth knowing who.
The decision is not automatic: blocking an AI crawler is legitimate if you do not want your content ending up in somebody else's answers. But it should be a choice, not an inheritance.
Second: do you have an llms.txt, and is it current?
It is a recent convention: a text file that compactly explains what you offer and where the pages that matter live. It is not a standard anyone enforces and it guarantees nothing, but it costs an hour and removes the need for engines to guess your structure.
The practical problem is not creating it, it is maintaining it. On our own site it had existed for a while and contained no case studies, no guides, and of the Italian pages only the home link: it described a site that no longer existed. An out-of-date llms.txt is worse than none, because it confidently points at the wrong version.
The practical rule: if you publish a page that matters, it goes in at the same moment. Otherwise in six months it describes last year's site.
Third: do your pages declare what they are?
Structured data is a block of code telling a machine "this is a product page, this is the price, these are the frequently asked questions". Without it, an engine has to infer everything from the text, and sometimes infers wrong.
The check is to look at the page source and search for application/ld+json. If there is nothing, no page on your site is declaring what it is.
Here too something we found on ourselves works as a warning: the schema was there, but only on the pages that explain the topic, meaning blog and guides. The commercial pages, the ones you want cited when somebody searches for what you sell, had none at all. Forty pages out of seventy-eight. So the check is not "do I have structured data", it is "do the pages that matter have it".
What structured data does not do
Marking up does not create. If the information is not on the page, declaring it in the schema does not make it exist: engines compare the two, and a mismatch does more harm than an absence.
Content first, markup second. It is boring and it saves a great deal of wasted time.
What still cannot be checked
These three checks tell you whether you are readable. They do not tell you whether you are cited, and that measurement does not exist today in any reliable form: no widely available tool tells you how many times a generative engine used one of your pages, for which question, in which language.
What you can do is more artisanal: run the questions that matter in your category by hand and see who gets cited, and watch the traffic arriving from AI chat referrers, which is small but growing. Anyone promising you a precise count is estimating.
How Uptonica handles this
Product fills the attributes and rewrites pages so they state the data plainly, which is the part structured data can then declare. Without the content on the page, marking it up achieves nothing.
See what Product doesFrequently asked questions
Should I block AI crawlers?
It depends what you sell. A publisher living off advertising on its own pages has serious reasons to block them. A store that wants its products recommended has the opposite interest. The wrong thing is having decided without knowing.
Is llms.txt genuinely useful, or a fad?
Today no engine guarantees it uses it, so take it for what it is: an hour of work with a possible benefit and no risk. The cost is so low that the right question is not whether it helps, but whether you have an hour.
Does Shopify already generate structured data?
Most themes generate some, usually on product pages. Almost never on collection pages, informational pages or FAQs. It is worth checking page type by page type rather than assuming the theme does everything.
How often should this be rechecked?
robots.txt when you change theme or install anything touching SEO. Schema when you add new page types. The llms.txt file every time you publish something that matters, because it is the one that decays fastest.
Where do I start if I am short on time?
With the second check inverted: see whether your commercial pages have structured data. It is the most common gap and the most expensive one, because it affects exactly the pages you want to be found on.