We audited 65 of Mexico's largest companies to answer one question: when someone asks ChatGPT, Perplexity, or Google's AI Overviews about their industry, is there anything on their website an AI engine can actually quote?
The average score was 47 out of 100. Sixty-three percent scored below 60. And not one company in the sample — zero out of 65 — had FAQ schema on its homepage.
But the most useful finding is the one that inverts the usual assumption.
They aren't blocking AI. They're handing it a blank page.
The story everyone expects is that companies are locking AI out — blocking GPTBot in robots.txt to protect their content. We found almost none of that. Sixty-four of 65 sites (98%) let every major AI crawler in: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended. Exactly one company blocked them.
So the doors are open. The problem is what the crawlers find when they walk in.
An AI engine answering a question doesn't rank pages — it extracts claims and cites sources. To do that reliably it needs machine-readable structure: schema that says this is a company, this is a service, this is a question and here is its answer. Most of these sites offer none of it. The crawler arrives, finds a marketing homepage with no structured layer, and moves on to a source it can parse.
Open door, empty room.
The seven signals, measured
We scored each site on the same seven signals our free AI visibility check uses, weighted by how much each one affects whether an answer engine can cite you.
| Signal | Weight | Passed | Failed |
|---|---|---|---|
| Structured data (Schema.org) | 22 | 35% | 57% |
| AI crawlers allowed | 18 | 98% | 2% |
| FAQ / Q&A schema | 16 | 0% | 100% |
| llms.txt | 12 | 11% | 89% |
| Clean heading structure | 12 | 43% | 35% |
| Title & meta description | 10 | 45% | 34% |
| Canonical URL | 10 | 72% | 28% |
Three results deserve attention.
FAQ schema: 0%. This is the single most extractable structure on a web page — a question and its answer, marked up so an answer engine can lift it verbatim. Every company in the sample answers customer questions somewhere on its site. None of them marked those answers up. That is a pure, unforced miss.
Structured data: 57% failing. More than half publish no JSON-LD at all. Without it, an AI engine has to infer what the company even is from prose. Entity recognition is the foundation of getting cited, and most of the sample skips it.
llms.txt: 11% adoption. This one is genuinely early — a file that points AI crawlers to your best content. Eleven percent is higher than we expected for a standard this new, and it's the cheapest signal on the list to add.
Where the gaps are worst
Scores split sharply by sector.
| Sector | Sites | Avg. score | With structured data |
|---|---|---|---|
| Education | 6 | 62 | 67% |
| Retail | 5 | 61 | 80% |
| Technology | 9 | 51 | 44% |
| Health | 4 | 50 | 25% |
| Industrial | 15 | 48 | 33% |
| Food & beverage | 6 | 44 | 33% |
| Logistics | 7 | 39 | 14% |
| Tourism | 5 | 37 | 20% |
| Construction | 3 | 35 | 0% |
| Financial services | 5 | 30 | 20% |
Financial services scored lowest (30/100) — the sector whose customers ask AI the most questions about products, rates, and requirements. Every one of those questions is being answered by someone else's content.
Education scored highest (62/100), which makes sense: universities have spent years structuring course and program data for search, and that work transfers directly to AI.
Construction had zero structured data across every site we could reach.
The single best score in the entire sample was 84/100. The worst was 18.
What this means if you're the one being asked about
AI search is a zero-sum answer. When someone asks "who does X in Mexico," the engine names two or three companies. Not ten. If your competitors are as unstructured as you are, whoever fixes this first takes the citation.
The good news is that the gap is mostly mechanical. These are not brand or content problems — they're markup problems, and they're fixable in weeks:
- Add FAQ schema to the questions you already answer. You wrote the answers years ago. Mark them up so an answer engine can quote them. Nobody in this sample has done it — it's the fastest differentiation available.
- Publish structured data for your core entities. Organization, Service, Product, LocalBusiness. This is how an AI engine learns what you are.
- Fix the heading structure. One
h1, meaningfulh2s. Extraction follows document outline; a page with no outline has nothing to extract. - Add llms.txt. A plain-text file pointing crawlers to your best pages. It takes an hour and 89% of your competitors don't have one.
- Write self-contained answers. Forty to sixty words that stand alone without the surrounding page. That's the unit an AI engine quotes.
Methodology
We selected 87 medium and large Mexican companies across ten sectors, weighted
toward the Nuevo León industrial corridor, and requested each homepage plus its
/robots.txt and /llms.txt in July 2026. Sixty-five responded (75%); the
other 22 did not resolve, timed out, or refused our crawler and were excluded
from all analysis.
Scoring is deterministic and involves no AI APIs — it inspects the returned
HTML and text files for the seven signals above, with pass earning full weight,
warn half, and fail zero, summing to a 0–100 score. It's the same logic that
powers our public AI visibility check, so any company in the sample
can reproduce its own result.
Two honest limitations. First, we scored homepages only — a company with rich schema on interior pages would score lower here than it deserves. Second, this is a purposive sample of large firms, not a random sample of Mexican business, so read it as a directional benchmark for the enterprise segment rather than a national statistic.
The short version
Mexican companies are not being shut out of AI search. They're opting out of it by omission — publishing pages with nothing an answer engine can hold onto, at exactly the moment their customers started asking machines instead of typing queries.
The fix isn't a rebrand. It's structure.
Want your own number? Run the free AI visibility check — same seven signals, same scoring, results in seconds. If you'd rather we fix the gaps for you, that's what our generative engine optimization service does.
























