Two small model doors standing side by side on a desk, one plain white and one framed in scarlet, with a single brass key lying on the surface between them, an illustration of AI search visibility resting on the same website as Google search

Probably closer to ready than the market for AI-optimisation packages would have you believe. AI search visibility runs on the things a website has always needed: pages a crawler can reach, content that exists in the HTML before anything else runs, and answers stated plainly enough to be quoted. Google is explicit that there are “no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary”.

What changed is who is reading. A site built for Google had one important reader that rendered JavaScript, crawled patiently and carried years of history about your business. It now has several, and the newer ones fetch the raw page, hold no memory of you and quote whichever source answers most directly. A Google-era build usually fails that in two or three specific places. This is a readiness check for those, not an argument for a new website.

What AI search visibility actually depends on

Start with the part Google has documented. To appear in its AI features, a page must be indexed and eligible to be shown with a snippet, meeting the ordinary Search technical requirements. Google adds that you do not need to create new machine-readable files, AI text files or markup, and that no special schema.org structured data is required. Eligibility for AI Overviews is not a separate track you opt into; it is the one you are already on.

The assistants are a second door into the same building. OpenAI documents four separate user agents, each with its own robots.txt decision: OAI-SearchBot for ChatGPT search answers, GPTBot for training data, ChatGPT-User for a live fetch when someone asks about you, and OAI-AdsBot for advertised landing pages. The settings are independent, and OpenAI notes that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Anthropic and Perplexity publish comparable guidance. None of it asks for a new file format; all of it assumes a crawler can fetch your page and find the answer in it.

Nothing on the readiness list is new. What has changed is the cost of failing it.

A small brass barrel bolt fixed to a plain pale wooden panel, drawn fully open, with one scarlet paper disc resting beside it
An open card file box holding a close row of plain cream index cards on edge, with one scarlet divider standing taller than the rest near the front of the row

Where a website built for Google quietly fails

Google renders JavaScript and has done for years, which let a whole generation of sites get away with delivering an empty shell first and the content afterwards. The AI crawlers mostly do not: on current evidence they take the HTML as served, and none of the major crawler documentation claims otherwise. If your service copy, prices, opening hours or contact details only appear once a script has run, the assistant summarising your industry is reading a page that is effectively blank.

Five failures account for most of what we find on a first look:

  • Content that only exists after JavaScript runs. Tabs, accordions and sliders built by a page builder often render fine for a visitor and arrive empty for a crawler.
  • The answer spread across six paragraphs. A page that circles a question for 400 words before answering it is hard to quote. Pages that get cited answer first, then explain.
  • Business facts that disagree with each other. A different address on the site than on Google Business Profile, an old number in a directory, a name that does not match the SSM registration. Machines resolve those conflicts by trusting none of them.
  • Key information locked inside images or PDFs. Price lists exported as JPEGs, a brochure that is the only place your scope is written down, a menu rendered as a picture. None of it is readable text.
  • Pages that time out on a real connection. Six seconds on 4G is a problem for customers first and crawlers second, but it is a problem for both.

None of these is exotic and none needs a rebuild. They are the faults that cost a little in ordinary search and now cost more, because a summary that cannot find your answer will use somebody else’s. That is the mechanism behind the pattern described in our article on why rankings can hold while traffic falls.

Three identical plain white paper cups standing in an evenly spaced row on a desk, a thin scarlet band wrapped around the base of the middle cup

The robots.txt decision nobody made

Most Malaysian business sites have a robots.txt written by a plugin, or by a developer, several years ago. Nobody has opened it since, and it now governs whether four or five AI crawlers may read the site. That is a business decision sitting in a text file, made by default.

The useful distinction is between crawlers that feed answers and crawlers that feed training. Shutting out a search crawler is not withholding something valuable from a platform; it simply means that when a customer asks, the answer describes a competitor. Shutting out a training crawler is a different trade, and the platforms have started saying plainly that it carries no search cost. Apple updated its Applebot documentation on 7 September 2026 to state that site rules for Applebot-Extended are not considered in ranking for Apple Search, which matches Google’s position on Google-Extended.

Three decisions, made once and then left alone: allow the search crawlers if you want to be cited, choose separately whether your content trains models, and check that nothing in the file blocks a crawler you meant to welcome. A stray Disallow line left over from a staging site is the commonest reason a business is absent from AI answers entirely.

What is not worth paying for yet

An industry has grown up around AI search in the last eighteen months, and some of what it sells does not exist. The clearest example is llms.txt, a proposed file listing your content for language models. Google’s John Mueller has described llms.txt as “purely speculative for now”, pointing out that the file has existed for years and none of the AI systems use it. Adding one costs an hour and harms nothing. Paying a monthly retainer for it is another matter.

The same applies to “AI schema” and to audits that score a site against an invented readiness metric. Accurate structured data is worth having for the reasons it always was, but Google has said in writing that none of it is a requirement for AI features. Spend the money on the five faults above instead: they are measurable, and fixing them also improves the site for the people who do arrive.

A readiness check you can run this week

None of this needs an agency to start. Six checks, an afternoon, and a written list of what you find:

  • Read your own page source. Open a service page, view source, and search for a sentence you can see on screen. If it is not there, most AI crawlers will not see it either.
  • Open yoursite.com/robots.txt. Note every Disallow line and every named user agent. Decide each on purpose rather than inheriting it.
  • Ask an assistant about your business. What ChatGPT, Gemini or Perplexity gets wrong about what you do, where you are and what you charge for tells you which facts are missing from the open web.
  • Check the first screen of your three most important pages. Does each answer the question in its title before it starts selling? If not, that is a rewrite, not a redesign.
  • Compare your business facts in four places. The website footer, Google Business Profile, your SSM registration and any directory listing. Make them identical.
  • Load a service page on a mid-range phone on mobile data. Anything over three seconds to the headline is worth fixing before anything else here.

Where those checks point at page structure rather than plumbing, the work is the same as answer engine optimisation for service pages: one question per page, answered near the top, in language a customer would use.

Frequently asked questions

How can I tell whether ChatGPT or Perplexity can read my website?

View the page source of an important page and search for a sentence you can see on screen. If it is missing, the text is inserted by JavaScript that most AI crawlers do not run. Then check robots.txt for any rule naming OAI-SearchBot, GPTBot, ClaudeBot or PerplexityBot.

Does structured data help with AI search visibility?

It helps machines read a page accurately, which is worth having. It is not a requirement: Google states no special schema.org structured data is needed to appear in AI Overviews or AI Mode. Add it because it describes your business correctly, not because a tool promises citations.

Do I need a new website to be visible in AI search?

Rarely. The faults that keep a site out of AI answers are usually rendering, crawler access, buried answers and inconsistent business details, all fixable inside an existing WordPress build. A rebuild is the honest answer only when the platform cannot serve readable HTML or cannot be made fast.

What we check first

When a client asks whether their site is ready for AI search, we do not start with the AI. We start with what a crawler receives, what a service page says in its first hundred words, and whether the business describes itself the same way everywhere. Our SEO and search visibility work covers those three in order, and in most cases it turns out to be a revamp of pages that already rank rather than anything newer or more expensive.