64 Grade C

CrawlGap AI-visibility audit

Wikipedia

https://www.wikipedia.org/ · 2026-09-28T13:09:21.034Z · 200 in 737ms

Partly visible: crawlers can read the site, but the parts that win citations are missing.

Render gap 0% · raw 370 words · rendered 370 words · robots.txt found · sitemap missing

Homepage screenshot
Crawler access100
Readability without JavaScript100
Machine-readable structure33
Answer-readiness33
Technical foundations80

Findings

high Machine-readable structure

No Organization / LocalBusiness structured data

Evidence
No JSON-LD found in the raw HTML.
Why it costs citations
This is the block that tells an assistant the entity name, what kind of business it is, where it operates and how to contact it. Without it the assistant has to guess from prose, and it frequently guesses a competitor with clearer markup.
Fix
Add a server-rendered JSON-LD Organization (or LocalBusiness for a physical location) with name, url, logo, description, telephone, address, sameAs links to your profiles, and areaServed.
high Answer-readiness

No headings are phrased as questions people actually ask

Evidence
7 headings, none in question form.
Why it costs citations
Assistants retrieve by matching a user question to a passage. A heading that restates the question, followed immediately by a direct answer, is the highest-probability shape for being lifted and cited.
Fix
Convert or add headings like "How much does X cost in <city>?", "How long does X take?", "Do you offer X on weekends?" — each followed by a 2–3 sentence direct answer.
high Answer-readiness

No machine-readable way to contact the business

Evidence
tel: links 0, mailto/emails 0, WhatsApp links 0. Phone-shaped strings in text: 0.
Why it costs citations
When an assistant recommends a business, the next thing it tries to supply is how to reach it. A phone number that exists only inside an image or a JS widget cannot be handed to the customer, so the recommendation goes to whoever published a plain tel: link.
Fix
Add <a href="tel:+91…">, a mailto: link, and a wa.me click-to-chat link in the server-rendered HTML, plus telephone/email in your Organization schema.
medium Machine-readable structure

No FAQPage structured data

Evidence
No FAQPage / Question markup in the raw HTML.
Why it costs citations
Question-and-answer markup is the format answer engines lift most directly, because each answer is already a self-contained, attributable unit.
Fix
Publish 6–10 real customer questions with short direct answers, marked up as FAQPage, rendered server-side.
medium Technical foundations

No XML sitemap found

Evidence
Checked /sitemap.xml and the robots.txt Sitemap: entries.
Why it costs citations
A sitemap is how a crawler finds pages that are not well linked — exactly the pages that JavaScript navigation already hides.
Fix
Generate sitemap.xml with every indexable URL and reference it from robots.txt with "Sitemap: https://…/sitemap.xml".
low Machine-readable structure

No BreadcrumbList structured data

Evidence
Not found in raw HTML.
Why it costs citations
Breadcrumbs help a retrieval system understand where a page sits in the site hierarchy, which improves how it is described when cited.
Fix
Add BreadcrumbList JSON-LD on inner pages.
low Answer-readiness

Title is 9 characters

Evidence
"Wikipedia"
Why it costs citations
A very short title omits the service and location terms people ask about.
Fix
Aim for 45–60 characters, leading with what you do and where.
low Answer-readiness

No machine-readable dates on the page

Evidence
No <time datetime> elements and no dated schema.
Why it costs citations
Assistants prefer sources they can date, and prefer recent ones when a question is time-sensitive.
Fix
Add <time datetime="YYYY-MM-DD"> for published/updated, and datePublished/dateModified in schema.
low Answer-readiness

7 of 7 images have no alt text

Evidence
Images without alt attributes.
Why it costs citations
Alt text is the only description of an image a text-based crawler receives, and it is also an accessibility requirement.
Fix
Describe each meaningful image in its alt attribute; use alt="" for decorative ones.
low Technical foundations

No canonical URL declared

Evidence
Missing <link rel="canonical">.
Why it costs citations
Without a canonical, the same content reachable at several URLs (www/non-www, trailing slash, tracking parameters) can be treated as duplicates and split its authority.
Fix
Add a self-referential absolute canonical tag to every page.
pass Crawler access

All major answer-engine crawlers are allowed

Evidence
OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, bingbot, Applebot, Amazonbot, DuckAssistBot, YouBot
Why it costs citations
Nothing in robots.txt stands between the site and AI citations.
pass Readability without JavaScript

Content is present in the raw HTML

Evidence
Raw HTML carries 370 of 370 rendered words (100%).
Why it costs citations
AI crawlers receive substantially the same content a human sees. This is the single biggest thing most sites get wrong, and it is right here.

What an AI crawler reads

Wikipedia The Free Encyclopedia Donate now | I already donated Wikipedia still can't be sold. An important update for readers in the United States. You deserve an explanation, so please don't skip this 1-minute read. This year, Wikipedia has had fewer visitors and fewer new supporters, so your visit today means a lot. We hope that Wikipedia's given you at least $2.75 of knowledge recently. If everyone who found Wikipedia useful gave $2.75, we'd hit our goal in a few hours. After 25 years, Wikipedia is still the internet we were promised—created by people, not by machines. It's not perfect, but it's not here to push a point of view. It's run by a nonprofit, not a giant technology company or a billionaire. Less than 2% of our readers donate. The rare few who give do so because Wikipedia provides them with useful knowledge. If that sounds like you, please donate $2.75. Any contribution you make today helps. Please select an amount (USD). The average donation in the United States is around $13. Many first-time donors give $2.75. If you've never given, all that matters is that you're choosing to stand up for free, open information. For that, you have our gratitude. 5 10 20 30 50 100 Other How often would you like to donate? One time Give monthly Donate now We ask you, sincerely: don't skip this. Be one of the rare readers who gives. Proud host of Wikipedia and other free knowledge projects We never sell your information. By submitting, you are agreeing to our donor privacy policy 

What a human sees after JavaScript

Wikipedia The Free Encyclopedia Donate now | I already donated Wikipedia still can't be sold. An important update for readers in the United States. You deserve an explanation, so please don't skip this 1-minute read. This year, Wikipedia has had fewer visitors and fewer new supporters, so your visit today means a lot. We hope that Wikipedia's given you at least $2.75 of knowledge recently. If everyone who found Wikipedia useful gave $2.75, we'd hit our goal in a few hours. After 25 years, Wikipedia is still the internet we were promised—created by people, not by machines. It's not perfect, but it's not here to push a point of view. It's run by a nonprofit, not a giant technology company or a billionaire. Less than 2% of our readers donate. The rare few who give do so because Wikipedia provides them with useful knowledge. If that sounds like you, please donate $2.75. Any contribution you make today helps. Please select an amount (USD). The average donation in the United States is around $13. Many first-time donors give $2.75. If you've never given, all that matters is that you're choosing to stand up for free, open information. For that, you have our gratitude. 5 10 20 30 50 100 Other How often would you like to donate? One time Give monthly Donate now We ask you, sincerely: don't skip this. Be one of the rare readers who gives. Proud host of Wikipedia and other free knowledge projects We never sell your information. By submitting, you are agreeing to our donor privacy policy 

robots.txt

# robots.txt for http://www.wikipedia.org/ and friends
#
# Please note: There are a lot of pages on this site, and there are
# some misbehaved spiders out there that go _way_ too fast. If you're
# irresponsible, your access to the site may be blocked.
#

# Observed spamming large amounts of https://en.wikipedia.org/?curid=NNNNNN
# and ignoring 429 ratelimit responses, claims to respect robots:
# http://mj12bot.com/
User-agent: MJ12bot
Disallow: /

# advertising-related bots:
User-agent: Mediapartners-Google*
Disallow: /

# Wikipedia work bots:
User-agent: IsraBot
Disallow:

User-agent: Orthogaffe
Disallow:

# Crawlers that are kind enough to obey, but which we'd rather not have
# unless they're feeding search engines.
User-agent: UbiCrawler
Disallow: /

User-agent: DOC
Disallow: /

User-agent: Zao
Disallow: /

# Some bots are known to be trouble, particularly those designed to copy
# entire sites. Please obey robots.txt.
User-agent: sitecheck.internetseer.com
Disallow: /

User-agent: Zealbot
Disallow: /

User-agent: MSIECrawler
Disallow: /

User-agent: SiteSnagger
Disallow: /

User-agent: WebStripper
Disallow: /

User-agent: WebCopier
Disallow: /

User-agent: Fetch
Disallow: /

User-agent: Offline Explorer
Disallow: /

User-agent: Teleport
Disallow: /

User-agent: TeleportPro
Disallow: /

User-agent: WebZIP
Disallow: /

User-agent: linko
Disallow: /

User-agent: HTTrack
Disallow: /

User-agent: Microsoft.URL.Control
Disallow: /

User-agent: Xenu
Disallow: /

User-agent: larbin
Disallow: /

User-agent: libwww
Disallow: /

User-agent: ZyBORG
Disallow: /

User-agent: Download Ninja
Disallow: /

# Misbehaving, requests much too fast
User-agent: fast
Disallow: /

# Sorry, wget in its recursive mode is a frequent problem.
# Please read the man page and use it properly; there is a
# --wait option you can use to set the delay between hits,
# for instance.
User-agent: wget
Disallow: /

# The 'grub' distributed client has been *very* poorly behaved.
User-agent: grub-client
Disallow: /

# Doesn't follow robots.txt anyway, but...
User-agent: k2spider
Disallow: /

# Hits many times per second, not acceptable
# http://www.nameprotect.com/botinfo.html
User-agent: NPBot
Disallow: /

# A capture bot, downloads gazillions of pages with no public benefit
# http://www.webreaper.net/
User-agent: WebReaper
Disallow: /

# Per their statement, semrushbot respects crawl-delay directives
# We want them to overall stay within reasonable request rates to
# the backend (20 rps); keeping in mind that the crawl-delay will
# be applied by site and not globally by the bot, 5 seconds seem
# like a reasonable approximation
User-agent: SemrushBot
Crawl-delay: 5

#
# Friendly, low-speed bots are welcome viewing article pages, but not
# dynamically-generated pages please.
#
# Inktomi's "Slurp" can read a minimum delay between hits; if your
# bot supports such a thing using the 'Crawl-delay' or another
# instruction, please let us know.
#
# There is a special exception for API mobilev