SiteRecon · methodology
What is checked, and how it is scored
A module score is the share of its check weight that passed, worked out in code. The overall score is the weighted mean of the modules that could be measured. A model never sets a score. This page is generated from the checks themselves, so it cannot drift from them.
SEO · 30% of the overall score
A plain fetch of the homepage, robots.txt, sitemap and response headers. No AI.
| The homepage returns HTTP 200 | 10 |
| The site is served over HTTPS | 8 |
| The page is not blocked from search (no noindex) | 10 |
| Has a title tag | 8 |
| The title is 15 to 60 characters | 4 |
| Has a meta description | 6 |
| The description is 70 to 160 characters | 3 |
| Has exactly one H1 heading | 6 |
| Has a canonical link | 4 |
| Declares its language | 3 |
| Has a mobile viewport tag | 6 |
| Images have alt text | 5 |
| Has a robots.txt file | 4 |
| Has an XML sitemap | 6 |
| Has structured data (JSON-LD) | 6 |
| Has social preview tags | 4 |
| Sends an HSTS header | 3 |
| The HTML is under 2 MB | 4 |
AI visibility · 20% of the overall score
The same fetch, read for AI crawler rules, the text in the HTML and structured data. No AI.
| GPTBot (ChatGPT) is allowed in robots.txt | 10 |
| ClaudeBot (Claude) is allowed in robots.txt | 10 |
| PerplexityBot is allowed in robots.txt | 10 |
| Google-Extended (Gemini) is allowed in robots.txt | 10 |
| The content is in the HTML, not only added by JavaScript | 20 |
| Organisation schema identifies the business | 12 |
| Questions are answered under question headings | 10 |
| Has a 40 to 200 word paragraph that can be quoted | 12 |
| Has an llms.txt file (a small bonus) | 6 |
Content and conversion · 20% of the overall score
Fixed checks on the page text (80 points), plus a small AI review (20 points). The AI's answers only count when they quote a specific passage that is really on the page.
| Has a clear call to action | 25 |
| The headline is 3 to 20 words | 15 |
| Has a way to contact the business | 15 |
| Shows social proof (reviews, ratings, guarantees) | 10 |
| Has a menu with three or more links | 15 |
| AI review: the value is clearAI, must quote the page | 5 |
| AI review: it says who it is forAI, must quote the page | 5 |
| AI review: the next step is clearAI, must quote the page | 5 |
| AI review: something concrete sets it apartAI, must quote the page | 5 |
Speed · 15% of the overall score
Google PageSpeed Insights, one mobile lab run. Needs a free API key.
| Mobile performance score of 90 or more | 40 |
| The main content appears within 2.5 seconds | 20 |
| The layout does not jump (shift under 0.1) | 15 |
| The page is not frozen by scripts (under 200 ms) | 15 |
| Something appears within 1.8 seconds | 10 |
Social media · 15% of the overall score
Social links found on the homepage, then each public profile page read without logging in. Pages that need a login are reported as unchecked.
| Links two or more social profiles | 25 |
| No placeholder links to a platform's front page | 15 |
| Links Facebook, Instagram or LinkedIn | 15 |
| Lists the profiles in the Organization sameAs field | 10 |
| Every linked profile exists | 25 |
| A linked profile posted in the last six months | 10 |
Where the data goes
- Your pages are fetched by SiteRecon itself with the user agent
SiteReconBot, respecting robots.txt. The optional browser step then opens your homepage in Chromium, which identifies itself as an ordinary browser and does not check robots.txt for the files the page itself loads. - Text from your homepage is sent to a free AI service (NVIDIA or Google) for the content review, to name competitors, and to suggest marketing ideas.
- Your site address is sent to Google PageSpeed Insights for the speed measurements.
- Public social profile addresses are opened through GitHub's public API, yt-dlp (for YouTube) and Jina Reader (for everything else), so they can be read without logging in.
- Your domain is sent to the Hacker News search to count public mentions, and to an optional web search (Tavily) to find competitors.
- Up to three competitor sites that the AI names are fetched (their homepage, robots.txt, sitemap and llms.txt), at most three times an hour for any one site. The AI chooses them from your page, so treat them as suggestions.
- Reports are stored for 30 days, and a repeat scan of the same site within 24 hours reuses the stored report. Each finished scan is also logged, with your domain as a label, in a private MLflow.
What SiteRecon will not do
- Log in to any platform. Pages that need a login are listed as not checked.
- Get around a site that blocks crawlers or serves a bot challenge. That is reported, and the site is not scanned.
- Scan private or internal addresses.
- Present the AI's opinion as a measurement. Anything the AI says must quote your page or it is dropped.