How we verify
Every listing on this site is the output of a rule, and every rule is on this page. Where a number appears next to a firm, this explains what it counts, what it refuses to count, and what would move it.
Inclusion ruleset 1.0.0 · Scorecard ai-seo.v1, effective 2026-08-06
1. Who gets listed at all
We list AI-specialist agencies only. On a general directory the specialist and the everything-shop appear in the same list, which makes the specialist invisible — so the first thing this site does is turn most applicants away. Of the firms we have examined, we list 6 of 15.
There are three routes in. A firm needs one.
AI work is 70% or more of the business
Measured against the firm’s own public materials, not the figure it gives us. A declared share is multiplied by a credibility factor and capped near what the record supports; where the two disagree, we publish the gap rather than picking one.
Or 2 of 5 AI-first signals, each with a public URL
A pure percentage floor would empty this category: almost every good AI SEO firm is an SEO firm that went AI-first. So the floor stays strict and this route stays generous in reach — but a signal without a source does not count, which is what stops the route from becoming a checkbox.
- A named, publicly documented AI/GEO methodology the firm authored — not a blog post about someone else's.
- AI tooling the firm built and operates itself, publicly demonstrable.
- The service menu itself is AI work only — no separately sold non-AI service lines.
- A publicly named AI-visibility measurement stack, i.e. they can show what they claim to move.
- Named individuals whose stated roles are AI-specific, not generalists with an AI title.
Or the firm is AI-native or AI-vertical
A company whose own product is the AI has no meaningful “share of AI work”. An agency that does fairly ordinary marketing exclusively for AI companies is the specialist a buyer in that market is looking for. Both are AI-specific in the way that matters; neither survives a percentage test. Every competitor collapses these into “AI agency” and loses the distinction.
We publish what we turn away, with the measured figure and the sources — but only when we actually measured it. “We could not check” is never published as “they failed”. That distinction is the difference between a directory and a liability.
2. What we check
Two ladders. A base ladder every firm faces — is this a real, trading, honest business — and a category ladder that differs for every category we cover. The AI SEOladder asks things an AI PPC firm would find irrelevant, and vice versa. A large generalist structurally cannot pass a specialist ladder, which is the point.
Is AI search actually their business?
- AI search is a named service, not a blog topic — I look for a dedicated AI search or GEO service page, not a paragraph about AI bolted onto the traditional SEO page.
- When the AI search page first appeared — I check the archive for the first time that page existed. Conviction before 2024 reads differently from a page that went up six weeks ago.
- The traditional-SEO vs AI-search split is knowable — I work out from their own site how much of the practice is AI search and how much is classic SEO wearing a new label.
Can they measure the thing they sell?
- A named AI-visibility measurement stack — I check whether they name how they measure AI visibility at all — a tracking product or their own tooling. Anyone selling AI search visibility with no stated measurement method is selling a feeling.
- A defined AI-search KPI — I look for a stated metric — citation share, prompt coverage, answer presence. Rankings is not an AI-search KPI.
- A case study reporting an AI-search number — I check that at least one result is an AI-search figure rather than an organic-traffic figure. Selling AI search and reporting Google traffic is this category's signature swap.
Who would actually do the work?
- A named search lead with a traceable public record — I look for a real person with published work, talks or studies attached to their name — not a team of experts.
- Team size against client count — I compare the team page against the client wall to see whether the ratio is plausible.
What does an engagement cost, and how do you leave?
- Stated pricing model — I look for a named model — retainer, project, performance or hybrid — rather than one hidden behind a call.
- Discoverable contract length and exit terms — I check whether the commitment and how to leave are findable without signing anything first.
Five outcomes, never flattened into a pass rate
- Verified
- We found it in public record.
- Stated
- The firm says so and we have not corroborated it.
- Pending
- We have not run this check yet. Our backlog, not their failing — so it stays in the denominator.
- Unavailable
- We tried and the source blocked us. This never renders as a failure.
- Not applicable
- No such source exists here. Leaves the denominator entirely.
Conflating “a bot wall blocked us” with “this agency failed” is how a verification engine becomes a defamation problem. The five labels exist so that cannot happen by accident.
3. The TF Score
It measures how complete and how well corroborated our evidence is — not how good the agency is.A firm can be excellent and score low because we have not established much about it yet. We say so on every page the number appears on, because a 0–100 number next to a company name will be read as a rating unless it is relentlessly explained.
60 points are universal and 40 are specific to AI SEO.
| Pillar | Weight | What it asks |
|---|---|---|
| Business legitimacy | 12 | Is this a real, trading, identifiable company — registered, contactable and continuously operating? |
| Evidence quality and delivery record | 16 | Can the work be corroborated by someone other than the firm — named clients, reachable references, published results? |
| Independent validation | 10 | Does anything outside the firm's own website confirm what it says about itself? |
| Transparency | 10 | Does the firm publish what it charges, how long you are committed, and what it is actually doing for the money? |
| Evidence freshness | 6 | How recently was any of this last checked? In AI search a two-year-old finding is about a different product. |
| Review integrity | 6 | Are the reviews that exist relevant to this work and plausibly independent, rather than a wall of five-star ratings about something else? |
| Measurable AI-search outcomes | 10 | Can the firm say how it knows whether AI visibility moved — a named measurement stack, a defined KPI, and a result reported against it? |
| Technical AI SEO depth | 10 | Retrieval, rendering and structured data — whether the technical implementation is knowable rather than implied. We do not have a check for this yet, so it is excluded from the score rather than counted as a zero. Not yet scored — excluded from the denominator, never counted as zero. |
| Defined AI SEO practice | 8 | Is AI search a named service with a named person behind it and a history, or a page added last quarter? |
| Editorial and content governance | 6 | Who signs off before anything ships, and on what basis. We do not have a check for this yet, so it is excluded from the score rather than counted as a zero. Not yet scored — excluded from the denominator, never counted as zero. |
| Ongoing delivery and maintenance | 6 | Is there capacity to keep doing the work, and terms that describe an ongoing engagement rather than a one-off? |
2 pillars are declared but not yet scored, because we have no checks behind them yet. They are removed from the denominator rather than scored zero — counting an unasked question as a failure would defame every firm equally, and dropping it silently would inflate every score equally.
What a claim is worth
| How it was established | Value |
|---|---|
| _readme | What a claim is worth by how well established it is. `not_applicable` is absent on purpose — it leaves the denominator entirely rather than scoring zero, because a check that does not apply is not a failure. |
| Publicly verified | 1 |
| Confidentially reviewed | 0.85 |
| Partially established | 0.6 |
| Self-reported | 0.35 |
| Evidence requested | 0 |
| Awaiting review | 0 |
| Not submitted | 0 |
| Rejected | 0 |
| Out of date | 0 |
Evidence ages. Each check carries its own shelf life, and a result past it is worth less than a fresh one — a verified fact from three years ago is a historical note.
Bands
- Well evidenced — 80 and above
- Evidence-backed — 65 and above
- Developing evidence — 45 and above
- Limited verified evidence — 0 and above
How the order is decided
Firms are ordered by evidence confidence, then by category delivery record, then by how recently we checked. Scoring and ranking are separate on purpose: one number cannot honestly mean both “our evidence is strong” and “look at this firm first”.
Ranking is never for sale. The scorer selects from a database view that physically cannot see payment, subscription, claim status, featured slots or reviews, and the deploy gate asserts the order is byte-identical across every billing state. Sponsored placement exists, is capped, sits above the list, is labelled on every card, and never enters the structured data an AI assistant reads. Paid features are presentation only.
4. Proof tiers
Separate from the score: how strongly a firm’s own claims are established. Tiers are derived in code from what reviewers record, never set by hand — a badge an administrator can toggle is a badge for sale.
- Publicly verified
- We checked this firm's core claims against sources anyone can open, and they held.
- Confidentially reviewed
- A reviewer saw evidence for this firm's core claims that cannot be published — client work under NDA, usually.
- Self-reported
- These are the firm's own statements, published as such. We have not been able to corroborate them yet.
- Insufficient public evidence
- There is not enough on the public record to say much either way. That is a statement about the evidence, not about the firm.
5. Tools
A firm may tell us what is in its stack and we label that as its statement. Calling a tool a specialism is a claim about capability, so it needs corroboration; without it we show the tool as a stack entry and say on the profile that we walked the claim back. Declaring tools cannot move a firm up any list.
6. Getting it wrong
We will. When we do, correcting a factual error is free and open to everyone — claimed profile or not, paying or not. Charging a company to fix something untrue we published about it is not defensible in any of the countries we operate in. Only the promotional right of reply sits behind a claimed profile.
Reviews
Reviews never affect a firm’s position and never gate a badge. They are a paid-tier feature, so either would make both purchasable by proxy. We do not reprint competitors’ review scores — those are collected for their platforms, not ours. Where a firm has a profile elsewhere we say so and link it, and leave the number where it was earned.