Vault Mark
Business leader reviewing an AI citation capability scorecard covering evidence, measurement and data ownership

Agency Evaluation · AI Search · Evidence & Measurement

Before accepting a screenshot or a promise that an agency can make your brand “appear in ChatGPT,” verify whether it has a real system for selecting buyer questions, producing evidence, checking website accessibility, maintaining prompt logs, and connecting citations to referral traffic and qualified enquiries.

Terms such as AEO, GEO, LLMO and AI Search now appear in more agency proposals. The problem is not that these concepts are invalid. The problem is that many businesses cannot yet distinguish a verifiable operating system from existing SEO and content services that have simply been given new labels.

A screenshot showing that ChatGPT mentioned a brand once may be interesting, but it does not explain why the brand appeared, which page supported the answer, whether the result can be repeated, or whether the prompt reflects a real buying decision.

Businesses that need the fundamentals before evaluating a provider can first review six conditions that can improve a brand’s readiness to be cited by ChatGPT and the broader AI Search Optimization system.

Direct answer

An agency with genuine AI citation capability should be able to show more than screenshots. It should identify the buyer questions it intends to win, assign a primary page to each answer, verify crawl and index eligibility, maintain timestamped prompt logs with source URLs, separate mentions, citations, recommendations, referrals and leads, and explain both the limitations and the decision rules it will use when results do not develop as expected.

What AI Citation Means—and Which Signals Must Stay Separate

Before evaluating an agency, agree on what each result means. An AI system mentioning a brand, citing a page, recommending a provider and producing a qualified lead are different events and should not be combined into one visibility number.

Each signal answers a different business question
Signal What it means What it does not prove
Brand Mention An AI system names the brand It does not prove which page or source informed the answer
Direct Citation The system links to or identifies a page as a source It does not prove that the user clicked or intended to buy
Recommendation The system presents the brand as an option It does not prove that the user selected that option
Referral Visit A user clicks from an AI experience to the website It does not prove that the visitor is a qualified lead
Qualified Enquiry The contact fits the business and has a need that can progress to sales It is not yet confirmed revenue

OpenAI states that public websites can appear in ChatGPT Search. Allowing OAI-SearchBot to access a site helps make content eligible for discovery, summaries and linked results. Referral URLs from ChatGPT Search also include utm_source=chatgpt.com, which can support traffic measurement (OpenAI Publishers and Developers FAQ).

Google states that AI Overviews and AI Mode continue to rely on established SEO foundations. A page needs to be indexed and eligible to appear with a snippet before it can be considered as a supporting link. Google does not require a special AI file or special schema that guarantees inclusion (Google Search Central: AI features and your website).

Gate 1

Does the Agency Start With Buyer Questions or With Services It Wants to Sell?

1. “Which buyer questions are we trying to earn the right to answer?”

“Every question related to your business” is not a credible answer. The agency should define specific question arenas: comparison questions, pre-selection questions, pricing and risk questions, questions about value, and objections that the sales team repeatedly has to resolve.

Evidence to request A buyer-question universe, informational/commercial/transactional intent classification, branded and non-branded demand separation, and a clear explanation of why each question set matters to the business.
Red flag The agency starts with article volume, keyword volume or a service list before it can explain what decision the buyer is trying to make.

2. “Which page is the primary URL owner for each question?”

Each important question should have one clearly assigned primary answer page. If multiple pages address the same question with the same intent, the site can create duplication, cannibalisation and weak ownership rather than one authoritative source.

A strong system should include a query-to-page map, parent pillars, supporting articles and an internal-link direction similar to the cluster-first approach described in the AI Search Content Factory.

Evidence to request A query-to-page map, primary URLs, parent pillars, supporting articles, internal-link direction, canonicals and Thai/English page pairs.
Red flag “We will publish several articles and see which one ranks” without query ownership or a cannibalisation correction rule.

3. “How do you decide whether the answer should be an article, service page, tool, checklist or original dataset?”

AI citation work should not be reduced to producing a large number of blog posts. Some questions need a concise answer; others need a comparison matrix, calculator, checklist, decision tree or dataset with a transparent methodology.

Evidence to request Content-type decision rules, the reason for choosing an answer block, guide, checklist, dataset or decision tree, and a clear statement of the information gain the new asset will add.
Red flag Every question produces the same recommendation: “a 1,500–2,000 word SEO article,” because that is the easiest output for the agency to manufacture.
Gate 2

Does the Agency Understand Technical Eligibility?

4. “How will you verify that search engines and AI systems can access the page?”

Before discussing citations, the page must be accessible, processable and indexable. The review should cover robots.txt, meta robots, X-Robots-Tag, canonicals, sitemaps, HTTP status, JavaScript rendering, CDN or WAF controls, internal links, Googlebot, Bingbot and OAI-SearchBot.

For a broader view of technical SEO, entities and schema, see the AI On-page, Technical & Experience SEO OS.

Evidence to request A technical crawl report, robots.txt review, rendered HTML, a list of URLs with noindex or incorrect canonicals, and a named owner for each required correction.
Red flag The agency discusses prompts and content but does not inspect crawling, indexing, rendering or security layers that may block bots.

5. “How will you determine whether the problem is technical, content-related, entity-related or caused by weak external evidence?”

A brand may fail to appear because a page cannot be accessed, no page answers the question directly, the content adds no new information, the brand entity is inconsistent, supporting external sources are weak, or the tested questions do not reflect meaningful demand.

Evidence to request A diagnostic matrix that separates problem layers, the evidence used to identify the bottleneck, and the dependencies that must be fixed before producing more content.
Red flag Every symptom leads to the same recommendation: more content and more backlinks, without a root-cause diagnosis.

6. “Will the schema you add match what users can actually see on the page?”

Structured data should describe visible content, not create the appearance that a page has expertise, features or information that it does not show. Google requires structured data to match visible content and does not offer special schema that guarantees appearance in AI Overviews or AI Mode.

Evidence to request A schema map, visible-content-to-schema comparison, validator results, matching author/reviewer/date information, and a method for preventing duplicate schema from multiple plugins.
Red flag The agency claims that adding FAQ, HowTo or an invented “AI schema” will make ChatGPT or Google select the site directly.
Gate 3

Is the Content Evidence-Led—or Merely an AI-Generated Summary?

7. “Which source supports each claim, and who is accountable for accuracy?”

Readable content is not automatically citeable content. The agency should separate sourced facts, evidence-based conclusions, professional judgement, assumptions, publishable internal information and information restricted by confidentiality.

Evidence to request A source log, claim register, named author, specialist reviewer, fact-check date, correction policy and version history.
Red flag AI drafts the article and a writer only checks the language, with no claim or source verification.

8. “What information gain will this page provide that other sources do not?”

Repackaging existing answers may create a comprehensive-looking article, but it does not make the page necessary as a source. Information gain can come from a framework, comparison matrix, checklist, original survey, anonymised audit, calculator or market-specific evidence.

Evidence to request A one-sentence unique contribution, the supporting method, a comparison with existing pages and competitors, and the limitations of the information.
Red flag The proposed information gain is simply “make it longer,” “add more keywords” or “ask AI to expand the detail.”

9. “How will you build a clear and consistent brand entity?”

Systems should find consistent information about who the brand is, what it specialises in, who it serves, which markets it serves and which page is the source of truth. Entity building is not repeating the brand name in every paragraph. It is aligning the homepage, about page, service pages, author profiles, schema and external profiles.

Evidence to request An entity glossary, source-of-truth page map, Organization and Person data review, external-profile consistency audit and a correction plan for conflicting information.
Red flag The only recommendation is to mention the brand name more often without auditing entity relationships or source consistency.
Gate 4

Does the Agency Separate Signals—or Treat Screenshots as Results?

10. “How do you separate mentions, citations, recommendations, referrals and leads?”

These metrics belong in different fields because they answer different business questions. Bing Webmaster Tools provides AI Performance reporting for citation activity, cited pages and grounding queries, but Bing states that citation data is not a ranking, authority score, click metric or proof that one specific page change caused the result (Bing Webmaster Tools: AI Performance).

Each stage in the measurement chain should be reported separately
Stage Question answered Example evidence
Mention Did the system name the brand? Timestamped output
Citation Which page did the system cite? Source URL or Bing AI Performance
Recommendation Was the brand presented as an option? Prompt result with context
Referral Did the user click through? Analytics referral or UTM
Qualified Lead Was the contact commercially relevant? CRM lead status or sales feedback
Revenue Was revenue confirmed? Closed-won record or finance confirmation

The method for connecting search signals to leads and sales should align with a wider measurement system such as AI Data & Measurement, rather than collapsing everything into one “AI visibility” figure.

11. “What does your prompt log record?”

A verifiable prompt log should record at least:

  • The exact prompt
  • The platform and language
  • The date and time
  • The relevant country or context
  • The account state when it may affect the response
  • Brand mentions, direct citations and recommendations
  • Source URLs and surfaced competitors
  • A screenshot or saved output
  • The person who performed the test
  • Notes about volatility
Evidence to request Ask to see a real, anonymised prompt log—not an empty template.
Red flag Screenshots are provided without the exact prompt, date, platform, source URL or negative results.

12. “How do you test repeatability and volatility?”

AI-generated answers can change with prompt wording, language, time, context, platform, model, website updates and the source set available to the system. Google states that AI Overviews and AI Mode can use different models and techniques, while Bing notes that citation volume can change with demand, content and system updates.

Evidence to request Prompt variants, a repeated-test schedule, a baseline period, definitions for stable/emerging/inconclusive results, and a change log for the relevant pages.
Red flag Only the most favourable outputs appear in the report, while tests in which the brand did not appear are omitted.

13. “How will you connect a citation to a page, referral, lead and revenue?”

The measurement chain should look like this:

Measurement chain

Prompt → AI answer → Cited page → Referral → Behaviour → Enquiry → Qualified lead → Revenue

The agency does not need to pretend that every stage can be attributed perfectly. It should identify what is observed data, modelled data, sales-confirmed data, assumption and unknown.

Red flag Revenue is attributed to AI citations without CRM source data, lead history, sales confirmation or an attribution caveat.

14. “What decisions will you make if the expected signals do not appear?”

A credible plan requires decision rules, not only a content calendar. A page that is not indexed needs technical correction before more content is produced. A mention without citation requires a source-strength review. Traffic without suitable leads requires a review of intent, offer and commercial routing.

Hypotheses, measurement and decision rules should be built into the operating method, not invented after a campaign fails. This is consistent with the approach used in AI Growth Experimentation.

Evidence to request Stop/continue/revise/scale rules, review cadence, decision ownership and criteria for refreshing, merging or retiring pages.
Red flag Every monthly answer is “we need more time,” with no condition that would trigger a change in direction.
Gate 5

Will the Business Still Own the System?

15. “Who owns the content, prompt logs, analytics, accounts and source files?”

When the engagement ends, the business should retain access to the website, domain, Google Search Console, Bing Webmaster Tools, GA4, GTM, CRM data, prompt baseline, citation logs, source files, research notes, content briefs, schema documentation, dashboards and measurement definitions needed to continue operating.

The agreement should distinguish between owner, administrator, manager access, licensed assets, agency-owned tools and confidential methodology.

Evidence to request An ownership matrix, access list, exit and handover process, delivery file formats, retention and deletion rules, and clearly defined intellectual-property boundaries.
Red flag All data remains inside agency-owned accounts and the business receives only PDF reports without access to raw data or core systems.

Vault Mark AI Citation Capability Scorecard

Score the agency’s answer to each question using the same standard:

0 points No answer, a promise without evidence, or avoidance of the question
1 point A process is described, but there is no example, owner or acceptance method
2 points The process, evidence, owner, limitation and acceptance criteria are clear
Maximum score: 30 points
Score What it indicates Appropriate next decision
26–30 The operating system is strong enough for deeper due diligence Verify the delivery team, scope, evidence, ownership and commercial fit
20–25 Some capability exists, but evidence gaps remain Document the gaps and acceptance criteria before signing
12–19 The agency understands the terminology, but the operating system is incomplete Request a pilot or proof pack before a large engagement
0–11 There is a high risk of buying screenshots or content volume rather than a system Do not approve the scope without stronger evidence
Scorecard limitation

A high score does not mean the agency should automatically be hired. The business must still review the delivery team, scope, budget, industry understanding, internal collaboration requirements and real work evidence. Vault Mark’s work can be reviewed through the Showcase, but the method, baseline and limitations of each case should still be assessed independently.

AI Citation Proof Stack: Which Evidence Deserves More Weight?

  1. Screenshot
    Shows that a result occurred at least once, but does not verify the prompt, source, date or repeatability
  2. Timestamped Prompt Log
    Records the exact prompt, platform, date, output and source URL
  3. Repeated Prompt Testing
    Uses several prompt variations over time and reports both positive and negative results
  4. Query-to-Page Evidence
    Connects the question, answer-owning page, page changes and detected citations
  5. Business Measurement
    Separates mentions, citations, referrals, leads and revenue while disclosing attribution limitations

An agency does not need level-five evidence for every programme from day one. It should, however, explain how the system will progress from low-confidence evidence to more verifiable evidence over time.

How to Test an Agency’s Evidence in 30 Minutes

Step 1: Select three to five real buyer questions

Use questions customers ask before contacting or buying. Do not use prompts designed to force the brand name into the answer. Examples include:

  • How should a company like ours choose a marketing agency?
  • How can we verify that an agency measures lead quality?
  • Should a B2B company begin with SEO, ads or content?
  • Why is website traffic increasing while leads are not?
  • What information should we prepare before hiring a digital marketing agency?

Step 2: Request the relevant prompt log

Check whether it contains the exact prompt, date, platform, source URL and negative results.

Step 3: Open the cited source page

Confirm that the page answers the question rather than merely mentioning the keyword. Review the author, sources, dates, internal links, visible evidence and commercial route.

Step 4: Ask the agency to explain causation and limitations

A credible answer should acknowledge that one isolated page change cannot usually be proven as the sole cause of a citation.

Step 5: Ask for the decision rule

Ask what the agency will inspect, change, stop or measure next if the desired signals do not appear within the agreed 30–90 day review window.

Example: Two Agencies Propose Very Different AI Citation Programmes

Assume a B2B company sells a high-value system with a long sales cycle.

Agency A proposes 20 articles per month and promises that the brand will start “appearing in AI.” It shows two screenshots but provides no prompt log and does not assign an answer-owning page to each question.

Agency B begins with 25 buyer questions, groups them by buying-committee role, assigns one primary URL to each cluster, checks technical eligibility, creates four articles and one checklist, and then records a prompt baseline that separates citations, referrals and qualified leads.

Agency B does not guarantee citations. It does show a system that lets the business understand:

  • Which buyer questions it is trying to win
  • Which evidence it will create
  • Which page owns each answer
  • Which signal will be measured
  • When the plan should change

The difference is not which agency speaks more confidently about GEO. The difference is which one makes its decisions auditable. This evaluation should be applied to any digital marketing agency, not only providers that describe themselves as AI agencies.

What the First 90 Days Should Deliver

Actual deliverables must reflect the business baseline and diagnosed problem. A verifiable scope should still define tangible outputs such as:

  1. A buyer-question universe
  2. Branded and non-branded prompt baselines
  3. A query ownership map
  4. A technical crawl and index review
  5. An entity and source-of-truth audit
  6. Priority content and citation-asset briefs
  7. Published or updated priority pages
  8. Source and claim logs
  9. Internal-link implementation
  10. A prompt test log
  11. Citation, referral and lead measurement definitions
  12. A decision review covering what to continue, fix, pause or test next
Do not turn deliverables into a guarantee

This list does not promise that citations will appear within 90 days. It describes work that a business can inspect and accept, rather than receiving a report that merely states that the team is “working on GEO.”

Frequently Asked Questions

Can an agency guarantee that ChatGPT will cite a brand?

An agency should not guarantee a particular prompt result or citation. It can improve technical eligibility, content quality, entity clarity, evidence and measurement, but AI and search systems select their own answers and sources. Google explicitly states that meeting its requirements does not guarantee crawling, indexing or display.

Can screenshots be used as evidence?

Screenshots can support an evidence pack, but they are not sufficient on their own. They should be accompanied by the exact prompt, platform, date, source URL, context, repeated tests and the occasions on which the brand did not appear.

Do we need an ai.txt file or special schema to appear in Google AI features?

Google states that no special machine-readable AI file or schema is required for AI Overviews or AI Mode. The core requirements remain crawlability, indexability, useful content, internal links, visible text and structured data that accurately matches the page.

Does an AI citation equal traffic or a lead?

No. A citation indicates that a page was used as a source, but it does not prove that the user clicked or submitted an enquiry. Bing states that citation metrics are not rankings, authority scores, traffic or engagement, so referral and business outcomes must be measured separately.

Should a business work on AI citations before SEO?

These should not automatically be treated as separate systems. Crawlability, indexing, content clarity, internal linking and search eligibility support both traditional search and AI-assisted discovery. The first priority should be the actual bottleneck, not whichever term is currently attracting attention.

Can Thai-language content be cited by AI systems?

Public content does not need to be written in English to be eligible for ChatGPT Search. Relevance, quality, supporting evidence and the competitive source set still matter. Businesses should test Thai and English prompts separately rather than assuming that both languages will produce the same result.

Do Not Begin by Asking Whether an Agency “Does GEO”

Begin by asking which business decisions it can help you make more confidently. Valuable AI citation work is not about forcing a brand name into a carefully selected demonstration prompt. It is about creating useful sources for the questions buyers ask before deciding, then verifying whether those sources are accessible, cited and connected to commercial outcomes.

When a business does not yet know whether the real constraint is search visibility, content, entity clarity, website implementation, measurement or the route from attention to lead, it should not immediately purchase a larger execution stack.

Vault Mark’s Customer Growth Blueprint connects customer, offer, channel and performance evidence to identify what should happen first. Only then should the business determine the role of SEO, AI Search, content, paid media, website work or measurement.

Facebook
Threads
X
LinkedIn
Reddit
Telegram