← Blog

    How to Audit What AI Assistants Say About Your Tours

    There is no Search Console for ChatGPT. Here is a repeatable, no-cost method for measuring how AI assistants describe and recommend your experiences — and how to diagnose what is holding you back.

    The short version

    You cannot fix AI visibility you have not measured, and no assistant gives you a Search Console. There is no impressions report for ChatGPT, no rank tracker for Gemini, no dashboard telling you that a traveller asked for the best cooking class in Bologna and got sent to your competitor.

    So you audit it manually. It takes an afternoon, costs nothing, and produces a baseline you can actually act on. This is the method: build a prompt set, run it properly, score what comes back, then work out why.

    If you have not read why AI visibility is becoming the distribution channel for tour and activity suppliers, start there. This post is the practical follow-up.

    What are you actually measuring?

    Three different things, and conflating them is the most common mistake.

    Mentions — the assistant names your business in its answer. This comes from the model's training data or its live search, and it is the strongest signal that the model knows you exist.

    Citations — the assistant links to your website as a source. You can be mentioned without being cited, and cited without being mentioned. Citations are what drive referral traffic.

    Position — where you sit when the answer is a list. Being third in a list of eight is not the same as being first, and travellers behave the same way they do on a search results page.

    Track all three separately. A supplier who is mentioned often but never cited has a content problem. One who is cited but never mentioned has an entity recognition problem. The fixes are different.

    Which assistants are worth testing?

    Not all of them, and not equally. Weight your effort by where the traffic actually comes from.

    AssistantShare of AI referral trafficWhy it matters
    ChatGPT78–87% depending on the dataset (Statcounter via Passionfruit, Conductor via AirOps)900m weekly active users (Search Engine Land). If you test one thing, test this
    Google Gemini8.65%, and rising (Statcounter via Passionfruit)Feeds AI Overviews, so it overlaps with your existing SEO
    Perplexity7.07% (Statcounter via Passionfruit)Heavily citation-driven — the easiest place to see which of your pages AI can actually read
    Claude2.91% (Statcounter via Passionfruit)Small volume, but the highest reported sign-up conversion rate of the group

    Volume is not the whole story. AI referral traffic is still only around 1.08% of total sessions across 13,770 domains (Conductor, via AirOps) — but those visitors convert at roughly 4.4x the rate of traditional organic search (Semrush, via Contently). Small channel, disproportionately valuable traffic. That is precisely the profile worth auditing early.

    Step 1: Build a prompt set that matches real traveller questions

    Twenty prompts is enough for a first audit. Do not write them as keywords — write them the way a traveller types into a chat window, in full sentences, with context.

    Cover four intent types:

    • Category discovery — "What are the best food tours in Seville?"
    • Trip planning — "I have three days in Seville with two teenagers, what should we book?"
    • Comparison — "Is a walking tapas tour or a cooking class better in Seville?"
    • Branded — "Tell me about [your business name]" and "Is [your business name] any good?"

    The branded prompts matter more than they look. If an assistant gets your cancellation policy, duration or meeting point wrong, that is a live misinformation problem affecting travellers right now, and it is usually fixable in a day.

    Write the prompts down in a spreadsheet before you run any of them. An audit you improvise is an audit you cannot repeat next quarter.

    Step 2: Run them properly

    This is where most DIY audits go wrong, and the errors all push in the same direction — they make your visibility look better than it is.

    Run each prompt at least three times. A Washington State University study running ten identical prompts against 700+ hypotheses found ChatGPT returned consistent answers only 73% of the time (via Merciv). Separate research on reproducibility found substantively different outputs even at temperature zero (arXiv, February 2026). One run is an anecdote. Three runs is a measurement.

    Use a logged-out or incognito session. Chat history and memory personalise answers. If you have spent six months asking ChatGPT about your own business, it will helpfully surface your own business, and you will conclude everything is fine.

    Set the geography deliberately. Answers differ by location. Test from your source markets, not from where you happen to be sitting. A VPN is enough for a first pass.

    Record the full answer, not just whether you appeared. Copy the whole response into your sheet with a timestamp. You need the wording later to diagnose why competitors are being chosen, and you need the timestamp because these systems change under you without notice.

    Step 3: Score what comes back

    For each prompt, log these columns:

    ColumnWhat to record
    MentionedYes/no
    CitedYes/no, plus which URL
    PositionWhere you appear in the list, or n/a
    Competitors namedEvery business the assistant recommended instead
    Sources citedWhich domains the answer leaned on — OTA listings, TripAdvisor, blogs, your own site
    AccuracyAny wrong price, duration, meeting point or policy
    SentimentHow the answer describes you

    Then calculate one headline number: inclusion rate — the percentage of your twenty prompts where you appeared at all. That is your baseline. Everything you do afterwards is measured against it.

    Most suppliers running this for the first time find an inclusion rate in the low single digits on non-branded prompts. That is normal, and it is the point of doing the audit.

    Step 4: Read the citations, not just the mentions

    The "sources cited" column is the most useful thing in the sheet, and almost everyone skips it.

    Look at which domains the assistant actually leaned on. If every answer about your city cites the same three OTA listing pages and a 2019 blog post, that tells you exactly where the model is getting its picture of your market. If your own site never appears, the model either cannot read it or does not trust it.

    Then check the competitors who were named. Look at what their pages have that yours do not: structured data, clear pricing, live availability, review volume, consistent details across sources. You are reverse-engineering the model's selection criteria from its output, which is the closest thing to a ranking factors document you are going to get.

    Step 5: Diagnose the cause

    Match the pattern to the fix.

    1. Never mentioned, never cited. The model does not know you exist as an entity. Priority is structured data, schema markup and getting consistent references to your business across the web.
    2. Mentioned but never cited. The model knows you but does not trust or cannot parse your site. Priority is on-page structure — clear, machine-readable product data rather than prose.
    3. Cited but with wrong details. The model is reading stale or conflicting information. Priority is auditing your details across your own site and every OTA listing until they match exactly.
    4. Appears on branded prompts only. You are visible to people who already know you, which is not distribution. Priority is category and trip-planning content that answers the non-branded question.
    5. Appears but ranked below competitors. You are in the consideration set and losing on signals. Priority is availability, response speed and review recency.

    How often should you repeat it?

    Quarterly for the full twenty-prompt set. Monthly for a five-prompt subset covering your highest-value categories.

    Two reasons for the cadence. Models are updated without announcement, so a result from January does not describe today's behaviour. And you need a before-and-after to know whether any of your fixes worked — otherwise you are optimising blind, which is the situation you started in.

    Keep every run in the same sheet with dates. The trend line is worth more than any single audit.

    Where Tixxly fits

    Doing this manually is the right first move. It costs nothing, and the diagnostic value of reading twenty real AI answers about your own market is genuinely high.

    It also does not scale. Twenty prompts across four assistants at three runs each is 240 queries per audit, and it is a snapshot that starts ageing immediately.

    Tixxly is AI infrastructure for travel commerce. We connect tour and activity suppliers to AI ecosystems — ChatGPT, Gemini, Perplexity, Claude and others — so their products can be found, understood and booked when travellers ask an assistant for recommendations. That means structured, machine-readable product data, live availability, and a Trust Score built from data quality, response speed, booking success rate and reputation sentiment.

    We are not an OTA and we charge no commission, so suppliers keep 100% of booking revenue less card processing fees. Connect once, reach everywhere.

    Run the audit first. Whatever it tells you, you will make better decisions with a baseline than without one.