datagrit

datagrit › Guides › Google AI Overview Citation & Brand Tracker

Guide

How to track which sources Google AI Overview cites for your queries

Check a list of Google queries for AI Overview citations, brand mentions and organic rank, per country, in Python or cURL and without a browser.

Published 2026-10-08 · uses the Google AI Overview Citation & Brand Tracker Actor

For many searches Google now answers above the results with an AI Overview: a short generated answer with a handful of cited pages. If your page is one of the cited sources, you get a visible link at the top of the page. If it is not, a page that ranks second in the organic results can still be invisible for the reader who stops at the answer. Rank trackers were built for the ten blue links and say nothing about this block, so teams that care about visibility in AI answers (often called AEO or GEO) end up checking queries by hand.

This guide shows how to check a list of queries automatically with the Google AI Overview Citation & Brand Tracker Actor, what each row contains, and where the method has limits.

What you need to measure

Three questions cover most reporting needs:

The first question is more slippery than it looks, because Google decides per request whether to show an AI Overview. The Actor reports what the page it received contained: a page with an AI Overview is present, with its text and citations; a page without one is absent, with the organic results still in the row. A query can flip between checks, so the useful setup is a scheduled run over the same list, compared over time.

Run it from Python

The Actor takes a list of queries, the market (gl for country, hl for language), and optionally the domains and brand names you track.


from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("datagrit/google-ai-overview-citation-tracker").call(run_input={
    "queries": [
        "what is kubernetes",
        "how to lower blood pressure",
        "best project management software",
    ],
    "gl": "us",
    "hl": "en",
    "trackedDomains": ["kubernetes.io", "mayoclinic.org"],
    "brandNames": ["Kubernetes", "Mayo Clinic"],
    "maxItems": 50,
})

rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())

The same call with cURL:


curl -X POST "https://api.apify.com/v2/acts/datagrit~google-ai-overview-citation-tracker/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"queries": ["what is kubernetes"], "gl": "us", "hl": "en", "trackedDomains": ["kubernetes.io"]}'

A run reads several queries in parallel, and a list of ten queries usually finishes in about two minutes, because Google answers some requests slowly.

The fields that matter

FieldMeaningExample
aiOverviewPresentTrue when the page had an AI Overview, false when it did nottrue
aiOverviewTextThe answer as plain text, one line per paragraph or list itemKubernetes (often called K8s) is ...
citationCountDistinct source pages cited6
citedDomainsDomains in order of first citation["kubernetes.io", "redhat.com"]
trackedDomainCitedWhether any tracked domain is citedtrue
bestCitationPositionPosition of the first citation from a tracked domain1
brandMentionCountMentions of your brand names in the answer text2
trackedDomainOrganicRankFirst organic position of a tracked domain1

Tracking fields are null when you pass no domains or brands, so an empty list is never confused with "not cited". Rows with found: false are status rows and are not charged.

Three recipes

A weekly visibility table

Schedule the Actor weekly with the same query list and load the rows into a sheet or a database table keyed by query and date. The interesting columns are trackedDomainCited and bestCitationPosition. Plot the share of queries where you are cited, per week, per market. Because Google varies AI Overviews, one check is a sample; the trend over several weeks is the signal.

Cited but not ranked, ranked but not cited

Filtering rows with pandas separates two groups that need different work:


import pandas as pd

df = pd.DataFrame(rows)
df = df[df["found"] == True]

cited_not_ranked = df[(df["trackedDomainCited"] == True) & (df["trackedDomainOrganicRank"].isna())]
ranked_not_cited = df[(df["trackedDomainCited"] == False) & (df["trackedDomainOrganicRank"] <= 3)]

Pages in the second group already rank well and are not used as a source. They are the first candidates for a clearer definition, a table, or a section that answers the question in two sentences.

Brand mentions without a link

brandMentionCount counts names inside the answer text, independent of the citations. A competitor can be named in the answer while your page is cited, or the other way around. Tracking both tells you whether you are only a source or also part of the answer.

Limits worth knowing

Getting started

Open the Google AI Overview Citation & Brand Tracker on the Apify Store, paste ten of your own queries and add your domain. Pricing is pay per result, and a status row for a query without results or with an unreadable page costs nothing. Documentation and examples are also on the project page.

Try Google AI Overview Citation & Brand Tracker on ApifyActor documentation

More guides