datagrit

datagrit › Data › Google News Monitor - Full Text & Real URLs

Data

Google News Monitor - Full Text & Real URLs

Monitor Google News for keywords, site: queries and topics in any country; decoded publisher URLs, optional full text, only-new mode.

Run it on Apify StoreUse the APIfrom $3.50 per 1,000 results + $10 per run · no code needed
from $3.50 per 1,000 results + $10 per runpay only for articles you get
JSON · CSV · Excelexport or call via API
Scheduled runsdaily or weekly feeds with Apify schedules
v0.2updated 2026-09-30

Google News Monitor watches Google News for your keywords, site: searches and topic sections in any country and language, and returns every article as a clean row: headline, source, publication time, the real publisher URL decoded from the Google News link and, if you want it, the full article text, author and image. Switch on Only new articles and schedule it, and each run returns just the articles you have not received yet. It is built for brand and competitor monitoring, PR and media tracking, market news for trading or research, and news feeds for AI agents and n8n, Make or Zapier workflows.

Use cases

Sample output


{
  "query": "openai",
  "matchedQueries": ["openai"],
  "title": "FTC launches broad investigation into Anthropic, OpenAI",
  "sourceName": "washingtonpost.com",
  "sourceDomain": "washingtonpost.com",
  "sourceHomepage": "https://www.washingtonpost.com",
  "publishedAt": "2026-09-30T17:44:14.000Z",
  "hoursSincePublished": 0.6,
  "originalUrl": "https://www.washingtonpost.com/technology/2026/09/30/ftc-launches-broad-investigation-into-anthropic-openai/",
  "urlStatus": "resolved",
  "googleNewsUrl": "https://news.google.com/rss/articles/CBMirAFBVV95cUxPNERSOFU3...?oc=5",
  "articleId": "CBMirAFBVV95cUxPNERSOFU3...",
  "snippet": "The probe reflects a focus by Trump administration officials on using existing laws to police AI.",
  "language": "en",
  "country": "US",
  "matchedEditions": ["US:en", "GB:en"],
  "fullText": "The Federal Trade Commission has opened a broad investigation into the safety of artificial intelligence systems...",
  "fullTextStatus": "partial",
  "author": "Ian Duncan",
  "imageUrl": "https://www.washingtonpost.com/wp-apps/imrs.php?src=...",
  "found": true,
  "scrapedAt": "2026-09-30T18:20:00.000Z"
}

Every field is described in the dataset schema of the Actor. The Articles view shows the main columns and the Full text view shows author, snippet, text and image.

How does URL decoding work?

Google News links point to news.google.com and hide the publisher URL in an encoded article ID. The Actor decodes each returned article through the same public endpoint the Google News website uses: one request for the article page and one decoding request per 10 articles, about 1.1 requests per article. The result goes into originalUrl. The run status reports how many links were decoded and how many decoded URLs are on the source's own domain ("Original URL resolved for 20 of 20 articles, on the source's own domain for 20 of 20"). If Google stops decoding links, or if most decoded URLs point away from the article's source, the run fails with a clear message before any article is returned or charged, instead of returning rows with missing or wrong URLs.

Output fields

Every result is one flat record, so it drops straight into a spreadsheet, a database or a CRM.

FieldTypeDescriptionExample
querystringThe first query or topic (topic:NAME) in your input that found the article. On the single status row (found = false) it lists all queries and topics that were read, comma-separated.openai
matchedQueriesarrayEvery query and topic that found this article, in input order. Duplicates across queries and editions are merged into one row, so this list shows all of them. Empty on the status row.["openai","topic:TECHNOLOGY"]
titlestringHeadline as Google News shows it, without the " - Source" suffix.FTC launches broad investigation into Anthropic, OpenAI
sourceNamestringPublisher name from the Google News feed.The Washington Post
sourceDomainstringDomain of the publisher homepage from the feed, without www.washingtonpost.com
sourceHomepagestringPublisher homepage URL from the feed.https://www.washingtonpost.com
publishedAtstringPublication time from the feed, ISO 8601 in UTC.2026-09-30T17:44:14.000Z
hoursSincePublishednumberHours between publication and the start of the run, one decimal. The time range filter is checked against this value.0.6
originalUrlstringPublisher URL of the article decoded from the Google News link. Null when decoding is off or failed (see urlStatus).https://www.washingtonpost.com/technology/2026/09/30/ftc-launches-broad-investig
urlStatusstringresolved = originalUrl decoded; failed = Google did not decode this link; notRequested = Decode original article URLs was off. Null only on the status row.resolved
googleNewsUrlstringArticle link as published in the Google News feed (news.google.com). It redirects to the publisher in a browser.https://news.google.com/rss/articles/CBMirAFBVV95cUxPNERSOFU3M09XYWNrVTdSdHhkcVl
articleIdstringGoogle News article identifier (the guid of the feed item). The same article has the same ID in every edition and query.CBMirAFBVV95cUxPNERSOFU3M09XYWNrVTdSdHhkcVl4NkxXeUU2dHItLW41QnJ5TjhOX1VhaUh0Wmc1
snippetstringArticle summary from the publisher page (structured data description, og:description or meta description). Filled only with Full text on: the Google News feed itself carries no summary, only the headline.The probe reflects a focus by Trump administration officials on using existing l
languagestringLanguage of the edition in which the article was found first, for example en, de or es-419.en
countrystringCountry code of the edition in which the article was found first.US
matchedEditionsarrayEvery edition (COUNTRY:language) whose feed contained this article.["US:en","GB:en"]
fullTextstringArticle text extracted from the publisher page (up to 100,000 characters), paragraphs separated by blank lines. Null unless Full text is on and text was found.The Federal Trade Commission has opened a broad investigation into the safety of
fullTextStatusstringok = 500 or more characters extracted; partial = shorter text (paywall teaser, video page or short item); empty = page without readable paragraphs; blocked = publisher answered 401, 402, 403, 429 or 451; failed = other HTTP error or not an HTML page; noUrl = original URL could not be decoded; notRequested = Full text was off. Null only on the status row.ok
authorstringAuthor names from the publisher page, separated by semicolons. Filled only with Full text on.Ian Duncan
imageUrlstringMain image of the article from the publisher page. Filled only with Full text on.https://www.washingtonpost.com/wp-apps/imrs.php?src=https://arc-anglerfish-washp
foundbooleanTrue for every article. False only on the single status row written when a run returns no article; its query field and the run status say why.true
scrapedAtstringISO 8601 time when the run started reading the feeds.2026-09-30T18:20:00.000Z

Sample record

{
  "query": "openai",
  "matchedQueries": [
    "openai",
    "topic:TECHNOLOGY"
  ],
  "title": "FTC launches broad investigation into Anthropic, OpenAI",
  "sourceName": "The Washington Post",
  "sourceDomain": "washingtonpost.com",
  "sourceHomepage": "https://www.washingtonpost.com",
  "publishedAt": "2026-09-30T17:44:14.000Z",
  "hoursSincePublished": 0.6,
  "originalUrl": "https://www.washingtonpost.com/technology/2026/09/30/ftc-launches-broad-investigation-into-anthropic-openai/",
  "urlStatus": "resolved",
  "googleNewsUrl": "https://news.google.com/rss/articles/CBMirAFBVV95cUxPNERSOFU3M09XYWNrVTdSdHhkcVl4NkxXeUU2dHItLW41QnJ5TjhOX1VhaUh0Wmc1Y3ZFVm9qeXhfd0gxRVNNclo3UURDbGFrb19KS2RhSWlrbFl1RTZ5V2dlbEN4WXZ0SkwyTy1mRFc4M0lDYVRCZ0lBQnJzZ2xmenA1cDNaX294MzJYVmhydXlBQ1ZQbG40LW9lY1UtQ2MzRmZsdDJ2MWplNVAx?oc=5",
  "articleId": "CBMirAFBVV95cUxPNERSOFU3M09XYWNrVTdSdHhkcVl4NkxXeUU2dHItLW41QnJ5TjhOX1VhaUh0Wmc1Y3ZFVm9qeXhfd0gxRVNNclo3UURDbGFrb19KS2RhSWlrbFl1RTZ5V2dlbEN4WXZ0SkwyTy1mRFc4M0lDYVRCZ0lBQnJzZ2xmenA1cDNaX294MzJYVmhydXlBQ1ZQbG40LW9lY1UtQ2MzRmZsdDJ2MWplNVAx",
  "snippet": "The probe reflects a focus by Trump administration officials on using existing laws to police AI.",
  "language": "en",
  "country": "US",
  "matchedEditions": [
    "US:en",
    "GB:en"
  ],
  "fullText": "The Federal Trade Commission has opened a broad investigation into the safety of artificial intelligence systems made by Anthropic and OpenAI...",
  "fullTextStatus": "ok",
  "author": "Ian Duncan",
  "imageUrl": "https://www.washingtonpost.com/wp-apps/imrs.php?src=https://arc-anglerfish-washpost-prod-washpost.s3.amazonaws.com/public/example.jpg",
  "found": true,
  "scrapedAt": "2026-09-30T18:20:00.000Z"
}

Input

FieldNameTypeWhat it does
queriesSearch queriesarrayGoogle News searches, one per line. Google News operators work: "exact phrase", OR, -exclude, site:reuters.com, intitle:word, when:1d, after:2026-09-01 and before:2026-09-30. Each query is read in every edition below. An article found by several queries or editions is returned once, with all of them in matchedQueries and matchedEditions. Leave empty when you only monitor topics.
topicsTopic sectionsarrayOptional Google News sections to monitor in every edition: TOP (top stories), WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE or HEALTH. Rows from a section have the query topic:NAME, for example topic:TECHNOLOGY. Unknown names are skipped and listed in the run status.
editionsEditions (country:language)arrayGoogle News editions to read, as COUNTRY:language codes: US:en, GB:en, IN:en, AU:en, CA:en, CA:fr, DE:de, FR:fr, ES:es, IT:it, NL:nl, PL:pl, JP:ja, BR:pt-419, MX:es-419 and other editions Google News offers. en-US style codes are accepted and UK is read as GB. Each edition is one feed per query, so 3 queries x 2 editions read 6 feeds. Google serves another edition for a pair it does not offer (for example PL:en returns US:en); such an edition is skipped, returns no articles, is not charged and is named in the run status. If none of the editions exists, the run fails.
timeRangeTime rangestringKeep only articles published within this period before the run. The Actor adds the matching when: operator to every query that has no when:, after: or before: of its own, and then checks the publication date of every article again against the time range and the query's own when:, after: and before: (the last two with one day of tolerance), because Google News sometimes returns older articles (on 30 September 2026, 8 of 100 results of site:bbc.co.uk when:1d were from 2011 to 2025). Topic sections are filtered by the same date check.
maxItemsMaximum articlesintegerMost articles to return in one run, newest first, across all queries, topics and editions. One Google News feed holds about 100 articles at most, so this is the practical maximum per query and edition.
onlyNewOnly articles new since the last runbooleanReturn only articles that an earlier run with the same queries, topics, editions and time range has not returned yet. The memory is kept per combination of those settings in your account, so two monitors with different queries never hide each other's articles, and only articles that were actually returned are remembered. Use it with a schedule to get a stream of fresh news.
resolveUrlsDecode original article URLsbooleanGoogle News links point to news.google.com, not to the publisher. When on, the Actor decodes every returned article to its original publisher URL (originalUrl) and checks site: queries against it too. This takes one request per article plus one per 10 articles to Google News (about 1.1 per article), so a run of 100 articles takes one to two minutes longer. Always on when Full text is on.
fullTextFull text, author and imagebooleanOpen each article on the publisher's site and extract the article text, author, main image and summary (snippet). Plain HTTP, no browser: paywalled and script-rendered pages return partial or no text, which fullTextStatus reports per article. Adds roughly one second per article.
proxyConfigurationProxy configurationobjectOptional proxy for requests to Google News and publishers. Not needed in normal use.

Call it from your code

Run the Actor and get the results in one request. Replace YOUR_APIFY_TOKEN with the token from your Apify account settings.

curl -X POST "https://api.apify.com/v2/acts/datagrit~google-news-rss-monitor/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"queries":["openai","\"electric vehicles\""],"editions":["US:en"],"timeRange":"1d","maxItems":20}'
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagrit/google-news-rss-monitor').call({
  "queries": [
    "openai",
    "\"electric vehicles\""
  ],
  "editions": [
    "US:en"
  ],
  "timeRange": "1d",
  "maxItems": 20
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.length, items[0]);

Install with npm i apify-client.

from apify_client import ApifyClient
import os

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("datagrit/google-news-rss-monitor").call(run_input={
  "queries": [
    "openai",
    "\"electric vehicles\""
  ],
  "editions": [
    "US:en"
  ],
  "timeRange": "1d",
  "maxItems": 20
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(len(items), items[0] if items else None)

Install with pip install apify-client.

Try Google News Monitor - Full Text & Real URLs on Apify

Related Actors

Company data

French Company Finder - Sirene Financials

French company lead lists from Sirene screened by net result and revenue, with net margin, size, matching establishment and optional directors.

from $5.60 / 1,000 results
Company data

Poland KRS New Company Registrations Feed

Newly registered Polish companies, foundations and associations from the official KRS court register: NIP, address, PKD, capital, email, with filters and change detection.

from $10.50 / 1,000 results