Scraping Google with Node.js is mostly a fight with Google’s delivery, and this guide is built around what each route returned when its scripts ran. A plain HTTP request gets a 200 with no results inside, a stealth-patched Puppeteer gets further or gets the block page depending on the IP it exits from, and a SERP API returns parsed JSON. All three routes are below, as complete scripts, with the run outputs shown.
What Can We Scrape from Google?
Google is many surfaces, and the two this guide scrapes are the ones people ask for most, with two more noted for orientation.
Search results
The classic SERP, that’s organic results with their titles, links, and snippets, plus the blocks around them. This is the data behind rank tracking and content research.
Google Maps
Business listings with names, ratings, review counts, and contacts, the raw material of local lead lists. It’s also the surface that behaves best under a browser. Maps behaves differently from Search under scraping, and this guide measures the difference.
Google News
Headlines and sources per query, useful for monitoring mentions of a brand or a topic. The same techniques as SERP scraping apply, with different selectors, and the dedicated News endpoint returns the headlines already parsed.
Other Google services
Shopping, Images, and Trends each have their own layout and their own API coverage. The approaches below transfer, and each surface has a dedicated endpoint in scraping APIs.
Why Should We Scrape Google?
SEO monitoring wants positions per query over time, and it wants them daily, which is where rank tracking tools come from. Market research wants to know who appears for which queries and with what messaging, both in organic results and in the ad slots above them. Lead generation wants the business data Maps holds, the names and ratings and contacts that become a call list. Price and product teams watch Shopping. In every case the volume is what forces automation, since checking fifty queries by hand stops being a plan on day two, and five hundred isn’t a plan on day one.
How to scrape Google data
Three ways need no code at all, and they split by who is doing the asking. A no-code Google scraper takes the queries in a form and returns a results file, which covers one-off collections without a project. The API playground generates the request in whatever language the project speaks, Node.js included, so it doubles as a starting point for the scripts below.
The third is for when the caller is a model rather than a person. Point any MCP client at https://mcp.hasdata.com/api/mcp with an x-api-key header, and the client picks up 63 tools, 22 of them Google:
{
"mcpServers": {
"hasdata": {
"url": "https://mcp.hasdata.com/api/mcp",
"headers": { "x-api-key": "YOUR_API_KEY" }
}
}
}Asking for hasdata_google_serp_serp_getSearchResults with {"q": "web scraping node js"} returns the same parsed JSON the API route below returns, 8 organic results with aiOverview, relatedQuestions and pagination beside them. The model gets fields rather than a page, so there is no scraper in the loop to maintain and nothing to parse on the way in. That route ends where this guide’s does not, because it cannot be dropped into an existing Node service the way a request can.
Inside a Node.js project, four routes matter, and here’s what each one delivered in this run:
| Route | What the run returned | Where it fits |
|---|---|---|
| Axios + Cheerio | HTTP 200, title “Google Search”, 0 results in the markup | Nothing to parse; Google renders results with JavaScript |
| Puppeteer (plain) | Announces itself through navigator.webdriver, so the stealth plugin is the baseline below | A base to build on, and detectable as automation until patched |
| Puppeteer + stealth plugin | Maps answered with data; Search still served /sorry from a datacenter exit | The standard baseline; the exit IP decides on Search |
| SERP API | Parsed JSON, 8 organic results with positions | The route with the parsing on the service side, priced per request |
The rest of the guide builds each row into a working script.
Preparing for Scraping
The setup comes down to two installs and two selector notes.
Installing the environment
Node.js 18 or newer from nodejs.org covers everything below, and node -v confirms the install. Create a project folder and initialize it:
npm init -yThat writes the package.json the installs below attach to.
Installing the libraries
The three routes need these packages:
npm i axios cheerio puppeteer-extra puppeteer-extra-plugin-stealth puppeteeraxios makes HTTP requests, both to Google and to the API. cheerio parses HTML with jQuery-like selectors (the pairing has its own guide, and the wider JavaScript library overview maps the alternatives), and puppeteer-extra wraps Puppeteer so plugins like the stealth patch can hook into it. The @scrapeit-cloud/* SDK packages on npm still work but carry the pre-rename brand and have had no release since 2024, so the API sections below call the endpoints directly with axios and skip that dependency.
Google SERP page analysis
Search results live in JavaScript-rendered markup with generated class names that rotate between visits. Titles sit in h3 tags inside result links, which is the one stable hook, and counting h3 elements is the quickest check of whether a fetch got real results at all. That check is what the first script below relies on.

The highlighted boxes are the whole contract, a link and its h3, and everything else on the page is generated noise.
Google Map page analysis
Maps renders its results list into a container with role="feed", with one div.Nv2PK card per place. The name is in .qBF1Pd and the rating in the aria-label of a span (“4.9 stars 306 Reviews”), the same markup our Google Maps guide works with from Python.

Aria attributes survive redesigns better than the class soup around them, which is why the rating reads from the label.
Scraping Google SERP with NodeJS
The three routes from the comparison table, in the same order.
Scrape Google SERP Using Axios and Cheerio
The plain-HTTP route is worth running once, because its failure is quiet. The script is small:
const axios = require("axios");
const cheerio = require("cheerio");
(async () => {
const response = await axios.get(
"https://www.google.com/search?q=web+scraping+node+js&hl=en",
{
headers: {
"User-Agent":
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/153.0.0.0 Safari/537.36",
},
validateStatus: () => true,
}
);
console.log("status:", response.status);
const $ = cheerio.load(response.data);
console.log("title:", $("title").text());
console.log("results:", $("h3").length);
})();It prints:
status: 200
title: Google Search
results: 0That’s the trap. The request “succeeds”, the status is 200, and there’s nothing to parse, because Google sent the JavaScript shell rather than rendered results. Every Cheerio selector returns empty from here, so the route’s real role is this check, and the comparison table already said the rest.
Scrape Google SERP using Puppeteer
Rendering is Puppeteer’s job, and the stealth plugin is the baseline rather than an add-on. Plain Puppeteer announces itself through navigator.webdriver and a dozen other properties that Google reads. puppeteer-extra patches them:
const puppeteer = require("puppeteer-extra");
const StealthPlugin = require("puppeteer-extra-plugin-stealth");
puppeteer.use(StealthPlugin());
(async () => {
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.goto("https://www.google.com/search?q=web+scraping+node+js&hl=en", {
waitUntil: "domcontentloaded",
timeout: 45000,
});
// results render after the document is parsed, so read the DOM a beat later
await new Promise((r) => setTimeout(r, 4000));
// EU and VPN exits meet a consent screen before any results
if (page.url().includes("consent")) {
await page.evaluate(() => document.querySelector("button#L2AGLb")?.click());
await new Promise((r) => setTimeout(r, 5000));
}
const results = await page.evaluate(() =>
[...document.querySelectorAll("h3")].map((h) => h.textContent)
);
console.log("results:", results.length);
console.log(results.slice(0, 5));
await browser.close();
})();The plugin ships a set of evasions that rewrite what Google’s checks read: navigator.webdriver returns undefined, the headless user agent loses its HeadlessChrome marker, the plugins and languages lists stop being empty, and the WebGL vendor string reports real-looking hardware. Each evasion is a separate module, so a specific one can be switched off when it breaks a site.
One measured caveat belongs next to this script. From a VPN exit on a datacenter ASN, Google answered the stealth-patched browser with its /sorry block page and zero results. The plugin patches the browser’s fingerprint, and the IP class stays visible, so on Search this route depends on the quality of the exit. Residential connections pass where datacenter ranges get walled. The anti-bot systems in our Cloudflare guide split traffic the same way.
Scrape using Google SERP API
The API route trades the fight for a request. This calls the Google SERP API directly with axios, no SDK involved:
const axios = require("axios");
const API_KEY = "YOUR-API-KEY";
(async () => {
const response = await axios.get("https://api.hasdata.com/scrape/google/serp", {
params: { q: "web scraping node js", location: "United States" },
headers: { "x-api-key": API_KEY },
});
const results = response.data.organicResults ?? [];
console.log(`Got ${results.length} organic results`);
for (const r of results.slice(0, 3)) {
console.log(`${r.position}. ${r.title}`);
}
})();The call returns:
Got 8 organic results
1. How to do web scraping? : r/node
2. Web Scraping with Node.js: The Best Tools to Use
3. The best Node.js web scrapers for your use caseThe response also carries the SERP’s other blocks (ads, related questions, and an AI Overview token where one appears), each as structured fields, so the parsing this article spent two sections fighting simply arrives done.
Paginate and save the results
Deeper pages come from the start parameter, ten results at a time, and Node’s own fs handles the export. This loop collected two pages and wrote both formats:
const axios = require("axios");
const fs = require("fs");
const API_KEY = "YOUR-API-KEY";
async function fetchPage(start) {
const response = await axios.get("https://api.hasdata.com/scrape/google/serp", {
params: { q: "web scraping node js", location: "United States", num: 10, start },
headers: { "x-api-key": API_KEY },
});
return response.data.organicResults ?? [];
}
(async () => {
const all = [];
for (const start of [0, 10]) {
const page = await fetchPage(start);
console.log(`start=${start}: ${page.length} results`);
all.push(...page);
}
fs.writeFileSync("serp_results.json", JSON.stringify(all, null, 2));
const csv = ["position,title,link"]
.concat(all.map((r) => `${r.position},"${r.title.replaceAll('"', '""')}",${r.link}`))
.join("\n");
fs.writeFileSync("serp_results.csv", csv);
console.log(`Saved ${all.length} rows`);
})();The run collected 8 results on the first page and 9 on the second, 17 rows into both files. Google’s page sizes wobble around ten, so exit the loop on an empty page rather than on a count you expected.
Scraping Google Maps using NodeJS
Maps repeats the same three routes with different behavior in the middle one.
Scrape Google Maps Using Axios and Cheerio
The measured answer from the comparison table applies double here. Maps is a JavaScript application through and through, so the plain-HTTP route returns even less than on Search, and this guide doesn’t spend a section on it.
Scrape Google Maps using Puppeteer
Maps treated the same stealth browser better than Search did. This script ran from the same exit that /sorry’d the SERP attempt and came back with places:
const puppeteer = require("puppeteer-extra");
const StealthPlugin = require("puppeteer-extra-plugin-stealth");
puppeteer.use(StealthPlugin());
(async () => {
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.goto(
"https://www.google.com/maps/search/coffee+shop+in+New+York?hl=en",
{ waitUntil: "domcontentloaded" }
);
if (page.url().includes("consent")) {
await page.evaluate(() => {
const b = [...document.querySelectorAll("button")].find((x) =>
x.textContent.includes("Accept all")
);
b?.click();
});
await new Promise((r) => setTimeout(r, 8000));
}
await page.waitForSelector('div[role="feed"]');
const places = await page.evaluate(() =>
[...document.querySelectorAll('div[role="feed"] div.Nv2PK')].map((card) => ({
name: card.querySelector(".qBF1Pd")?.textContent ?? "",
rating:
card.querySelector('span[aria-label*="stars"]')?.getAttribute("aria-label") ?? "",
}))
);
console.table(places);
await browser.close();
})();The run returned the first batch of cards, “Gold Star Coffee, 4.9 stars 306 Reviews” leading. Scrolling the feed container loads more places, and the class names here rotate on Google’s schedule, so the role="feed" anchor is the part of the selector worth trusting.
Scrape using Google Map API
The API route for Maps returns place data as JSON through the Google Maps API, which is its own endpoint, separate from the SERP one:
const axios = require("axios");
const API_KEY = "YOUR-API-KEY";
(async () => {
const response = await axios.get(
"https://api.hasdata.com/scrape/google-maps/search",
{
params: { q: "coffee shop in New York" },
headers: { "x-api-key": API_KEY },
}
);
const places = response.data.localResults ?? [];
console.log(`Got ${places.length} places`);
for (const p of places.slice(0, 3)) {
console.log(`${p.position}. ${p.title} | rating ${p.rating} (${p.reviews} reviews)`);
}
})();The output:
Got 20 places
1. 787 coffee | rating 4.9 (3564 reviews)
2. Down Under Coffee | rating 4.7 (76 reviews)
3. GRIND THE NYC COFFEE SHOP & BAGEL HOUSE | rating 4.8 (2936 reviews)Each place record carries the id fields (placeId, dataId) that other endpoints take, so reviews and place details chain from this response, the pattern our Maps reviews guide builds on.
Challenges in Web Scraping Google
The five below are the ones these runs kept meeting.
Anti-Scraping Mechanisms
Google fingerprints the client at several layers at once, and the measurements above show them stacking. The same browser passed Maps and got blocked on Search, because the decision weighs the surface, the exit IP and the fingerprint together.
CAPTCHAs
The /sorry page the SERP run met is Google’s challenge wall. Solving it inside a script needs a solver service, and provoking it less needs better exits and slower request patterns.
Rate Limiting
Volume from one IP gets throttled before it’s blocked. Spacing requests and rotating exits both raise the ceiling, and an API route moves the whole problem to the provider’s side.
Changing HTML Structure
Google’s class names are generated and rotate without notice, which is why the selectors above anchor on tags (h3), roles (role="feed") or aria attributes rather than on class names. The one class that could not be avoided, Nv2PK, is the piece most likely to need updating first.
IP Blocking
Datacenter ranges carry a reputation before the first request, which the stealth run demonstrated, and the reputation doesn’t reset quickly once a range is scored. Residential and mobile exits are the standard answer, covered in our rotating proxies overview, and rotating them per session rather than per request keeps the sessions coherent.
Conclusion and takeaways
The measured summary of scraping Google with Node.js fits in three lines. Axios and Cheerio get a 200 with nothing inside, so their role is the five-line check that proves it. Puppeteer with the stealth plugin is the real browser route, and its results track the quality of the IP behind it, Maps being measurably more forgiving than Search. The SERP and Maps APIs return the parsed JSON both fights were about, priced per request, with no SDK needed beyond axios.


