HasData
Back to all posts

Web Scraping Google Search Results in Python

Scraping Google search results looks simple. Send a request, parse the HTML, extract titles and links. In practice, it rarely works that way. Google SERPs are dynamic, full of changing selectors, and protected by strong anti-bot systems. A basic Requests + BeautifulSoup script won’t get you far. Let’s walk through a setup that gets past these roadblocks.

Live Demo

Before we start coding, you can try a live extraction below. This shows the kind of structured data (JSON) successful SERP scraping delivers.

That playground runs on the SERP API. The rest of this article builds the same thing by hand with Selenium, then compares the two.

Building Your Own Google SERP Scraper

To scrape SERPs manually using Python, we will need a headless browser to render the JavaScript and handle the dynamic layout.

Code Overview

Google frequently updates its CSS selectors, so make sure to verify and update them before running the script.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import NoSuchElementException, TimeoutException
import pandas as pd
import json
import urllib.parse

def init_driver():
    options = Options()
    options.add_argument("--headless=new")
    options.add_argument("--window-size=1920,1080")
    driver = webdriver.Chrome(options=options)
    return driver

def build_search_url(query: str):
    encoded = urllib.parse.quote_plus(query)
    return f"https://www.google.com/search?q={encoded}"

def extract_ai_overview(driver):
    try:
        block = driver.find_element(By.CSS_SELECTOR, "div[jsname='dvXlsc']")
    except NoSuchElementException:
        return ""
    return block.text


def extract_people_also_ask(driver):
    questions = []
    for b in driver.find_elements(By.CSS_SELECTOR, "div[jsname='N760b']"):
        found = b.find_elements(By.CSS_SELECTOR, "div.JlqpRe span")
        if found:
            questions.append(found[0].text)
    return questions


def extract_related_searches(driver):
    return [a.text for a in
            driver.find_elements(By.CSS_SELECTOR, "span.dg6jd.JGD2rd")]


def parse_serp(driver, query, max_pages=1):
    results = []
    base_url = build_search_url(query)


    for page in range(max_pages):
        url = base_url + (f"&start={page*10}" if page > 0 else "")
        driver.get(url)

        # wait for the results column instead of sleeping a fixed fifteen
        # seconds, which is both slower than it needs to be and too short
        # whenever the page is slow
        try:
            container = WebDriverWait(driver, 20).until(
                EC.presence_of_element_located((By.ID, "center_col"))
            )
        except TimeoutException:
            continue


        # div.MjjYud also wraps the AI Overview, the PAA block and video
        # carousels, so a card is only organic when it carries all three parts
        for block in container.find_elements(By.CSS_SELECTOR, "div.MjjYud"):
            title = block.find_elements(By.CSS_SELECTOR, "h3")
            link = block.find_elements(By.CSS_SELECTOR, "a")
            snippet = block.find_elements(By.CSS_SELECTOR, "div.VwiC3b")
            if not (title and link and snippet):
                continue
            results.append({
                "Title": title[0].text,
                "Link": link[0].get_attribute("href"),
                "Snippet": snippet[0].text
            })


    ai_overview = extract_ai_overview(driver)
    people_also_ask = extract_people_also_ask(driver)
    related_searches = extract_related_searches(driver)
    return {
        "organic_results": results,
        "ai_overview": ai_overview,
        "people_also_ask": people_also_ask,
        "related_searches": related_searches
    }


def save_data(data, json_filename="serp_full.json"):


    if "organic_results" in data and data["organic_results"]:
        df_organic = pd.DataFrame(data["organic_results"])
        df_organic.to_csv("organic_results.csv", index=False, encoding="utf-8")
        print(f"Saved {len(df_organic)} organic results to organic_results.csv")


    if "ai_overview" in data and data["ai_overview"]:
        df_ai = pd.DataFrame([data["ai_overview"]])
        df_ai.to_csv("ai_overview.csv", index=False, encoding="utf-8")
        print("Saved AI overview to ai_overview.csv")


    if "people_also_ask" in data and data["people_also_ask"]:
        df_paa = pd.DataFrame(data["people_also_ask"], columns=["Question"])
        df_paa.to_csv("people_also_ask.csv", index=False, encoding="utf-8")
        print(f"Saved {len(df_paa)} People Also Ask questions to people_also_ask.csv")


    if "related_searches" in data and data["related_searches"]:
        df_related = pd.DataFrame(data["related_searches"], columns=["Related_Search"])
        df_related.to_csv("related_searches.csv", index=False, encoding="utf-8")
        print(f"Saved {len(df_related)} related searches to related_searches.csv")


    with open(json_filename, "w", encoding="utf-8") as f:
        json.dump(data, f, ensure_ascii=False, indent=4)
    print(f"Saved full SERP data to {json_filename}")


def main():
    query = "what is web scraping"

    driver = init_driver()
    try:
        data = parse_serp(driver, query, max_pages=3)
        save_data(data)
    finally:
        driver.quit()


if __name__ == "__main__":
    main()

The rest of this section takes that script apart, function by function, so you can see which selector does what before you point it at your own query.

Setup and Environment

Install the required libraries:

pip install selenium pandas

Import the necessary modules:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time
import pandas as pd
import json
import urllib.parse

If you are new to scraping, start with our Beginner’s Guide to Web Scraping in Python.

Page Structure Analysis

Extract data from the main blocks on the page:

  1. AI overview, when the query triggers one
  2. Organic results, meaning title, link and snippet
  3. People also ask
  4. People also search for, which Google also labels related searches

A Google results page for "what is coffee" with four blocks boxed and numbered, the AI Overview at the top, an organic result beneath it, the People also ask accordion, and the People also search for grid at the bottom

For each element, find the right CSS selector using Chrome DevTools (press F12 or right-click and Inspect).

Here is the table with the selectors for this project:

TitleSelectorDescription
AI overview containerdiv[jsname=‘dvXlsc’]Google AI overview block.
People also askdiv[jsname=‘N760b’]Each expandable question card in the PAA.
PAA question textdiv.JlqpRe spanThe visible text of the question inside the PAA block.
Related search itemspan.dg6jd.JGD2rdEach suggested query in the related searches section.
Main results container#center_colGoogle SERP’s core results container.
Organic result blockdiv.MjjYudEach individual organic search result card.
Result titleh3 (inside div.MjjYud)The clickable title of an organic search result.
Result linka (inside div.MjjYud)The URL hyperlink pointing to the result’s website.
Result snippetdiv.VwiC3bThe short description/preview text shown under each result title.

Every one of those is a class Google generates, so they change without notice and a scraper built on them needs a check that fails loudly. Writing them is its own skill, covered in the CSS selectors cheat sheet and the XPath guide. If you’d rather not maintain them at all, the SERP API returns the same blocks as structured JSON.

Launch a Headless Browser

Set up a webdriver instance and set options:

def init_driver():
    # Initialize Chrome WebDriver with options
    options = Options()
    options.add_argument("--headless=new")
    # a window size matters headless, since Google serves a narrower layout
    # to a small viewport and the selectors below change with it
    options.add_argument("--window-size=1920,1080")
    driver = webdriver.Chrome(options=options)
    return driver

Build a search URL from the keyword:

def build_search_url(query: str):
    # Encode the query and build a Google search URL
    encoded = urllib.parse.quote_plus(query)
    return f"https://www.google.com/search?q={encoded}"

quote_plus is what keeps a multi-word query intact. Building the URL with an f-string instead breaks the moment a keyword carries an ampersand or a hash, and it won’t tell you it did.

Scrape the AI Overview First

The AI Overview comes before everything else on the page and it is the part of a SERP that changed most in the last year, so it is the block worth reading before the ten blue links. It’s also the one most often absent, since Google decides per query whether to generate one, and an extraction that assumes it’s there fails on the queries that need it least.

def extract_ai_overview(driver):
    # Returns the block's text, or an empty string when there is no overview
    try:
        block = driver.find_element(By.CSS_SELECTOR, "div[jsname='dvXlsc']")
    except NoSuchElementException:
        return ""
    return block.text

An absent overview returns an empty string rather than raising, which matters because Google won’t generate one for every query.

Scrape Organic Search Results

Navigate to the page, wait for it to load, and extract the organic results:

def parse_serp(driver, query, max_pages=1):
    # Parse Google SERP organic results
    results = []
    base_url = build_search_url(query)


    for page in range(max_pages):
        # Add pagination parameter if needed (&start=10, &start=20, etc.)
        url = base_url + (f"&start={page*10}" if page > 0 else "")
        driver.get(url)
        # wait for the results column rather than sleeping a fixed interval
        container = WebDriverWait(driver, 20).until(
            EC.presence_of_element_located((By.ID, "center_col"))
        )

        # div.MjjYud wraps more than organic cards, so keep only the ones
        # carrying a title, a link and a snippet
        for block in container.find_elements(By.CSS_SELECTOR, "div.MjjYud"):
            title = block.find_elements(By.CSS_SELECTOR, "h3")
            link = block.find_elements(By.CSS_SELECTOR, "a")
            snippet = block.find_elements(By.CSS_SELECTOR, "div.VwiC3b")
            if not (title and link and snippet):
                continue
            results.append({
                "Title": title[0].text,
                "Link": link[0].get_attribute("href"),
                "Snippet": snippet[0].text
            })

The three-part check is what keeps the list organic. div.MjjYud also wraps the overview, the PAA accordion and video carousels, and none of those carry all three.

Scrape People Also Ask

The PAA section may not always appear, so wrap the extraction in try/except:

def extract_people_also_ask(driver):
    # Extract "People Also Ask" questions
    questions = []
    for b in driver.find_elements(By.CSS_SELECTOR, "div[jsname='N760b']"):
        found = b.find_elements(By.CSS_SELECTOR, "div.JlqpRe span")
        if found:
            questions.append(found[0].text)
    return questions

Reading the question text with find_elements rather than find_element means a card built differently from the rest is skipped instead of ending the run.

Extract related searches, if they appear:

def extract_related_searches(driver):
    # Extract "Related Searches" suggestions
    return [a.text for a in
            driver.find_elements(By.CSS_SELECTOR, "span.dg6jd.JGD2rd")]

This block is the shortest of the four because related searches are plain text in a single span, with no nesting to walk.

Export Results to CSV/JSON

Save the data as JSON and store each SERP section (organic results, related searches, etc.) in separate CSV files. Print row counts for each file:

def save_data(data, json_filename="serp_full.json"):

    # Save all data into a JSON file
    with open(json_filename, "w", encoding="utf-8") as f:
        json.dump(data, f, ensure_ascii=False, indent=4)
    print(f"Saved full SERP data to {json_filename}")

    # Save organic results to CSV
    if "organic_results" in data and data["organic_results"]:
        df_organic = pd.DataFrame(data["organic_results"])
        df_organic.to_csv("organic_results.csv", index=False, encoding="utf-8")
        print(f"Saved {len(df_organic)} organic results to organic_results.csv")

    # Save AI Overview to CSV
    if "ai_overview" in data and data["ai_overview"]:
        df_ai = pd.DataFrame([data["ai_overview"]])
        df_ai.to_csv("ai_overview.csv", index=False, encoding="utf-8")
        print("Saved AI overview to ai_overview.csv")

    # Save People Also Ask questions to CSV
    if "people_also_ask" in data and data["people_also_ask"]:
        df_paa = pd.DataFrame(data["people_also_ask"], columns=["Question"])
        df_paa.to_csv("people_also_ask.csv", index=False, encoding="utf-8")
        print(f"Saved {len(df_paa)} People Also Ask questions to people_also_ask.csv")

    # Save related searches to CSV
    if "related_searches" in data and data["related_searches"]:
        df_related = pd.DataFrame(data["related_searches"], columns=["Related_Search"])
        df_related.to_csv("related_searches.csv", index=False, encoding="utf-8")
        print(f"Saved {len(df_related)} related searches to related_searches.csv")

Each section gets its own CSV and the whole response one JSON, so a run can be diffed against the previous one without reparsing anything.

The Same Data Through an API

The SERP API returns the same blocks as JSON, with region set by a parameter rather than by where the request comes from. No browser to drive and no selectors to keep up with.

Get Your API Key

To use the API, register on the HasData website and get your API key. The key is activated after email confirmation (or instantly, if you sign up with Google or GitHub).

Basic Google Search Results Scraper

Replace the API key with your own and set the request parameters before running the script.

import requests
import json
import os
import pandas as pd
from urllib.parse import urlencode


BASE_URL = "https://api.hasdata.com/scrape/google/serp"
api_key = "YOUR-API-KEY"


QUERY = "Coffee"
LOCATION = "Austin,Texas,United States"
DEVICE_TYPE = "desktop"
LANG = "en"
GL = "us"


HEADERS = {
    "Content-Type": "application/json",
    "x-api-key": api_key
}


def build_url():
    params = {}
    if QUERY:
        params["q"] = QUERY
    if LOCATION:
        params["location"] = LOCATION
    if DEVICE_TYPE:
        params["deviceType"] = DEVICE_TYPE
    if LANG:
        params["hl"] = LANG
    if GL:
        params["gl"] = GL
    return f"{BASE_URL}?{urlencode(params)}"


def fetch_data():
    url = build_url()
    response = requests.get(url, headers=HEADERS)
    if response.status_code != 200:
        raise Exception(f"Error {response.status_code}: {response.text}")
    return response.json()


def save_csv(items, filename):
    if not items: return
    df = pd.DataFrame(items)
    df.to_csv(filename, index=False, encoding="utf-8")
    print(f"{filename} saved, {len(df)} rows")


def save_json(data, filename):
    with open(filename, "w", encoding="utf-8") as f:
        json.dump(data, f, ensure_ascii=False, indent=4)
    print(f"{filename} saved")


def main():
    data = fetch_data()
    save_csv(data.get("organicResults"), "organic_results.csv")
    local_places = data.get("localResults", {}).get("places")
    save_csv(local_places, "local_places.csv")
    save_csv(data.get("relatedSearches"), "related_searches.csv")
    paa = [{"question": q["question"]} for q in data.get("relatedQuestions", [])]
    save_csv(paa, "people_also_ask.csv")
    kg = data.get("knowledgeGraph", {})
    if kg:
        main_info = {k: v for k, v in kg.items() if k != "nutritionInformation" and k != "headerImages"}
        save_csv([main_info], "knowledge_graph.csv")
        nutrition = kg.get("nutritionInformation")
        if nutrition:
            nutrients = nutrition.get("nutrient", {})
            save_csv([{"description": nutrition.get("description"), **nutrients}], "nutrition.csv")
        save_csv(kg.get("headerImages"), "knowledge_graph_images.csv")


    save_csv(data.get("perspectives"), "perspectives.csv")
    save_json(data, "full_serp.json")


if __name__ == "__main__":
    main()

The same script follows in pieces, one section per block.

Import Libraries

Import the libraries to the project:

import requests
import json
import os
from urllib.parse import urlencode
import pandas as pd

pandas is the one import the script can’t run without and the easiest to forget, since it’s only used in the save step at the very bottom.

Set Parameters

Set your API key and the list of desired parameters (you can find the full list of available Google SERP API parameters in the documentation).

# API base URL
BASE_URL = "https://api.hasdata.com/scrape/google/serp"
api_key = "YOUR-API-key"


# Optional parameters (leave empty if not needed)
QUERY = "Coffee"
LOCATION = "Austin,Texas,United States"
DEVICE_TYPE = "desktop"
LANG = "en"
GL = "us"


# API headers
HEADERS = {
    "Content-Type": "application/json",
    "x-api-key": api_key
}

Leaving a parameter empty keeps it out of the query string entirely, which is why the builder checks each one before adding it.

Make a Request

Build the API request URL from the parameters (some may be optional or missing).

# Build request URL
def build_url():
    params = {}
    if QUERY:
        params["q"] = QUERY
    if LOCATION:
        params["location"] = LOCATION
    if DEVICE_TYPE:
        params["deviceType"] = DEVICE_TYPE
    if LANG:
        params["hl"] = LANG
    if GL:
        params["gl"] = GL
    return f"{BASE_URL}?{urlencode(params)}"

Send the request and receive a JSON response with the search results:

# Get JSON data from API
def fetch_data():
    url = build_url()
    response = requests.get(url, headers=HEADERS)
    if response.status_code != 200:
        raise Exception(f"Error {response.status_code}: {response.text}")
    return response.json()

A non-200 raises with the body attached, so a spent quota or a bad key says so instead of failing later on an empty dictionary. That’s the difference between a message and a stack trace.

Process and Save SERP Sections

Add universal functions to save the data as JSON or CSV:

# Save CSV
def save_csv(items, filename):
    if not items: return
    df = pd.DataFrame(items)
    df.to_csv(filename, index=False, encoding="utf-8")
    print(f"{filename} saved, {len(df)} rows")

# Save JSON
def save_json(data, filename):
    with open(filename, "w", encoding="utf-8") as f:
        json.dump(data, f, ensure_ascii=False, indent=4)
    print(f"{filename} saved")

Example sections to parse and save:

  • Organic results
  • Local places
  • Related searches
  • People also ask (related questions)
  • Knowledge graph (main and additional info, header images)
  • Perspectives

These appear in the sample response, but you can extend the script to parse all the available sections listed in the Google SERP API documentation.

def main():
    data = fetch_data()

    # Organic results
    save_csv(data.get("organicResults"), "organic_results.csv")

    # Local places
    local_places = data.get("localResults", {}).get("places")
    save_csv(local_places, "local_places.csv")

    # Related searches
    save_csv(data.get("relatedSearches"), "related_searches.csv")

    # People Also Ask
    paa = [{"question": q["question"]} for q in data.get("relatedQuestions", [])]
    save_csv(paa, "people_also_ask.csv")

    # Knowledge Graph
    kg = data.get("knowledgeGraph", {})
    if kg:
        # Save main info
        main_info = {k: v for k, v in kg.items() if k != "nutritionInformation" and k != "headerImages"}
        save_csv([main_info], "knowledge_graph.csv")
        # Save nutrition info
        nutrition = kg.get("nutritionInformation")
        if nutrition:
            nutrients = nutrition.get("nutrient", {})
            save_csv([{"description": nutrition.get("description"), **nutrients}], "nutrition.csv")
        # Save header images
        save_csv(kg.get("headerImages"), "knowledge_graph_images.csv")

    # Perspectives
    save_csv(data.get("perspectives"), "perspectives.csv")

    # Full JSON
    save_json(data, "full_serp.json")

The knowledge graph gets special handling because nutrition data and header images are nested one level deeper than everything else in the response.

Using the SERP API Through MCP

Every route above assumes your own code consumes the data. When the consumer is an AI client instead, the same endpoint is reachable over the Model Context Protocol, and there is no scraper to write at all.

Authentication is one header. Point any MCP client at https://mcp.hasdata.com/api/mcp and give it the key:

{
  "mcpServers": {
    "hasdata": {
      "type": "http",
      "url": "https://mcp.hasdata.com/api/mcp",
      "headers": { "x-api-key": "YOUR_API_KEY" }
    }
  }
}

Clients that prefer browser sign-in can authenticate that way instead. Either way the client picks up 63 tools, 22 of them Google. The one matching this article is hasdata_google_serp_serp_getSearchResults, which takes the same q parameter the API route takes.

After the config, the request is a sentence:

Terminal window where a one-line question asks for the first five Google results for best noise cancelling headphones 2026, answered with a five-row markdown table of position, title and domain

The model received parsed fields rather than a page, so nothing between the question and the table touched HTML. Swapping getSearchResults for getAiOverviewResponse returns the AI Overview block the same way.

One limit worth knowing. Ask an agent for thousands of rows and a good one will not call this tool a thousand times. It will write the plain API call from the sections above into a script and run it, because code is cheaper for it than tokens, and a script’s output does not vary. At volume, even the AI writes the script.

Pick the Method That Works Best for You

Scraping Google search results with a browser requires constant selector updates and anti-bot handling. The Google SERP API gives you ready-to-use JSON, so you don’t need to parse HTML or handle captchas. You can also pick the region you want, since the API uses proxies for localization.

Valentina Skakun
Valentina Skakun
Valentina is a software engineer who builds data extraction tools before writing about them. With a strong background in Python, she also leverages her experience in JavaScript, PHP, R, and Ruby to reverse-engineer complex web architectures.If data renders in a browser, she will find a way to script its extraction.
Articles

Might Be Interesting