Scraping Google search results looks simple. Send a request, parse the HTML, extract titles and links. In practice, it rarely works that way. Google SERPs are dynamic, full of changing selectors, and protected by strong anti-bot systems. A basic Requests + BeautifulSoup script won’t get you far. Let’s walk through a setup that gets past these roadblocks.
Live Demo
Before we start coding, you can try a live extraction below. This shows the kind of structured data (JSON) successful SERP scraping delivers.
That playground runs on the SERP API. The rest of this article builds the same thing by hand with Selenium, then compares the two.
Building Your Own Google SERP Scraper
To scrape SERPs manually using Python, we will need a headless browser to render the JavaScript and handle the dynamic layout.
Code Overview
Google frequently updates its CSS selectors, so make sure to verify and update them before running the script.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import NoSuchElementException, TimeoutException
import pandas as pd
import json
import urllib.parse
def init_driver():
options = Options()
options.add_argument("--headless=new")
options.add_argument("--window-size=1920,1080")
driver = webdriver.Chrome(options=options)
return driver
def build_search_url(query: str):
encoded = urllib.parse.quote_plus(query)
return f"https://www.google.com/search?q={encoded}"
def extract_ai_overview(driver):
try:
block = driver.find_element(By.CSS_SELECTOR, "div[jsname='dvXlsc']")
except NoSuchElementException:
return ""
return block.text
def extract_people_also_ask(driver):
questions = []
for b in driver.find_elements(By.CSS_SELECTOR, "div[jsname='N760b']"):
found = b.find_elements(By.CSS_SELECTOR, "div.JlqpRe span")
if found:
questions.append(found[0].text)
return questions
def extract_related_searches(driver):
return [a.text for a in
driver.find_elements(By.CSS_SELECTOR, "span.dg6jd.JGD2rd")]
def parse_serp(driver, query, max_pages=1):
results = []
base_url = build_search_url(query)
for page in range(max_pages):
url = base_url + (f"&start={page*10}" if page > 0 else "")
driver.get(url)
# wait for the results column instead of sleeping a fixed fifteen
# seconds, which is both slower than it needs to be and too short
# whenever the page is slow
try:
container = WebDriverWait(driver, 20).until(
EC.presence_of_element_located((By.ID, "center_col"))
)
except TimeoutException:
continue
# div.MjjYud also wraps the AI Overview, the PAA block and video
# carousels, so a card is only organic when it carries all three parts
for block in container.find_elements(By.CSS_SELECTOR, "div.MjjYud"):
title = block.find_elements(By.CSS_SELECTOR, "h3")
link = block.find_elements(By.CSS_SELECTOR, "a")
snippet = block.find_elements(By.CSS_SELECTOR, "div.VwiC3b")
if not (title and link and snippet):
continue
results.append({
"Title": title[0].text,
"Link": link[0].get_attribute("href"),
"Snippet": snippet[0].text
})
ai_overview = extract_ai_overview(driver)
people_also_ask = extract_people_also_ask(driver)
related_searches = extract_related_searches(driver)
return {
"organic_results": results,
"ai_overview": ai_overview,
"people_also_ask": people_also_ask,
"related_searches": related_searches
}
def save_data(data, json_filename="serp_full.json"):
if "organic_results" in data and data["organic_results"]:
df_organic = pd.DataFrame(data["organic_results"])
df_organic.to_csv("organic_results.csv", index=False, encoding="utf-8")
print(f"Saved {len(df_organic)} organic results to organic_results.csv")
if "ai_overview" in data and data["ai_overview"]:
df_ai = pd.DataFrame([data["ai_overview"]])
df_ai.to_csv("ai_overview.csv", index=False, encoding="utf-8")
print("Saved AI overview to ai_overview.csv")
if "people_also_ask" in data and data["people_also_ask"]:
df_paa = pd.DataFrame(data["people_also_ask"], columns=["Question"])
df_paa.to_csv("people_also_ask.csv", index=False, encoding="utf-8")
print(f"Saved {len(df_paa)} People Also Ask questions to people_also_ask.csv")
if "related_searches" in data and data["related_searches"]:
df_related = pd.DataFrame(data["related_searches"], columns=["Related_Search"])
df_related.to_csv("related_searches.csv", index=False, encoding="utf-8")
print(f"Saved {len(df_related)} related searches to related_searches.csv")
with open(json_filename, "w", encoding="utf-8") as f:
json.dump(data, f, ensure_ascii=False, indent=4)
print(f"Saved full SERP data to {json_filename}")
def main():
query = "what is web scraping"
driver = init_driver()
try:
data = parse_serp(driver, query, max_pages=3)
save_data(data)
finally:
driver.quit()
if __name__ == "__main__":
main()The rest of this section takes that script apart, function by function, so you can see which selector does what before you point it at your own query.
Setup and Environment
Install the required libraries:
pip install selenium pandasImport the necessary modules:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import time
import pandas as pd
import json
import urllib.parseIf you are new to scraping, start with our Beginner’s Guide to Web Scraping in Python.
Page Structure Analysis
Extract data from the main blocks on the page:
- AI overview, when the query triggers one
- Organic results, meaning title, link and snippet
- People also ask
- People also search for, which Google also labels related searches

For each element, find the right CSS selector using Chrome DevTools (press F12 or right-click and Inspect).
Here is the table with the selectors for this project:
| Title | Selector | Description |
|---|---|---|
| AI overview container | div[jsname=‘dvXlsc’] | Google AI overview block. |
| People also ask | div[jsname=‘N760b’] | Each expandable question card in the PAA. |
| PAA question text | div.JlqpRe span | The visible text of the question inside the PAA block. |
| Related search item | span.dg6jd.JGD2rd | Each suggested query in the related searches section. |
| Main results container | #center_col | Google SERP’s core results container. |
| Organic result block | div.MjjYud | Each individual organic search result card. |
| Result title | h3 (inside div.MjjYud) | The clickable title of an organic search result. |
| Result link | a (inside div.MjjYud) | The URL hyperlink pointing to the result’s website. |
| Result snippet | div.VwiC3b | The short description/preview text shown under each result title. |
Every one of those is a class Google generates, so they change without notice and a scraper built on them needs a check that fails loudly. Writing them is its own skill, covered in the CSS selectors cheat sheet and the XPath guide. If you’d rather not maintain them at all, the SERP API returns the same blocks as structured JSON.
Launch a Headless Browser
Set up a webdriver instance and set options:
def init_driver():
# Initialize Chrome WebDriver with options
options = Options()
options.add_argument("--headless=new")
# a window size matters headless, since Google serves a narrower layout
# to a small viewport and the selectors below change with it
options.add_argument("--window-size=1920,1080")
driver = webdriver.Chrome(options=options)
return driverBuild a search URL from the keyword:
def build_search_url(query: str):
# Encode the query and build a Google search URL
encoded = urllib.parse.quote_plus(query)
return f"https://www.google.com/search?q={encoded}"quote_plus is what keeps a multi-word query intact. Building the URL with an f-string instead breaks the moment a keyword carries an ampersand or a hash, and it won’t tell you it did.
Scrape the AI Overview First
The AI Overview comes before everything else on the page and it is the part of a SERP that changed most in the last year, so it is the block worth reading before the ten blue links. It’s also the one most often absent, since Google decides per query whether to generate one, and an extraction that assumes it’s there fails on the queries that need it least.
def extract_ai_overview(driver):
# Returns the block's text, or an empty string when there is no overview
try:
block = driver.find_element(By.CSS_SELECTOR, "div[jsname='dvXlsc']")
except NoSuchElementException:
return ""
return block.textAn absent overview returns an empty string rather than raising, which matters because Google won’t generate one for every query.
Scrape Organic Search Results
Navigate to the page, wait for it to load, and extract the organic results:
def parse_serp(driver, query, max_pages=1):
# Parse Google SERP organic results
results = []
base_url = build_search_url(query)
for page in range(max_pages):
# Add pagination parameter if needed (&start=10, &start=20, etc.)
url = base_url + (f"&start={page*10}" if page > 0 else "")
driver.get(url)
# wait for the results column rather than sleeping a fixed interval
container = WebDriverWait(driver, 20).until(
EC.presence_of_element_located((By.ID, "center_col"))
)
# div.MjjYud wraps more than organic cards, so keep only the ones
# carrying a title, a link and a snippet
for block in container.find_elements(By.CSS_SELECTOR, "div.MjjYud"):
title = block.find_elements(By.CSS_SELECTOR, "h3")
link = block.find_elements(By.CSS_SELECTOR, "a")
snippet = block.find_elements(By.CSS_SELECTOR, "div.VwiC3b")
if not (title and link and snippet):
continue
results.append({
"Title": title[0].text,
"Link": link[0].get_attribute("href"),
"Snippet": snippet[0].text
})The three-part check is what keeps the list organic. div.MjjYud also wraps the overview, the PAA accordion and video carousels, and none of those carry all three.
Scrape People Also Ask
The PAA section may not always appear, so wrap the extraction in try/except:
def extract_people_also_ask(driver):
# Extract "People Also Ask" questions
questions = []
for b in driver.find_elements(By.CSS_SELECTOR, "div[jsname='N760b']"):
found = b.find_elements(By.CSS_SELECTOR, "div.JlqpRe span")
if found:
questions.append(found[0].text)
return questionsReading the question text with find_elements rather than find_element means a card built differently from the rest is skipped instead of ending the run.
Scrape Related Searches
Extract related searches, if they appear:
def extract_related_searches(driver):
# Extract "Related Searches" suggestions
return [a.text for a in
driver.find_elements(By.CSS_SELECTOR, "span.dg6jd.JGD2rd")]This block is the shortest of the four because related searches are plain text in a single span, with no nesting to walk.
Export Results to CSV/JSON
Save the data as JSON and store each SERP section (organic results, related searches, etc.) in separate CSV files. Print row counts for each file:
def save_data(data, json_filename="serp_full.json"):
# Save all data into a JSON file
with open(json_filename, "w", encoding="utf-8") as f:
json.dump(data, f, ensure_ascii=False, indent=4)
print(f"Saved full SERP data to {json_filename}")
# Save organic results to CSV
if "organic_results" in data and data["organic_results"]:
df_organic = pd.DataFrame(data["organic_results"])
df_organic.to_csv("organic_results.csv", index=False, encoding="utf-8")
print(f"Saved {len(df_organic)} organic results to organic_results.csv")
# Save AI Overview to CSV
if "ai_overview" in data and data["ai_overview"]:
df_ai = pd.DataFrame([data["ai_overview"]])
df_ai.to_csv("ai_overview.csv", index=False, encoding="utf-8")
print("Saved AI overview to ai_overview.csv")
# Save People Also Ask questions to CSV
if "people_also_ask" in data and data["people_also_ask"]:
df_paa = pd.DataFrame(data["people_also_ask"], columns=["Question"])
df_paa.to_csv("people_also_ask.csv", index=False, encoding="utf-8")
print(f"Saved {len(df_paa)} People Also Ask questions to people_also_ask.csv")
# Save related searches to CSV
if "related_searches" in data and data["related_searches"]:
df_related = pd.DataFrame(data["related_searches"], columns=["Related_Search"])
df_related.to_csv("related_searches.csv", index=False, encoding="utf-8")
print(f"Saved {len(df_related)} related searches to related_searches.csv")Each section gets its own CSV and the whole response one JSON, so a run can be diffed against the previous one without reparsing anything.
The Same Data Through an API
The SERP API returns the same blocks as JSON, with region set by a parameter rather than by where the request comes from. No browser to drive and no selectors to keep up with.
Get Your API Key
To use the API, register on the HasData website and get your API key. The key is activated after email confirmation (or instantly, if you sign up with Google or GitHub).
Basic Google Search Results Scraper
Replace the API key with your own and set the request parameters before running the script.
import requests
import json
import os
import pandas as pd
from urllib.parse import urlencode
BASE_URL = "https://api.hasdata.com/scrape/google/serp"
api_key = "YOUR-API-KEY"
QUERY = "Coffee"
LOCATION = "Austin,Texas,United States"
DEVICE_TYPE = "desktop"
LANG = "en"
GL = "us"
HEADERS = {
"Content-Type": "application/json",
"x-api-key": api_key
}
def build_url():
params = {}
if QUERY:
params["q"] = QUERY
if LOCATION:
params["location"] = LOCATION
if DEVICE_TYPE:
params["deviceType"] = DEVICE_TYPE
if LANG:
params["hl"] = LANG
if GL:
params["gl"] = GL
return f"{BASE_URL}?{urlencode(params)}"
def fetch_data():
url = build_url()
response = requests.get(url, headers=HEADERS)
if response.status_code != 200:
raise Exception(f"Error {response.status_code}: {response.text}")
return response.json()
def save_csv(items, filename):
if not items: return
df = pd.DataFrame(items)
df.to_csv(filename, index=False, encoding="utf-8")
print(f"{filename} saved, {len(df)} rows")
def save_json(data, filename):
with open(filename, "w", encoding="utf-8") as f:
json.dump(data, f, ensure_ascii=False, indent=4)
print(f"{filename} saved")
def main():
data = fetch_data()
save_csv(data.get("organicResults"), "organic_results.csv")
local_places = data.get("localResults", {}).get("places")
save_csv(local_places, "local_places.csv")
save_csv(data.get("relatedSearches"), "related_searches.csv")
paa = [{"question": q["question"]} for q in data.get("relatedQuestions", [])]
save_csv(paa, "people_also_ask.csv")
kg = data.get("knowledgeGraph", {})
if kg:
main_info = {k: v for k, v in kg.items() if k != "nutritionInformation" and k != "headerImages"}
save_csv([main_info], "knowledge_graph.csv")
nutrition = kg.get("nutritionInformation")
if nutrition:
nutrients = nutrition.get("nutrient", {})
save_csv([{"description": nutrition.get("description"), **nutrients}], "nutrition.csv")
save_csv(kg.get("headerImages"), "knowledge_graph_images.csv")
save_csv(data.get("perspectives"), "perspectives.csv")
save_json(data, "full_serp.json")
if __name__ == "__main__":
main()The same script follows in pieces, one section per block.
Import Libraries
Import the libraries to the project:
import requests
import json
import os
from urllib.parse import urlencode
import pandas as pdpandas is the one import the script can’t run without and the easiest to forget, since it’s only used in the save step at the very bottom.
Set Parameters
Set your API key and the list of desired parameters (you can find the full list of available Google SERP API parameters in the documentation).
# API base URL
BASE_URL = "https://api.hasdata.com/scrape/google/serp"
api_key = "YOUR-API-key"
# Optional parameters (leave empty if not needed)
QUERY = "Coffee"
LOCATION = "Austin,Texas,United States"
DEVICE_TYPE = "desktop"
LANG = "en"
GL = "us"
# API headers
HEADERS = {
"Content-Type": "application/json",
"x-api-key": api_key
}Leaving a parameter empty keeps it out of the query string entirely, which is why the builder checks each one before adding it.
Make a Request
Build the API request URL from the parameters (some may be optional or missing).
# Build request URL
def build_url():
params = {}
if QUERY:
params["q"] = QUERY
if LOCATION:
params["location"] = LOCATION
if DEVICE_TYPE:
params["deviceType"] = DEVICE_TYPE
if LANG:
params["hl"] = LANG
if GL:
params["gl"] = GL
return f"{BASE_URL}?{urlencode(params)}"Send the request and receive a JSON response with the search results:
# Get JSON data from API
def fetch_data():
url = build_url()
response = requests.get(url, headers=HEADERS)
if response.status_code != 200:
raise Exception(f"Error {response.status_code}: {response.text}")
return response.json()A non-200 raises with the body attached, so a spent quota or a bad key says so instead of failing later on an empty dictionary. That’s the difference between a message and a stack trace.
Process and Save SERP Sections
Add universal functions to save the data as JSON or CSV:
# Save CSV
def save_csv(items, filename):
if not items: return
df = pd.DataFrame(items)
df.to_csv(filename, index=False, encoding="utf-8")
print(f"{filename} saved, {len(df)} rows")
# Save JSON
def save_json(data, filename):
with open(filename, "w", encoding="utf-8") as f:
json.dump(data, f, ensure_ascii=False, indent=4)
print(f"{filename} saved")Example sections to parse and save:
- Organic results
- Local places
- Related searches
- People also ask (related questions)
- Knowledge graph (main and additional info, header images)
- Perspectives
These appear in the sample response, but you can extend the script to parse all the available sections listed in the Google SERP API documentation.
def main():
data = fetch_data()
# Organic results
save_csv(data.get("organicResults"), "organic_results.csv")
# Local places
local_places = data.get("localResults", {}).get("places")
save_csv(local_places, "local_places.csv")
# Related searches
save_csv(data.get("relatedSearches"), "related_searches.csv")
# People Also Ask
paa = [{"question": q["question"]} for q in data.get("relatedQuestions", [])]
save_csv(paa, "people_also_ask.csv")
# Knowledge Graph
kg = data.get("knowledgeGraph", {})
if kg:
# Save main info
main_info = {k: v for k, v in kg.items() if k != "nutritionInformation" and k != "headerImages"}
save_csv([main_info], "knowledge_graph.csv")
# Save nutrition info
nutrition = kg.get("nutritionInformation")
if nutrition:
nutrients = nutrition.get("nutrient", {})
save_csv([{"description": nutrition.get("description"), **nutrients}], "nutrition.csv")
# Save header images
save_csv(kg.get("headerImages"), "knowledge_graph_images.csv")
# Perspectives
save_csv(data.get("perspectives"), "perspectives.csv")
# Full JSON
save_json(data, "full_serp.json")The knowledge graph gets special handling because nutrition data and header images are nested one level deeper than everything else in the response.
Using the SERP API Through MCP
Every route above assumes your own code consumes the data. When the consumer is an AI client instead, the same endpoint is reachable over the Model Context Protocol, and there is no scraper to write at all.
Authentication is one header. Point any MCP client at https://mcp.hasdata.com/api/mcp and give it the key:
{
"mcpServers": {
"hasdata": {
"type": "http",
"url": "https://mcp.hasdata.com/api/mcp",
"headers": { "x-api-key": "YOUR_API_KEY" }
}
}
}Clients that prefer browser sign-in can authenticate that way instead. Either way the client picks up 63 tools, 22 of them Google. The one matching this article is hasdata_google_serp_serp_getSearchResults, which takes the same q parameter the API route takes.
After the config, the request is a sentence:

The model received parsed fields rather than a page, so nothing between the question and the table touched HTML. Swapping getSearchResults for getAiOverviewResponse returns the AI Overview block the same way.
One limit worth knowing. Ask an agent for thousands of rows and a good one will not call this tool a thousand times. It will write the plain API call from the sections above into a script and run it, because code is cheaper for it than tokens, and a script’s output does not vary. At volume, even the AI writes the script.
Pick the Method That Works Best for You
Scraping Google search results with a browser requires constant selector updates and anti-bot handling. The Google SERP API gives you ready-to-use JSON, so you don’t need to parse HTML or handle captchas. You can also pick the region you want, since the API uses proxies for localization.


