Walmart runs on Next.js, and that decides how to scrape it. Every product page carries its full application state in a <script id="__NEXT_DATA__"> tag, and on the 110 product cards we measured this month that state held a median of 907 data fields per product, against 4 of five fields for CSS selectors and zero cards with a Product JSON-LD schema. This guide builds the scraper on that source, measures all three ways of reading a card, and covers the two things that actually block Walmart scrapers: Akamai Bot Manager and HUMAN Security.
The Complete Scraper
The script below reads product URLs from links.csv, fetches every page through the Web Scraping API with a US residential exit, parses the __NEXT_DATA__ state, and writes five fields per product to result.csv. The API key comes from the dashboard after signing up, and the fetch runs through the API because a direct request gets the challenge page (measured below).
import csv
import json
import os
import requests
from bs4 import BeautifulSoup
API_KEY = os.environ["HASDATA_API_KEY"]
def fetch(url):
# residential exit required: Walmart serves the challenge to datacenter IPs
r = requests.post(
"https://api.hasdata.com/scrape/web",
headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
json={"url": url, "proxyType": "residential", "proxyCountry": "US",
"outputFormat": ["html"]},
timeout=120,
)
return r.text
def parse_card(html):
soup = BeautifulSoup(html, "lxml")
tag = soup.find("script", id="__NEXT_DATA__")
product = (json.loads(tag.text)["props"]["pageProps"]
["initialData"]["data"]["product"])
price = (product.get("priceInfo") or {}).get("currentPrice") or {}
return {
"title": product.get("name"),
"price": price.get("price"),
"rating": product.get("averageRating"),
"reviews": product.get("numberOfReviews"),
"image": (product.get("imageInfo") or {}).get("thumbnailUrl"),
}
with open("links.csv") as f:
links = [line.strip() for line in f if line.strip()]
with open("result.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["title", "price", "rating", "reviews", "image"])
writer.writeheader()
for link in links:
try:
writer.writerow(parse_card(fetch(link)))
except (KeyError, TypeError, ValueError) as e:
print(f"skipped {link}: {e}")
print(f"done, {len(links)} links processed")Each part of this script gets its own section below. Where __NEXT_DATA__ comes from and what else it holds, how to fill links.csv from search pages, and why the request goes through a residential proxy.
Three Ways to Read a Walmart Product Page
The same product card carries its data in three places, and they don’t agree. We fetched 110 product cards from 18 search categories and extracted each card three ways. The table is the whole argument:
| Source | Cards where it works | Fields it yields | Notes |
|---|---|---|---|
| CSS selectors (five core fields) | 97% of cards | 4 of 5 fields | the rating selector returned nothing on any card |
| JSON-LD Product schema | 0 of 110 cards | 0 | the Product schema is gone from product pages |
__NEXT_DATA__ state | 107 of 110 cards | 446 to 2,943 per card, median 907 | the page’s own application state |
One note keeps these numbers honest. The 110 cards came from 18 search categories, electronics through groceries, and each was fetched once through a US residential exit with no JavaScript rendering, and “fields” counts leaf values in the extracted structure: five for the selector route by construction, every leaf of the Product schema for JSON-LD, and every leaf of the product node for __NEXT_DATA__. Zero fetches hit the challenge page on this route.
Way 1. Parsing the Page Body
CSS selectors read the rendered markup, and they still work for four of the five core fields. Find the element in DevTools (F12, then Ctrl + Shift + C and click the element), note the tag and its attributes, and query them with BeautifulSoup:

The title lives in the h1 with an itemprop of name, and the price sits in a neighbouring span marked the same way:

The review count carries an itemprop of ratingCount, and the image is most reliable from the og:image meta tag, which exists on every card even when the visible img markup shifts:
import os
import requests
from bs4 import BeautifulSoup
API_KEY = os.environ["HASDATA_API_KEY"]
r = requests.post(
"https://api.hasdata.com/scrape/web",
headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
json={"url": "https://www.walmart.com/ip/771229626",
"proxyType": "residential", "proxyCountry": "US", "outputFormat": ["html"]},
timeout=120,
)
soup = BeautifulSoup(r.text, "lxml")
title = soup.find("h1", attrs={"itemprop": "name"})
price = soup.find("span", attrs={"itemprop": "price"})
reviews = soup.find("a", attrs={"itemprop": "ratingCount"})
rating = soup.find("span", attrs={"class": "rating-number"})
image = soup.find("meta", property="og:image")
for label, node in [("title", title), ("price", price), ("reviews", reviews),
("rating", rating), ("image", image)]:
value = None
if node is not None:
value = node.get("content") if label == "image" else node.get_text(strip=True)
print(f"{label}: {value}")On our 110 cards the title, price, reviews, and image selectors each returned data on 97% of pages, and span.rating-number returned nothing on any of them, so the script prints the rating as None. That’s the trade selectors make. Each field is one line, and every field lives or dies with the markup. A redesign that renames one class removes that field from your dataset until you notice.
Way 2. Parsing the JSON-LD Structured Data
Look closely at the page code and you’ll see the data isn’t only in the page body. Let’s pay attention to the <head>…</head> tag, more precisely to its <script nonce type="application/ld+json">…</script> tag.

This is the “Product” Schema Markup in JSON format. The product schema allows adding specific product attributes to product listings that can appear as rich results on the search engine results page (SERP). Let’s copy it and format it into a convenient form for research:
{
"@context": "https://schema.org",
"@type": "Product",
"image": "https://i5.walmartimages.com/asr/ce0f57f7-ad6f-4e0b-a7ae-f751068597c2_1.b7e1f1bab1fd7f98cb9aef1ae9b783fb.png",
"name": "Canon EOS Rebel T100 Digital SLR Camera with 18-55mm Lens Kit, 18 Megapixel Sensor, Wi-Fi, DIGIC4+, SanDisk 32GB Memory Card and Live View Shooting",
"sku": "771229626",
"gtin13": "013803300550",
"description": "<p>Creating distinctive stories with DSLR quality photos and Full HD movies is easier than you think with the 18 Megapixel Canon EOS Rebel T100. Share instantly and shoot remotely via your compatible smartphone with Wi-Fi and the Canon Camera Connect app. The powerful 18 Megapixel sensor has up to 19 times more surface area than many smartphones, and you can instantly transfer photos and movies to your smart device. The Canon EOS Rebel T100 has a Scene Intelligent Auto feature that allows you to simply frame and shoot for great results. It also features Guided Live View shooting with Creative Auto mode, and you can add unique finishes with Creative Filters. The Canon EOS Rebel T100 makes it fast and easy to share all the moments that matter.</p>",
"model": "T100",
"brand": {
"@type": "Brand",
"name": "Canon"
},
"offers": {
"@type": "Offer",
"url": "https://www.walmart.com/ip/Canon-EOS-Rebel-T100-Digital-SLR-Camera-with-18-55mm-Lens-Kit-18-Megapixel-Sensor-Wi-Fi-DIGIC4-SanDisk-32GB-Memory-Card-and-Live-View-Shooting/771229626",
"priceCurrency": "USD",
"price": 299,
"availability": "https://schema.org/InStock",
"itemCondition": "https://schema.org/NewCondition",
"availableDeliveryMethod": "https://schema.org/OnSitePickup"
},
"review": [
{
"@type": "Review",
"name": "Great camera for beginners",
"datePublished": "January 4, 2020",
"reviewBody": "Love this camera....",
"reviewRating": {
"@type": "Rating",
"worstRating": 1,
"ratingValue": 5,
"bestRating": 5
},
"author": {
"@type": "Person",
"name": "Sparkles"
}
},
{
"@type": "Review",
"name": "Perfect for beginners",
"datePublished": "January 7, 2020",
"reviewBody": "I am so in love with this camera!...",
"reviewRating": {
"@type": "Rating",
"worstRating": 1,
"ratingValue": 5,
"bestRating": 5
},
"author": {
"@type": "Person",
"name": "Brazilchick32"
}
},
{
"@type": "Review",
"name": "Great camera",
"datePublished": "January 17, 2020",
"reviewBody": "I really love all the features this camera has. Every time I use it, I'm discovering a new one. I'm pretty technologically challenged, but this hasn't hindered me. The zoom and focus give very detailed and sharp images. I cannot wait to take it on my next trip as right now I've only photographed the dog a million times",
"reviewRating": {
"@type": "Rating",
"worstRating": 1,
"ratingValue": 5,
"bestRating": 5
},
"author": {
"@type": "Person",
"name": "userfriendly"
}
}
],
"aggregateRating": {
"@type": "AggregateRating",
"ratingValue": 4.5,
"bestRating": 5,
"reviewCount": 172
}
}That’s enough data, and it’s in a friendlier shape than the markup. Read the contents of the <head>…</head> tag, pick out the tag that holds the JSON, then set the variables:
- Title value of the
[name]attribute. - Price value of the
[offers][price]attribute. - Reviews value of the
[aggregateRating][reviewCount]attribute. - Rating value of the
[aggregateRating][ratingValue]attribute. - Image value of the
[image]attribute.
For the JSON you only need the built-in library:
import jsonLet’s create a data variable in which we put the JSON data:
data = (json.loads(soup.find('script', attrs={'type': 'application/ld+json'}).text))Then put the data into the variables:
title = data['name']
price = data['offers']['price']
reviews = data['aggregateRating']['reviewCount']
rating = data['aggregateRating']['ratingValue']
image = data['image']The rest will remain the same. Check the script execution:

Full script:
from bs4 import BeautifulSoup
import requests
import json
with open("result.csv", "w") as f:
f.write("title; price; rating; reviews; image\n")
with open("links.csv", "r+") as links:
for link in links:
html_text = requests.get(link).text
soup = BeautifulSoup(html_text, 'lxml')
data = (json.loads(soup.find('script', attrs={'type': 'application/ld+json'}).text))
title = data['name']
price = data['offers']['price']
reviews = data['aggregateRating']['reviewCount']
rating = data['aggregateRating']['ratingValue']
image = data['image']
try:
with open("result.csv", "a") as f:
f.write(str(title)+"; "+str(price)+"; "+str(reviews)+"; "+str(rating)+"; "+str(image)+"\n")
except Exception as e:
print("There is no data")That section stays as the reference for the technique, and the numbers above say what happened to it on Walmart. Of the 110 product cards we measured, 0 carried a Product JSON-LD schema. The rich Product markup now appears only as a small WebPage stub, so on Walmart the technique’s successor is the block the framework itself writes.
Way 3. The NEXT_DATA Application State
Walmart’s pages are rendered by Next.js, and the framework embeds the full state of the page as JSON in a <script id="__NEXT_DATA__"> tag. The product record sits at props.pageProps.initialData.data.product, and the page is rendered from it. Our e-commerce scraping guide places this source at the top of its data-source hierarchy for exactly this reason, and Walmart is the cleanest example of it:
import json
import os
import requests
from bs4 import BeautifulSoup
API_KEY = os.environ["HASDATA_API_KEY"]
r = requests.post(
"https://api.hasdata.com/scrape/web",
headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
json={"url": "https://www.walmart.com/ip/771229626",
"proxyType": "residential", "proxyCountry": "US", "outputFormat": ["html"]},
timeout=120,
)
soup = BeautifulSoup(r.text, "lxml")
state = json.loads(soup.find("script", id="__NEXT_DATA__").text)
product = state["props"]["pageProps"]["initialData"]["data"]["product"]
def count_leaves(node):
if isinstance(node, dict):
return sum(count_leaves(v) for v in node.values())
if isinstance(node, list):
return sum(count_leaves(v) for v in node)
return 1
print(product["name"])
print(product["priceInfo"]["currentPrice"]["priceString"],
"|", product["averageRating"], "stars,", product["numberOfReviews"], "reviews")
print("fields in the product node:", count_leaves(product))The run prints the product name, price, and rating, then the field count:
Canon EOS Rebel T100 DSLR Camera EF-S 18-55mm f/3.5-5.6 IS II Lens Kit, 18 Megapixel CMOS (APS-C) Sensor, Full HD Videos, Built-in Wi-Fi, Beginner Photographers, Digital Camera, Black
$448.86 | 4.4 stars, 727 reviews
fields in the product node: 438Across the 110 measured cards the product node carried between 446 and 2,943 leaf fields, median 907. Variants, per-store availability, seller records, fulfillment options, badge and promotion state, review statistics. The five fields the selectors extract are all in there, plus roughly a thousand that never reach the page. One json.loads replaces every selector, and the path breaks only when Walmart changes its data model rather than its styling.
Collecting Product URLs
Product links come from search pages, and search pages are Next.js too. The visible result grid links products under /ip/ paths, and the same __NEXT_DATA__ trick works there as well, but anchor tags are enough for a link collector. The script reads keywords from keywords.txt, fetches one search page per keyword, and keeps every product link plus the related-searches carousel:
import os
import urllib.parse
import requests
from bs4 import BeautifulSoup
API_KEY = os.environ["HASDATA_API_KEY"]
with open("keywords.txt") as f:
keywords = [line.strip() for line in f if line.strip()]
product_links = []
related_searches = []
for keyword in keywords:
search_url = "https://www.walmart.com/search?q=" + urllib.parse.quote(keyword)
r = requests.post(
"https://api.hasdata.com/scrape/web",
headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
json={"url": search_url, "proxyType": "residential", "proxyCountry": "US",
"outputFormat": ["html"]},
timeout=120,
)
soup = BeautifulSoup(r.text, "lxml")
for a in soup.find_all("a", href=True):
href = a["href"].split("?")[0]
if "/ip/" in href:
full = "https://www.walmart.com" + href if href.startswith("/") else href
if full not in product_links:
product_links.append(full)
carousel = soup.find("ul", {"data-testid": "carousel-container"})
if carousel:
related_searches += [li.get_text(strip=True) for li in carousel.find_all("li")]
with open("links.csv", "w") as f:
f.write("\n".join(product_links))
print(len(product_links), "product links,", len(related_searches), "related searches")The related-searches carousel is worth keeping too, since its terms are Walmart’s own suggested keywords and make good seeds for the next run:

A run over our 18 test keywords collected 110+ unique product links, which is exactly the links.csv the complete scraper at the top consumes. For deeper collection, page through search results with the page URL parameter, and expect the same anti-bot wall as on product pages. The collection logic mirrors what we built for Amazon product scraping, where the search-to-product pipeline is covered in more depth.
Walmart Anti-Bot Protection
Walmart’s protection is run by Akamai Bot Manager with HUMAN Security providing the challenge, and the “Press & Hold” page most scrapers meet is HUMAN’s verification widget:

That check was measured rather than assumed, and headers alone do not clear it. A plain requests call with a current Chrome User-Agent:
import requests
from bs4 import BeautifulSoup
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36"}
r = requests.get("https://www.walmart.com/ip/771229626", headers=headers, timeout=40)
soup = BeautifulSoup(r.text, "lxml")
print(r.status_code, "|", soup.title.get_text(strip=True), "|", len(r.text), "bytes")The print line reports the status, the page title, and the byte count:
200 | Robot or human? | 15190 bytesThe response is HTTP 200 and a 15 KB “Robot or human?” page instead of the roughly 500 KB product page. The same request through a datacenter proxy returned the same challenge, so the defense reads more than the IP’s country. What it reads is the connection itself: Akamai fingerprints the TLS handshake, and a Python requests client doesn’t shake hands like Chrome no matter what its headers claim. Clients that impersonate a browser’s TLS fingerprint (curl_cffi is the common one in Python) address that layer, and a residential exit addresses the IP reputation layer. In our measurement, a residential proxy with no header tricks at all returned the full product page on 110 of 110 cards. The general blocking-avoidance guide covers the rest of the discipline, rate limits and request patterns included.
Scrape Walmart with the HasData API
Two routes exist and they answer different questions.
The dedicated Walmart endpoints cover search, product and reviews, and return the fields already parsed. One search call for laptop comes back with 52 products, so nothing in this article’s parsing code is needed for that path:
import os
import requests
API_KEY = os.environ["HASDATA_API_KEY"]
response = requests.get(
"https://api.hasdata.com/scrape/walmart/search/",
params={"q": "laptop"},
headers={"x-api-key": API_KEY},
timeout=120,
)
response.raise_for_status()
products = response.json()["productResults"]
print(len(products), "products")The Web Scraping API is the other route, and it’s the one this article’s approach needs. It returns the raw page rather than parsed fields, which is what the __NEXT_DATA__ work above reads. Use it when you want the roughly thousand fields in the page state rather than the handful a parsed response carries:
response = requests.post(
"https://api.hasdata.com/scrape/web",
headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
json={
"url": "https://www.walmart.com/ip/771229626",
"proxyType": "residential", # datacenter exits get the challenge page
"proxyCountry": "US", # prices and availability differ by region
"outputFormat": ["html"], # the response body is the raw page
},
timeout=120,
)
html = response.text
print(response.status_code, "|", len(html), "bytes of product page")proxyType must be residential here, because datacenter exits get the challenge page, and proxyCountry pins the store region, since prices and availability differ by location and a request from a random country returns plausible but wrong numbers. JavaScript rendering isn’t needed, since __NEXT_DATA__ is server-rendered into the initial HTML.
Conclusion
Scraping Walmart in 2026 is one JSON path, not thirty selectors. The page’s __NEXT_DATA__ state carries a median of 907 fields per product where selectors yield four of five and the JSON-LD schema yields nothing anymore, so parse the state and treat the markup as a fallback. The blocking layer is the real work, and it is fingerprint-deep, which is why this guide routes requests through a residential exit instead of collecting header tricks. Point the complete scraper at your own links.csv and the rest is a CSV file filling up.


