HasData
Back to all posts

How to Scrape Walmart with Python (2026)

Walmart runs on Next.js, and that decides how to scrape it. Every product page carries its full application state in a <script id="__NEXT_DATA__"> tag, and on the 110 product cards we measured this month that state held a median of 907 data fields per product, against 4 of five fields for CSS selectors and zero cards with a Product JSON-LD schema. This guide builds the scraper on that source, measures all three ways of reading a card, and covers the two things that actually block Walmart scrapers: Akamai Bot Manager and HUMAN Security.

The Complete Scraper

The script below reads product URLs from links.csv, fetches every page through the Web Scraping API with a US residential exit, parses the __NEXT_DATA__ state, and writes five fields per product to result.csv. The API key comes from the dashboard after signing up, and the fetch runs through the API because a direct request gets the challenge page (measured below).

import csv
import json
import os

import requests
from bs4 import BeautifulSoup

API_KEY = os.environ["HASDATA_API_KEY"]

def fetch(url):
    # residential exit required: Walmart serves the challenge to datacenter IPs
    r = requests.post(
        "https://api.hasdata.com/scrape/web",
        headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
        json={"url": url, "proxyType": "residential", "proxyCountry": "US",
              "outputFormat": ["html"]},
        timeout=120,
    )
    return r.text

def parse_card(html):
    soup = BeautifulSoup(html, "lxml")
    tag = soup.find("script", id="__NEXT_DATA__")
    product = (json.loads(tag.text)["props"]["pageProps"]
               ["initialData"]["data"]["product"])
    price = (product.get("priceInfo") or {}).get("currentPrice") or {}
    return {
        "title": product.get("name"),
        "price": price.get("price"),
        "rating": product.get("averageRating"),
        "reviews": product.get("numberOfReviews"),
        "image": (product.get("imageInfo") or {}).get("thumbnailUrl"),
    }

with open("links.csv") as f:
    links = [line.strip() for line in f if line.strip()]

with open("result.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["title", "price", "rating", "reviews", "image"])
    writer.writeheader()
    for link in links:
        try:
            writer.writerow(parse_card(fetch(link)))
        except (KeyError, TypeError, ValueError) as e:
            print(f"skipped {link}: {e}")

print(f"done, {len(links)} links processed")

Each part of this script gets its own section below. Where __NEXT_DATA__ comes from and what else it holds, how to fill links.csv from search pages, and why the request goes through a residential proxy.

Three Ways to Read a Walmart Product Page

The same product card carries its data in three places, and they don’t agree. We fetched 110 product cards from 18 search categories and extracted each card three ways. The table is the whole argument:

SourceCards where it worksFields it yieldsNotes
CSS selectors (five core fields)97% of cards4 of 5 fieldsthe rating selector returned nothing on any card
JSON-LD Product schema0 of 110 cards0the Product schema is gone from product pages
__NEXT_DATA__ state107 of 110 cards446 to 2,943 per card, median 907the page’s own application state

One note keeps these numbers honest. The 110 cards came from 18 search categories, electronics through groceries, and each was fetched once through a US residential exit with no JavaScript rendering, and “fields” counts leaf values in the extracted structure: five for the selector route by construction, every leaf of the Product schema for JSON-LD, and every leaf of the product node for __NEXT_DATA__. Zero fetches hit the challenge page on this route.

Way 1. Parsing the Page Body

CSS selectors read the rendered markup, and they still work for four of the five core fields. Find the element in DevTools (F12, then Ctrl + Shift + C and click the element), note the tag and its attributes, and query them with BeautifulSoup:

Product Title Selector

Product title selector

The title lives in the h1 with an itemprop of name, and the price sits in a neighbouring span marked the same way:

Product Price Selector

Product price selector

The review count carries an itemprop of ratingCount, and the image is most reliable from the og:image meta tag, which exists on every card even when the visible img markup shifts:

import os

import requests
from bs4 import BeautifulSoup

API_KEY = os.environ["HASDATA_API_KEY"]

r = requests.post(
    "https://api.hasdata.com/scrape/web",
    headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
    json={"url": "https://www.walmart.com/ip/771229626",
          "proxyType": "residential", "proxyCountry": "US", "outputFormat": ["html"]},
    timeout=120,
)
soup = BeautifulSoup(r.text, "lxml")

title = soup.find("h1", attrs={"itemprop": "name"})
price = soup.find("span", attrs={"itemprop": "price"})
reviews = soup.find("a", attrs={"itemprop": "ratingCount"})
rating = soup.find("span", attrs={"class": "rating-number"})
image = soup.find("meta", property="og:image")

for label, node in [("title", title), ("price", price), ("reviews", reviews),
                    ("rating", rating), ("image", image)]:
    value = None
    if node is not None:
        value = node.get("content") if label == "image" else node.get_text(strip=True)
    print(f"{label}: {value}")

On our 110 cards the title, price, reviews, and image selectors each returned data on 97% of pages, and span.rating-number returned nothing on any of them, so the script prints the rating as None. That’s the trade selectors make. Each field is one line, and every field lives or dies with the markup. A redesign that renames one class removes that field from your dataset until you notice.

Way 2. Parsing the JSON-LD Structured Data

Look closely at the page code and you’ll see the data isn’t only in the page body. Let’s pay attention to the <head>…</head> tag, more precisely to its <script nonce type="application/ld+json">…</script> tag.

Walmart “Product” Schema Markup in JSON format

Walmart “Product” Schema Markup in JSON format

This is the “Product” Schema Markup in JSON format. The product schema allows adding specific product attributes to product listings that can appear as rich results on the search engine results page (SERP). Let’s copy it and format it into a convenient form for research:

{
    "@context": "https://schema.org",
    "@type": "Product",
    "image": "https://i5.walmartimages.com/asr/ce0f57f7-ad6f-4e0b-a7ae-f751068597c2_1.b7e1f1bab1fd7f98cb9aef1ae9b783fb.png",
    "name": "Canon EOS Rebel T100 Digital SLR Camera with 18-55mm Lens Kit, 18 Megapixel Sensor, Wi-Fi, DIGIC4+, SanDisk 32GB Memory Card and Live View Shooting",
    "sku": "771229626",
    "gtin13": "013803300550",
    "description": "<p>Creating distinctive stories with DSLR quality photos and Full HD movies is easier than you think with the 18 Megapixel Canon EOS Rebel T100. Share instantly and shoot remotely via your compatible smartphone with Wi-Fi and the Canon Camera Connect app. The powerful 18 Megapixel sensor has up to 19 times more surface area than many smartphones, and you can instantly transfer photos and movies to your smart device. The Canon EOS Rebel T100 has a Scene Intelligent Auto feature that allows you to simply frame and shoot for great results. It also features Guided Live View shooting with Creative Auto mode, and you can add unique finishes with Creative Filters. The Canon EOS Rebel T100 makes it fast and easy to share all the moments that matter.</p>",
    "model": "T100",
    "brand": {
        "@type": "Brand",
        "name": "Canon"
    },
    "offers": {
        "@type": "Offer",
        "url": "https://www.walmart.com/ip/Canon-EOS-Rebel-T100-Digital-SLR-Camera-with-18-55mm-Lens-Kit-18-Megapixel-Sensor-Wi-Fi-DIGIC4-SanDisk-32GB-Memory-Card-and-Live-View-Shooting/771229626",
        "priceCurrency": "USD",
        "price": 299,
        "availability": "https://schema.org/InStock",
        "itemCondition": "https://schema.org/NewCondition",
        "availableDeliveryMethod": "https://schema.org/OnSitePickup"
    },
    "review": [
        {
            "@type": "Review",
            "name": "Great camera for beginners",
            "datePublished": "January 4, 2020",
            "reviewBody": "Love this camera....",
            "reviewRating": {
                "@type": "Rating",
                "worstRating": 1,
                "ratingValue": 5,
                "bestRating": 5
            },
            "author": {
                "@type": "Person",
                "name": "Sparkles"
            }
        },
        {
            "@type": "Review",
            "name": "Perfect for beginners",
            "datePublished": "January 7, 2020",
            "reviewBody": "I am so in love with this camera!...",
            "reviewRating": {
                "@type": "Rating",
                "worstRating": 1,
                "ratingValue": 5,
                "bestRating": 5
            },
            "author": {
                "@type": "Person",
                "name": "Brazilchick32"
            }
        },
        {
            "@type": "Review",
            "name": "Great camera",
            "datePublished": "January 17, 2020",
            "reviewBody": "I really love all the features this camera has. Every time I use it, I'm discovering a new one. I'm pretty technologically challenged, but this hasn't hindered me. The zoom and focus give very detailed and sharp images. I cannot wait to take it on my next trip as right now I've only photographed the dog a million times",
            "reviewRating": {
                "@type": "Rating",
                "worstRating": 1,
                "ratingValue": 5,
                "bestRating": 5
            },
            "author": {
                "@type": "Person",
                "name": "userfriendly"
            }
        }
    ],
    "aggregateRating": {
        "@type": "AggregateRating",
        "ratingValue": 4.5,
        "bestRating": 5,
        "reviewCount": 172
    }
}

That’s enough data, and it’s in a friendlier shape than the markup. Read the contents of the <head>…</head> tag, pick out the tag that holds the JSON, then set the variables:

  1. Title value of the [name] attribute.
  2. Price value of the [offers][price] attribute.
  3. Reviews value of the [aggregateRating][reviewCount] attribute.
  4. Rating value of the [aggregateRating][ratingValue] attribute.
  5. Image value of the [image] attribute.

For the JSON you only need the built-in library:

import json

Let’s create a data variable in which we put the JSON data:

data = (json.loads(soup.find('script', attrs={'type': 'application/ld+json'}).text))

Then put the data into the variables:

    title = data['name']
    price = data['offers']['price']
    reviews = data['aggregateRating']['reviewCount']
    rating = data['aggregateRating']['ratingValue']
    image = data['image']

The rest will remain the same. Check the script execution:

Results table with extracted product data

Results table with extracted product data

Full script:

from bs4 import BeautifulSoup
import requests
import json

with open("result.csv", "w") as f:
  f.write("title; price; rating; reviews; image\n")
with open("links.csv", "r+") as links:
  for link in links:
    html_text = requests.get(link).text
    soup = BeautifulSoup(html_text, 'lxml')
    data = (json.loads(soup.find('script', attrs={'type': 'application/ld+json'}).text))
    title = data['name']
    price = data['offers']['price']
    reviews = data['aggregateRating']['reviewCount']
    rating = data['aggregateRating']['ratingValue']
    image = data['image']
    try:
        with open("result.csv", "a") as f:
            f.write(str(title)+"; "+str(price)+"; "+str(reviews)+"; "+str(rating)+"; "+str(image)+"\n")
    except Exception as e:
        print("There is no data")

That section stays as the reference for the technique, and the numbers above say what happened to it on Walmart. Of the 110 product cards we measured, 0 carried a Product JSON-LD schema. The rich Product markup now appears only as a small WebPage stub, so on Walmart the technique’s successor is the block the framework itself writes.

Way 3. The NEXT_DATA Application State

Walmart’s pages are rendered by Next.js, and the framework embeds the full state of the page as JSON in a <script id="__NEXT_DATA__"> tag. The product record sits at props.pageProps.initialData.data.product, and the page is rendered from it. Our e-commerce scraping guide places this source at the top of its data-source hierarchy for exactly this reason, and Walmart is the cleanest example of it:

import json
import os

import requests
from bs4 import BeautifulSoup

API_KEY = os.environ["HASDATA_API_KEY"]

r = requests.post(
    "https://api.hasdata.com/scrape/web",
    headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
    json={"url": "https://www.walmart.com/ip/771229626",
          "proxyType": "residential", "proxyCountry": "US", "outputFormat": ["html"]},
    timeout=120,
)
soup = BeautifulSoup(r.text, "lxml")

state = json.loads(soup.find("script", id="__NEXT_DATA__").text)
product = state["props"]["pageProps"]["initialData"]["data"]["product"]

def count_leaves(node):
    if isinstance(node, dict):
        return sum(count_leaves(v) for v in node.values())
    if isinstance(node, list):
        return sum(count_leaves(v) for v in node)
    return 1

print(product["name"])
print(product["priceInfo"]["currentPrice"]["priceString"],
      "|", product["averageRating"], "stars,", product["numberOfReviews"], "reviews")
print("fields in the product node:", count_leaves(product))

The run prints the product name, price, and rating, then the field count:

Canon EOS Rebel T100 DSLR Camera EF-S 18-55mm f/3.5-5.6 IS II Lens Kit, 18 Megapixel CMOS (APS-C) Sensor, Full HD Videos, Built-in Wi-Fi, Beginner Photographers, Digital Camera, Black
$448.86 | 4.4 stars, 727 reviews
fields in the product node: 438

Across the 110 measured cards the product node carried between 446 and 2,943 leaf fields, median 907. Variants, per-store availability, seller records, fulfillment options, badge and promotion state, review statistics. The five fields the selectors extract are all in there, plus roughly a thousand that never reach the page. One json.loads replaces every selector, and the path breaks only when Walmart changes its data model rather than its styling.

Collecting Product URLs

Product links come from search pages, and search pages are Next.js too. The visible result grid links products under /ip/ paths, and the same __NEXT_DATA__ trick works there as well, but anchor tags are enough for a link collector. The script reads keywords from keywords.txt, fetches one search page per keyword, and keeps every product link plus the related-searches carousel:

import os
import urllib.parse

import requests
from bs4 import BeautifulSoup

API_KEY = os.environ["HASDATA_API_KEY"]

with open("keywords.txt") as f:
    keywords = [line.strip() for line in f if line.strip()]

product_links = []
related_searches = []

for keyword in keywords:
    search_url = "https://www.walmart.com/search?q=" + urllib.parse.quote(keyword)
    r = requests.post(
        "https://api.hasdata.com/scrape/web",
        headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
        json={"url": search_url, "proxyType": "residential", "proxyCountry": "US",
              "outputFormat": ["html"]},
        timeout=120,
    )
    soup = BeautifulSoup(r.text, "lxml")

    for a in soup.find_all("a", href=True):
        href = a["href"].split("?")[0]
        if "/ip/" in href:
            full = "https://www.walmart.com" + href if href.startswith("/") else href
            if full not in product_links:
                product_links.append(full)

    carousel = soup.find("ul", {"data-testid": "carousel-container"})
    if carousel:
        related_searches += [li.get_text(strip=True) for li in carousel.find_all("li")]

with open("links.csv", "w") as f:
    f.write("\n".join(product_links))

print(len(product_links), "product links,", len(related_searches), "related searches")

The related-searches carousel is worth keeping too, since its terms are Walmart’s own suggested keywords and make good seeds for the next run:

Related Searches

Related Searches

A run over our 18 test keywords collected 110+ unique product links, which is exactly the links.csv the complete scraper at the top consumes. For deeper collection, page through search results with the page URL parameter, and expect the same anti-bot wall as on product pages. The collection logic mirrors what we built for Amazon product scraping, where the search-to-product pipeline is covered in more depth.

Walmart Anti-Bot Protection

Walmart’s protection is run by Akamai Bot Manager with HUMAN Security providing the challenge, and the “Press & Hold” page most scrapers meet is HUMAN’s verification widget:

Walmart 'Press & Hold' CAPTCHA

Walmart 'Press & Hold' CAPTCHA

That check was measured rather than assumed, and headers alone do not clear it. A plain requests call with a current Chrome User-Agent:

import requests
from bs4 import BeautifulSoup

headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                          "(KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36"}

r = requests.get("https://www.walmart.com/ip/771229626", headers=headers, timeout=40)
soup = BeautifulSoup(r.text, "lxml")

print(r.status_code, "|", soup.title.get_text(strip=True), "|", len(r.text), "bytes")

The print line reports the status, the page title, and the byte count:

200 | Robot or human? | 15190 bytes

The response is HTTP 200 and a 15 KB “Robot or human?” page instead of the roughly 500 KB product page. The same request through a datacenter proxy returned the same challenge, so the defense reads more than the IP’s country. What it reads is the connection itself: Akamai fingerprints the TLS handshake, and a Python requests client doesn’t shake hands like Chrome no matter what its headers claim. Clients that impersonate a browser’s TLS fingerprint (curl_cffi is the common one in Python) address that layer, and a residential exit addresses the IP reputation layer. In our measurement, a residential proxy with no header tricks at all returned the full product page on 110 of 110 cards. The general blocking-avoidance guide covers the rest of the discipline, rate limits and request patterns included.

Scrape Walmart with the HasData API

Two routes exist and they answer different questions.

The dedicated Walmart endpoints cover search, product and reviews, and return the fields already parsed. One search call for laptop comes back with 52 products, so nothing in this article’s parsing code is needed for that path:

import os

import requests

API_KEY = os.environ["HASDATA_API_KEY"]

response = requests.get(
    "https://api.hasdata.com/scrape/walmart/search/",
    params={"q": "laptop"},
    headers={"x-api-key": API_KEY},
    timeout=120,
)
response.raise_for_status()
products = response.json()["productResults"]
print(len(products), "products")

The Web Scraping API is the other route, and it’s the one this article’s approach needs. It returns the raw page rather than parsed fields, which is what the __NEXT_DATA__ work above reads. Use it when you want the roughly thousand fields in the page state rather than the handful a parsed response carries:

response = requests.post(
    "https://api.hasdata.com/scrape/web",
    headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
    json={
        "url": "https://www.walmart.com/ip/771229626",
        "proxyType": "residential",   # datacenter exits get the challenge page
        "proxyCountry": "US",         # prices and availability differ by region
        "outputFormat": ["html"],     # the response body is the raw page
    },
    timeout=120,
)

html = response.text
print(response.status_code, "|", len(html), "bytes of product page")

proxyType must be residential here, because datacenter exits get the challenge page, and proxyCountry pins the store region, since prices and availability differ by location and a request from a random country returns plausible but wrong numbers. JavaScript rendering isn’t needed, since __NEXT_DATA__ is server-rendered into the initial HTML.

Conclusion

Scraping Walmart in 2026 is one JSON path, not thirty selectors. The page’s __NEXT_DATA__ state carries a median of 907 fields per product where selectors yield four of five and the JSON-LD schema yields nothing anymore, so parse the state and treat the markup as a fallback. The blocking layer is the real work, and it is fingerprint-deep, which is why this guide routes requests through a residential exit instead of collecting header tricks. Point the complete scraper at your own links.csv and the rest is a CSV file filling up.

Valentina Skakun
Valentina Skakun
Valentina is a software engineer who builds data extraction tools before writing about them. With a strong background in Python, she also leverages her experience in JavaScript, PHP, R, and Ruby to reverse-engineer complex web architectures.If data renders in a browser, she will find a way to script its extraction.
Articles

Might Be Interesting