HasData
Back to all posts

How to Scrape Etsy Data with Python

Every Etsy listing page ships a schema.org Product object inside a <script type="application/ld+json"> tag, and it holds the title, price, currency, availability, shop name, rating, review count, materials and image URLs already parsed. To scrape Etsy data in Python, read that object first. On 60 listings drawn from eight keyword queries, it returned the title, price, currency and availability on all 60, while a six-selector CSS route over the same pages returned all six of its fields on none of them.

So that is the order this guide follows. Read the JSON-LD, fall back to selectors for what it leaves out, and open a browser only where neither one reaches. Etsy data feeds competitor pricing and niche research, and it comes off three kinds of page.

What an Etsy Page Carries

Product pages, shop pages and search results carry different slices of the same record.

  1. Product pages hold the fullest version, with name, price, description, images, options, rating and review count, and materials on most listings.
  2. Shop pages hold the seller record plus a first page of that shop’s listings.
  3. Search results pages hold a card per product with name, price, rating and photo.

For a price sweep across a niche, search results are enough. Product pages are the only ones whose markup carries descriptions, materials and per-listing review counts, so you pull those once the sweep has narrowed to a shortlist.

Find these elements the usual way, by inspecting the page in DevTools (F12, or right-click and Inspect) and reading off the CSS selectors.

Etsy listing page showing the title, price, shop name and rating above the image gallery

Etsy product page example

Open the price first. The buy box wraps it in three data attributes that name the region, and those outlast the class names sitting next to them.

DevTools Elements panel on an Etsy listing, with the price text node selected inside a div carrying data-appears-component-name="price", data-selector="price-only" and data-buy-box-region="price", and the price paragraph's own class list visible

The buy box price region in the Elements panel

The same inspection on a shop page gives the seller record and the grid of that shop’s listings.

Etsy shop page showing the shop name, location, rating and sales count above the category sidebar and the featured items grid

Etsy shop page example

Search results drop the description and materials but keep enough for a price sweep.

Etsy search results page for a keyword, showing product cards with title, price, rating and shop name, one of them discounted with the original price struck through

Every field hides behind a different selector on each layout, and several of them exist on only one.

FieldProduct pageShop pageSearch card
Card containern/adiv[data-listing-id]div[data-listing-id]
Titleh1[data-buy-box-listing-title="true"].v2-listing-card__titleh3.wt-text-caption
Price region[data-buy-box-region="price"].n-listing-card__price pp.wt-text-title-01.lc-price
Price amountp.wt-text-title-larger in that regionin the cardspan.currency-value
Original pricespan.wt-text-strikethrough in that regionabsentspan.wt-text-strikethrough
Descriptiondiv[data-id="description-text"]absentabsent
Materials#legacy-materials-product-detailsabsentabsent
Ratinginput[name="initial-rating"]per-review inputs onlyin the card aria label
Review counth2.wt-text-heading-smallabsentabsent
Shop namein the buy boxh1.shop-namein the card
Shop locationabsentp.sb-shop-locationabsent
Sales countabsenta[href*="/sold"]absent
Shop IDabsent[data-shop-id]in the card attributes

Three of those rows will let you down. The review-count selector doesn’t match anything at all on a current listing page. The materials selector reaches only a fraction of the listings that carry the field. And the rating selector hands you a number that belongs to the shop, not to the listing. The measurement below puts figures on all three.

The price row names an attribute rather than a class, and it is worth saying why. A class selector on the price paragraph, p.wt-text-title-larger, resolves on the same 59 of 61 sampled pages as [data-buy-box-region="price"], [data-selector="price-only"] and [data-appears-component-name="price"] do, each with exactly one match, so availability isn’t what separates them. What separates them is that a class describes styling and gets renamed whenever the design system moves, while the attribute names the region. Scoping to the region also pays off directly, since the discounted price and the original sit in it as separate elements. 23 of the 61 listings were discounted, and all 23 gave up both amounts from inside that box.

This is the one field where the structured block comes up short. Its offers.price is always the current price, and the pre-discount figure arrives only when the offer publishes a priceSpecification entry typed as StrikethroughPrice, which happened on 12 of those 23 listings. Where it was there, it matched the page exactly. On the other 11 the original lives in the markup and nowhere else, which makes this the clearest case in the guide for reading a field both ways.

Which Method Fits Which Page

Each route reaches a different subset of the three page types, and the setup cost swings as widely as the coverage does.

RouteProduct pageShop pageSearch results
Requests and BeautifulSoup403 with a 776-byte stubnot checked herenot checked here
Web Scraping API with JS renderingfull page including the JSON-LDnot checked herenot checked here
SeleniumBase in UC Modenot checked herefull page with the listing gridfull page of cards
Official Open APIone listing by IDthe shop record and its listingskeyword search, newest first by default

Cells marked unchecked are simply outside what this guide measured. The ones carrying a result are the routes the code below actually uses.

Fetch a listing with plain requests and you get a 403 and a document of a few hundred bytes, never the page.

import requests

url = "https://www.etsy.com/listing/1447207452/tiny-handmade-resin-bagged-goldfish"
response = requests.get(url, timeout=60)
print(response.status_code, len(response.text))
print(response.text)

That document loads Etsy’s anti-bot script and nothing else.

<html lang="en"><head><title>etsy.com</title><style>#cmsg{animation: A 1.5s;}</style></head>
<body style="margin:0">
<p id="cmsg">Please enable JS and disable any ad blocker</p>
<script data-cfasync="false">var dd={'rt':'c','cid':'AHrlqAAAAAMAU1...','t':'bv','s':45946,
  'host':'geo.captcha-delivery.com','cookie':'ObHdvqjIv4vAcdKeD8vThY9b...'}</script>
<script data-cfasync="false" src="https://ct.captcha-delivery.com/c.js"></script>
</body></html>

Etsy may also show a CAPTCHA when it detects automated traffic, which can interrupt a plain Selenium run. UC Mode reduces how often that happens rather than removing it. If it happens, your options are to slow down and keep request volume low, use Etsy’s official API, or switch to a managed scraping API that returns the page data for you.

Etsy interstitial page asking to enable JavaScript and disable ad blockers

Etsy CAPTCHA

A browser also means you run and maintain Chrome and manage proxies yourself, which is the cost a managed route takes off your hands. Our guide to scraping with Selenium in Python covers that setup. The rest of the code needs requests, beautifulsoup4, lxml and seleniumbase from PyPI, a real Chrome for the browser route, and an API key in HASDATA_API_KEY. The snippets build up one module in the order they appear, so a function defined in one section stays in scope for the next.

Reading the Embedded JSON-LD

The Product object sits right in the page source, so parsing it takes one json.loads call and no selectors at all.

import json
import re


def read_product_ld(page):
    """Returns the schema.org Product object from an Etsy listing page."""
    blocks = re.findall(
        r'<script[^>]+application/ld\+json[^>]*>(.*?)</script>', page, re.S)
    for block in blocks:
        try:
            data = json.loads(block.strip())
        except ValueError:
            continue
        for node in (data if isinstance(data, list) else [data]):
            if isinstance(node, dict) and node.get("@type") == "Product":
                return node
    return None

Listing pages carry a second ld+json block holding the category breadcrumb, which is why the loop checks @type instead of grabbing the first block it finds. Parsing the same block in the browser console shows what the function gets back.

DevTools Console with the parsed ld+json Product object expanded, showing aggregateRating, brand, category, description, an image array, material, name, offers with price and currency, review and sku

The Product object as the page ships it

Every key in that object is a field you would otherwise chase through the markup. The function returns None when a page carries no Product object, and one page in the sample of 61 was exactly that, so the caller has to expect it.

Two fields need a shape check before you touch them. Price and rating nest one level down, under offers and aggregateRating. And image arrives as a list of ImageObject nodes on 59 of the 60 pages, then as a bare URL string on the sixtieth. String values are HTML-escaped, and three of the 60 titles contained &#39; or &quot;, so put html.unescape on every text field, not just the ones that look like they need it.

import html


def image_urls(value):
    """Etsy sends either a list of ImageObject nodes or a single URL string."""
    if isinstance(value, str):
        return [value]
    urls = [i.get("contentURL") if isinstance(i, dict) else i for i in value or []]
    return [u for u in urls if u]


def original_price(offers):
    """The pre-discount price, when the offer publishes one."""
    for spec in offers.get("priceSpecification") or []:
        if str(spec.get("priceType", "")).endswith("StrikethroughPrice"):
            return spec.get("price")
    return None


def flatten(product):
    """Pulls the fields worth keeping out of the Product object."""
    if not product:
        return None
    offers = product.get("offers") or {}
    if isinstance(offers, list):
        offers = offers[0] if offers else {}
    rating = product.get("aggregateRating") or {}
    brand = product.get("brand") or {}
    return {
        "title": html.unescape(product.get("name") or ""),
        "price": offers.get("price"),
        "original_price": original_price(offers),
        "currency": offers.get("priceCurrency"),
        "availability": (offers.get("availability") or "").rsplit("/", 1)[-1],
        "rating": rating.get("ratingValue"),
        "reviews": rating.get("reviewCount"),
        "shop": brand.get("name") if isinstance(brand, dict) else brand,
        "materials": product.get("material"),
        "description": html.unescape(product.get("description") or ""),
        "sku": product.get("sku"),
        "images": image_urls(product.get("image")),
    }

The sku field repeats the numeric listing ID from the URL, which makes it a handy join key when you collect the same products again later. Image lists ran from 1 to 18 URLs across the sample, 9 at the median. The Walmart guide applies the same technique to a marketplace whose structured block carries a different field set.

What CSS Selectors and JSON-LD Return

On every page in the sample the structured block reached more fields than the selectors did, and price, the one field both routes return, agrees to the cent. Both routes ran over the same 61 listing pages, gathered from eight keyword queries through Google search, of which 60 carried a Product object. One column applies six CSS extractions from the table above, the other applies flatten. A field counts as returned only when the code hands back a value, not when the element merely exists. The pages are the ones that rendered in full, so the sample is retrievable listings rather than a random draw of Etsy.

FieldCSS selectorsJSON-LD
Title59 / 6060 / 60
Price59 / 6060 / 60
Currencynot reachable60 / 60
Availabilitynot reachable60 / 60
Description59 / 6060 / 60
Shop namenot reachable60 / 60
Rating59 / 6059 / 60
Review count0 / 6059 / 60
Materials10 / 6048 / 60
Image URLsnot reachable60 / 60
All six CSS fields on one page0 / 60n/a

Prices matched on all 59 pages where both routes produced one, after stripping currency symbols and the + that marks a from-price, so the gap between the columns is coverage rather than correctness.

The review count selector, h2.wt-text-heading-small, matched nothing on any of the 60 pages. Neither did the three other element shapes that would plausibly carry the number, a link to the reviews anchor, a button labelled with it, or a span holding a review total, so the count now reaches a scraper only through the structured block. The materials selector, #legacy-materials-product-details, still resolves, but on 10 pages against the 48 that declare a material value, which leaves 38 listings carrying the field where the selector cannot see it.

Rating is the field you shouldn’t trust on either route. Pages carry up to 12 inputs named initial-rating, 4 at the median, because each individual review renders one of them, and one page in the sample carried none. soup.find returns the first, and on 58 of the 59 pages that have one, that first input sits inside a shop-scoped container, which makes it the seller’s rating rather than the listing’s. It matches the listing’s rating on 38 of the 59, and on the 21 where it does not, 9 pages carry the right number in some other input while 12 have it in no input at all. No fixed index gets you out of it, and a script that labels that value “product rating” is wrong on more than a third of the sample.

Shop pages went the same way, checked separately in a browser session. Of six selectors for the seller record, only the shop logo still resolves. The name, the location and the sales count have working replacements in the table above, while the shop title and the announcement have none in the markup and come back from the official API instead.

Scraping an Etsy Product Page

A product-page pass hands you the whole record for one listing ID, which is what a shortlist off a search sweep feeds on. It needs the page rendered, because the JSON-LD isn’t in the initial response, and it needs to arrive from an exit the site will serve.

HasData’s web scraping API renders the page from a residential exit and returns the HTML. Sign up to get a personal API key, then read it from the environment rather than pasting it into the script.

HasData dashboard on the API Keys page, with the Default key row highlighted

API Keys under Workspace

Resource blocking normally speeds a page up by skipping images and scripts. Etsy needs its own scripts to render, so blockResources and blockAds both stay off here. Four parameter sets on one listing put the difference at an 18 KB stub with them on against a 508 KB page with two ld+json blocks with them off.

import json
import os
import time

import requests

API = "https://api.hasdata.com/scrape/web"
KEY = os.environ["HASDATA_API_KEY"]


def fetch(url, attempts=4):
    payload = {
        "url": url,
        "proxyType": "residential",
        "proxyCountry": "US",
        "blockResources": False,
        "blockAds": False,
        "jsRendering": True,
    }
    for attempt in range(attempts):
        response = requests.post(
            API, headers={"Content-Type": "application/json", "x-api-key": KEY},
            data=json.dumps(payload), timeout=240)
        if response.status_code == 429:
            time.sleep(15)
            continue
        body = response.json()
        status = body.get("requestMetadata", {}).get("status")
        if status == "ok":
            return body.get("content", "")
        time.sleep(6 * (attempt + 1))
    raise RuntimeError(f"no page after {attempts} attempts")


product_url = "https://www.etsy.com/listing/1447207452/tiny-handmade-resin-bagged-goldfish"
record = flatten(read_product_ld(fetch(product_url)))
print(json.dumps(record, indent=2, ensure_ascii=False))

A 429 doesn’t mean the request was refused, only that the account’s concurrency is busy, which is why the loop waits and repeats instead of dropping through to the backoff. The record comes back with its fields typed and named.

{
  "title": "Tiny Handmade Resin Bagged Goldfish Figurine | Kawaii Paperweight | ...",
  "price": "15.00",
  "original_price": null,
  "currency": "USD",
  "availability": "InStock",
  "rating": "4.9",
  "reviews": 461,
  "shop": "GeekyGazelle",
  "materials": "Resin",
  "sku": "1447207452"
}

The full record also carries description and the image list, both trimmed out of the block above. Options and shipping estimates never make it into the structured block at all. To reach those you build a soup from the same page string and add one select_one call per field.

Scraping a Shop Page

Shop pages need two records, one for the seller and one per listing in the grid. There’s no structured block on them, so the seller half is selectors the whole way. A shop pass is also a low-volume job, a handful of requests per shop against thousands of listing pages, so it runs in a browser here rather than through managed infrastructure.

SeleniumBase in UC Mode drives Chrome through a patched driver based on undetected-chromedriver, and uc_open_with_reconnect reconnects the automation channel after the page has settled, which is what gets a shop page to render. The selectors below are the ones that resolve on the current layout. h1.wt-text-heading, h2.wt-text-caption, div.shop-location, div.shop-sales-reviews a and p#rmjs-1 return nothing on it.

from bs4 import BeautifulSoup
from seleniumbase import SB


def read_shop(html):
    soup = BeautifulSoup(html, "lxml")

    def text(selector):
        el = soup.select_one(selector)
        return el.get_text(" ", strip=True) if el else None

    holder = soup.select_one("[data-shop-id]")
    logo = soup.select_one("div.shop-icon img")
    shop = {
        "shop_id": holder["data-shop-id"] if holder else None,
        "name": text("h1.shop-name"),
        "location": text("p.sb-shop-location"),
        "sales": (text('a[href*="/sold"]') or "").split(" ")[0] or None,
        "logo": logo["src"] if logo else None,
    }

    products, seen = [], set()
    for card in soup.select("div[data-listing-id]"):
        listing_id = card["data-listing-id"]
        title = card.select_one(".v2-listing-card__title")
        price = card.select_one(".n-listing-card__price p")
        if listing_id in seen or not title:
            continue
        seen.add(listing_id)
        products.append({
            "listing_id": listing_id,
            "title": title.get_text(strip=True),
            "price": price.get_text(strip=True) if price else None,
        })
    return shop, products


with SB(uc=True) as sb:
    sb.uc_open_with_reconnect("https://www.etsy.com/shop/GeekyGazelle", 6)
    sb.sleep(5)
    shop, products = read_shop(sb.get_page_source())

print(shop["name"], shop["location"], shop["sales"], len(products), "listings on page 1")

The run prints the seller record and the size of the first grid.

GeekyGazelle Kentucky, United States 4728 39 listings on page 1

Both filters in that loop earn their keep. On the run above, div[data-listing-id] matched 41 nodes for 39 listings, because one card appears twice in the markup and one node carries no title element. The seen set and the title check drop both without a special case.

Only the first page of the grid arrives this way. For a shop’s full catalogue you either walk ?page=2 and onward until a request comes back with no cards, or you ask the official API for that shop’s listings, whose response carries a total count so the walk knows where to stop.

Scraping Etsy Search Results

Search results carry the heaviest bot checks of the three page types, and the same UC Mode session handles them. The filters Etsy exposes in its own UI are all URL parameters, so the query is assembled rather than clicked.

from urllib.parse import urlencode

params = {
    "q": "resin art",
    "free_shipping": "true",
    "is_discounted": "true",
    "min": 50,
    "max": 100,
}


def read_cards(html):
    soup = BeautifulSoup(html, "lxml")
    cards, seen = [], set()
    for card in soup.select("div[data-listing-id]"):
        listing_id = card["data-listing-id"]
        title = card.select_one("h3.wt-text-caption")
        if listing_id in seen or not title:
            continue
        seen.add(listing_id)
        amount = card.select_one("span.currency-value")
        was = card.select_one("span.wt-text-strikethrough")
        cards.append({
            "listing_id": listing_id,
            "url": f"https://www.etsy.com/listing/{listing_id}/",
            "title": title.get_text(strip=True),
            "price": amount.get_text(strip=True) if amount else None,
            "currency": (lambda e: e.get_text(strip=True) if e else None)(
                card.select_one("span.currency-symbol")),
            "original_price": was.get_text(strip=True) if was else None,
        })
    return cards


with SB(uc=True) as sb:
    sb.uc_open_with_reconnect(f"https://www.etsy.com/search?{urlencode(params)}", 6)
    sb.sleep(5)
    cards = read_cards(sb.get_page_source())

print(len(cards), "cards on page 1")

Duplication runs heavier here than on a shop page. That query matched div[data-listing-id] 129 times for 60 distinct listings, so most of what the selector hands back is the same cards rendered again for other viewport widths. Building each listing URL from the ID is worth the extra line, because those IDs are what a product-page pass and the official API’s batch endpoint both want.

Don’t read the price element’s text on a card. It holds a screen-reader label as well as the visible amount, so a discounted card’s text arrives with the same figure twice, and 36 of the 57 cards on that page read that way. The amount has an element of its own, span.currency-value, which carried a value on all 57, with the symbol beside it in span.currency-symbol. The pre-discount price is a third element, span.wt-text-strikethrough, present on the same 36 cards. So a card gives you the sale price and the reference price with no string splitting at all.

Results shift with the region the browser appears to be in, and ranking is only part of it. A session running from outside the US comes back with prices in the local currency and a localized page title, so if you’re comparing anything, pin the locale with a proxy. Like Selenium, SeleniumBase supports headless mode and user-agent management.

Etsy’s Official Open API

Etsy runs a documented REST API whose endpoints answer on https://openapi.etsy.com/v3/ and https://api.etsy.com/v3/, which its request standards describe as equivalent. A set of those endpoints needs nothing but an app key in an x-api-key header, sent as keystring:shared_secret, and registering an app in the developer portal produces both halves.

EndpointWhat it returns
/v3/application/listings/activekeyword search with price, taxonomy and shop-location filters, 100 per page
/v3/application/listings/{listing_id}one listing, optionally with Shop, Images, User or BuyerPrice attached
/v3/application/listings/batchup to 100 listings by ID in one call
/v3/application/listings/{listing_id}/reviewsreviews for a listing
/v3/application/shops/{shop_id}/listings/activea shop’s active listings with a total count
/v3/application/shops/{shop_id}one shop record, announcement and review_average included

Rate limits apply per API key as queries per second and queries per day, and the documented daily limit runs on a sliding 24-hour window rather than resetting at midnight. Every response reports the state of both in x-limit-per-second, x-remaining-this-second, x-limit-per-day and x-remaining-today, which are the numbers to log if a collection job needs to pace itself.

Two things stop the official route short of what the pages give you. Its keyword search sorts by creation date by default, and its nearest thing to relevance, sort_on=score, is Etsy’s own scoring rather than the ordering the public search page shows, so it answers a different question than a rank survey asks. Access beyond your own shop is also tiered. A Seller App is scoped to the shop that owns it, while a Personal App and then Commercial Access go through review before an app may serve other sellers. Among the criteria Etsy lists for that review is that an application must not sidestep the API to retrieve Etsy data, and the API Terms of Use set caching rules on top of it, so an app built on the official route commits to staying on it.

Etsy’s Quick Start tutorial is written in Node.js, so the request shapes transfer to Python’s HTTP clients unchanged while the code does not.

Saving the Results

CSV suits a flat sheet of listings, and JSON keeps the nested fields such as the image list intact. Both take the records you collected above exactly as they are.

import csv
import json

records = [record]

with open("etsy_products.json", "w", encoding="utf-8") as f:
    json.dump(records, f, ensure_ascii=False, indent=2)

fields = ["sku", "title", "price", "original_price", "currency", "availability",
          "rating", "reviews", "shop", "materials"]
with open("etsy_products.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=fields, extrasaction="ignore")
    writer.writeheader()
    writer.writerows(records)

extrasaction="ignore" lets the same record dictionaries feed both writers, so the description and the image list land in the JSON file and are dropped from the CSV without a second pass. A database becomes the better target once collection repeats on a schedule and rows need comparing over time, which is the setup the price scraping guide walks through.

Conclusion

Read the structured block first, and keep selectors for what it leaves out: options, shipping estimates, and the whole seller record. Don’t trust rating on either route. The input that looks like it carries the listing’s number is the shop’s on almost every page, and on 12 of 59 the listing’s own value isn’t in any input at all. For a price survey, search cards will answer the question until the shortlist is short enough to justify product pages, and it’s the listing IDs that carry the work from one stage to the next.

Valentina Skakun
Valentina Skakun
Valentina is a software engineer who builds data extraction tools before writing about them. With a strong background in Python, she also leverages her experience in JavaScript, PHP, R, and Ruby to reverse-engineer complex web architectures.If data renders in a browser, she will find a way to script its extraction.
Articles

Might Be Interesting