Playwright is a browser automation library with a full Python API, and for Python web scraping it is now the default choice. On PyPI it pulled 91.4 million downloads last month against Selenium’s 37.4 million, and Pyppeteer, the library it used to be compared with, is down to 0.88 million and dormant. This guide covers scraping with Playwright in Python end to end, from the first script to pagination, user agents, and a complete saved-to-disk example.
Why Playwright?
Besides letting you launch a browser and mimic real user behavior, Playwright has a few things that set it apart:
- Multiple browsers from one install.
playwright installbrings its own Chromium, Firefox, and WebKit builds, so there’s no driver to match against a local browser version. - Dynamic content. It runs a real browser, so pages that build their content with JavaScript read the same as static ones.
- Auto-waiting locators. A locator waits for its element before acting on it, which removes most of the manual
Nonechecks and sleeps that scraping code accumulates. - Codegen support. Playwright can record your actions on a page and turn them into Python code you can reuse for scraping.
The Python package offers the same API in async and sync flavors. This guide uses the async one, since parallel page loads are where a scraper feels the difference. For how Playwright compares with Selenium and Pyppeteer, our headless browser overview puts them side by side.
Requirements
Python itself comes first, and our guide covers setting it up if needed. With that done, install Playwright:
pip install playwrightThis installs just the library, no browsers included. To get the browsers, run:
playwright installThat downloads Chromium, Firefox, and WebKit. You only need to do this once after installing the library.
To see all available commands, run:
playwright helpThere you’ll find things like how to take a screenshot right from the terminal.
If you’re planning to use NodeJS instead of Python, we have a separate guide on Playwright with NodeJS.
Minimal Working Example (MWE)
If you just want the smallest Playwright Python script that scrapes a headline and text from a page, here it is:
import asyncio
from playwright.async_api import async_playwright
async def run():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto("https://example.com")
print("H1:", await page.locator("h1").inner_text())
for i, para in enumerate(await page.locator("p").all(), 1):
print(f"Paragraph {i}:", await para.inner_text())
await browser.close()
asyncio.run(run())It prints the page’s heading and both of its paragraphs. Note the headless setting in the launch call. With True the browser runs in the background without opening a window.
This API is async, so the script runs inside asyncio. A sync_api twin exists with the same method names if a plain script fits your project better.
Basic Scraping with Playwright
Let’s walk through a hands-on example and scrape product data from books.toscrape.com, a catalogue built for scraping practice. Each book card carries a title, a price, a stock flag, a cover image, and a rating:

The plan is to read all of that, card by card, and the first step is telling Playwright where each piece lives.
Locate Elements
Playwright’s current API for finding things is page.locator(), which takes a CSS selector or an XPath and represents the element without fetching it yet. One method covers both the single and the multiple case: .first or .nth() picks one match, .all() returns them as a list, and .count() says how many there are. Because a locator waits for its element before acting, the extraction code below has no manual existence checks at all.
The older query_selector calls still work, but they return None on a miss and put that check on every line of your code, which is why new Playwright code and this guide use locators.
We already compared CSS and XPath in another post, so if you’re unsure which to use, check that out. Open DevTools (F12 or right-click and Inspect) and pick out the selectors. For the catalogue above they come to:
| Element | CSS Selector | XPath Expression |
|---|---|---|
| Product card | article.product_pod | //article[@class=“product_pod”] |
| Title and link | h3 a | .//h3/a |
| Price | .price_color | .//p[@class=“price_color”] |
| Availability | .availability | .//p[contains(@class, “availability”)] |
| Cover image | .image_container img | .//div[@class=“image_container”]//img |
| Rating | p.star-rating | .//p[contains(@class, “star-rating”)] |
Every card is an article.product_pod container, so the plan is the usual one. Collect the containers, then read each field inside them.
Scrape Titles and Text
Getting text out of a located element is one call:
product = page.locator("article.product_pod").first
price = await product.locator(".price_color").inner_text()
availability = (await product.locator(".availability").inner_text()).strip()The title on this page is a small trap worth knowing. The link’s visible text is truncated (“A Light in the …” in the screenshot above), while the full name is in the title attribute, so attribute extraction is the right tool:
link_el = product.locator("h3 a")
title = await link_el.get_attribute("title")
link = await link_el.get_attribute("href")That returns the full “A Light in the Attic” with the price £51.77 and the In stock flag. Attributes also carry data that never renders as text. The rating here is encoded in a class name, star-rating Three, which get_attribute("class") reads directly.
Scrape Images
An image is one more attribute read, with one catch:
img_url = await product.locator(".image_container img").get_attribute("src")The catch is that src came back relative on this page, media/cache/... rather than a full URL. Join it against the page address with urllib.parse.urljoin before requesting it, then write the response bytes to a file with the right extension to save the picture.
For Multiple Elements
.all() turns the same extraction into a loop over every card on the page:
for product in await page.locator("article.product_pod").all():
link_el = product.locator("h3 a")
print(await link_el.get_attribute("title"))Today’s front page yields 20 products this way. For an element that may genuinely be absent, ask count() first instead of resurrecting the None checks:
if await product.locator(".promo-banner").count():
...The rest of the extraction stays the same.
Advanced Scraping Techniques
The basics cover a static catalogue. The techniques below cover forms, pagination, and identity checks.
Click Buttons and Text Input
page.fill() types into a field and page.click() presses things. The quotes sandbox has a login form that takes any credentials, which makes it a safe place to try the pattern:
await page.goto("https://quotes.toscrape.com/login")
await page.fill("input#username", "reader")
await page.fill("input#password", "example")
await page.click("input[type='submit']")After the click, the page carries a Logout link, which is how the run verified the login actually happened. On a real site only the selectors and values change.
Pagination and Infinite Scrolling
Pagination lets you move between pages of content, and the catalogue’s next button is li.next a. Target a “next” element rather than page numbers, and let count() end the loop:
while True:
await page.wait_for_selector("article.product_pod")
# ... extract the current page here ...
next_link = page.locator("li.next a")
if await next_link.count():
await next_link.click()
else:
breakThree pages of the catalogue collect 60 titles. Waiting for a selector beats sleeping for a guessed number of seconds, since the loop continues the moment the products exist.
Infinite scroll needs a different exit condition, because there’s no button. Scroll to the bottom, wait, and stop when the page height stops growing:
await page.goto("https://quotes.toscrape.com/scroll")
previous_height = await page.evaluate("document.body.scrollHeight")
while True:
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
await page.wait_for_timeout(1000)
new_height = await page.evaluate("document.body.scrollHeight")
if new_height == previous_height:
break
previous_height = new_heightOn the scrolling demo this loaded the full set of 100 quotes before the height went stable. page.evaluate() also runs any other JavaScript you need inside the page.
Add User Agents
Set the User Agent on the context, after launching the browser and before opening pages:
browser = await p.chromium.launch(headless=True)
user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36"
context = await browser.new_context(user_agent=user_agent)
page = await context.new_page()
await page.goto("https://books.toscrape.com/")You can use a real User Agent from your own browser or take one of the current ones from our auto-updated table. The string changes what the server reads in the header, and a server that inspects client hints can still recognize a headless browser, so treat the swap as one layer rather than a disguise.
Handling Errors
To keep the script alive when a page misbehaves, wrap the scraping in a try block and log what went wrong:
try:
# Your code here
...
except Exception as e:
print(f"Unexpected error: {e}")We covered the common failure types in the article about retrying requests in Python, so this guide keeps the catch-all. A screenshot taken in the same except block tells you what the page looked like when it failed.
except Exception as e:
await page.screenshot(path="error.png")A picture of the page at the moment of failure answers “what was actually there” faster than any log line.
Screenshot Capture of Web Pages
Screenshots are one call once the page is loaded, and a PDF export works the same way in headless Chromium:
await page.screenshot(path="screenshots/test.png")
await page.pdf(path="pdf/page.pdf")To avoid overwriting files, stamp the name with the current time:
from datetime import datetime
timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
await page.screenshot(path=f"screenshots/screenshot_{timestamp}.png")Each capture gets a unique name, and the filename itself records when it was taken, which turns a screenshots folder into a visual history of what you scraped.
Save Scraped Data
The collected rows are written as JSON for structure and CSV for spreadsheets:
import json
import csv
with open(json_filename, "w", encoding="utf-8") as f:
json.dump(scraped_data, f, indent=2, ensure_ascii=False)
with open(csv_filename, "w", encoding="utf-8", newline="") as f:
writer = csv.DictWriter(f, fieldnames=scraped_data[0].keys())
writer.writeheader()
writer.writerows(scraped_data)If you’re building a price tracker, append timestamped entries into one file instead of creating a new file per run, and the history stays queryable.
Full Example
Here’s everything put together, the version that saves 20 products:
import asyncio
import csv
import json
from datetime import datetime
from pathlib import Path
from playwright.async_api import async_playwright
async def run():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
user_agent = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36"
)
context = await browser.new_context(user_agent=user_agent)
page = await context.new_page()
Path("screenshots").mkdir(exist_ok=True)
Path("json").mkdir(exist_ok=True)
Path("csv").mkdir(exist_ok=True)
scraped_data = []
try:
await page.goto("https://books.toscrape.com/")
await page.wait_for_selector("article.product_pod")
for product in await page.locator("article.product_pod").all():
link_el = product.locator("h3 a")
scraped_data.append({
"title": await link_el.get_attribute("title"),
"link": await link_el.get_attribute("href"),
"price": await product.locator(".price_color").inner_text(),
"availability": (await product.locator(".availability").inner_text()).strip(),
"image": await product.locator(".image_container img").get_attribute("src"),
})
timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
await page.screenshot(path=f"screenshots/screenshot_{timestamp}.png")
except Exception as e:
print(f"Unexpected error: {e}")
finally:
await browser.close()
if scraped_data:
timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M-%S")
json_filename = f"json/products_{timestamp}.json"
csv_filename = f"csv/products_{timestamp}.csv"
with open(json_filename, "w", encoding="utf-8") as f:
json.dump(scraped_data, f, indent=2, ensure_ascii=False)
with open(csv_filename, "w", encoding="utf-8", newline="") as f:
writer = csv.DictWriter(f, fieldnames=scraped_data[0].keys())
writer.writeheader()
writer.writerows(scraped_data)
print(f"Saved {len(scraped_data)} products to {json_filename} and {csv_filename}")
asyncio.run(run())It collects every product card on the page into dictionaries, saves a screenshot, and writes the data to both JSON and CSV. When something breaks, the script prints the error instead of crashing, and the browser closes either way.
Useful Commands
Two Playwright tools run outside the scripts.
Record Actions Into Code
Playwright can write the code for you. Run this in your terminal:
playwright codegen https://example.comIt opens a browser window and loads the page you specified:

As you click around, Playwright generates the code for those actions in real time:

Testers use codegen a lot, since it automates UI interactions quickly. Scraping developers mostly write the selectors themselves, and the inspector still earns a look on a page whose structure resists reading, because it names the locator Playwright would pick for an element you click.
Manage Browser
To install a specific browser instead of all three, name it:
playwright install chromiumAnd playwright uninstall removes them, with the same optional browser argument. The builds belong to the library rather than to your system browser, so removing them leaves that browser alone.
Conclusion
Playwright covers Python scraping from a one-liner page read to a paginated catalogue crawl, and the locator API is what keeps the code short. Elements wait for themselves, and the manual None checks that query_selector code accumulates never enter the script. Launch, locate, extract, save. Those four steps carry every technique above.
If the page you’re after guards itself more closely than these demos, Scrapy users can run Playwright inside spiders through scrapy-playwright, and when the blocking itself becomes the work, our Web Scraping API renders the JavaScript, routes the request through its proxy pool, and returns the HTML for your parser.


