HasData
Back to all posts

Pyppeteer: The Puppeteer Alternative for Python Web Scraping

Pyppeteer is a Python port of the Puppeteer library from Node.js, and it’s unmaintained. The last release, 2.0.0, shipped in February 2024, and the project has been quiet since. For a new scraper, start with Playwright instead, which covers the same ground with a maintained API. This guide stays useful for the code that already exists, so it covers installing Pyppeteer on a current system (which takes one extra step now), complete runnable examples for every common task, and a migration map for moving the code to Playwright when the time comes.

What is Pyppeteer

Pyppeteer drives a Chromium browser from Python and simulates the actions of a real user. It arrived as a port of Puppeteer, so its method names are Puppeteer’s camelCase (newPage, waitForSelector) and its model is fully asynchronous. The feature set covers page management, selectors and XPath, JavaScript execution in the page’s context, and screenshots.

Getting Started with Pyppeteer

Pyppeteer 2.0.0 declares Python 3.8 or higher and below 4.0, and any code editor will do. The installation has two traps on a current system.

How to install Pyppeteer

The package itself installs normally:

pip install pyppeteer

The first trap is Chromium. Pyppeteer tries to download its own build on first launch, and that download fails on current systems, so point executablePath at a browser you already have. Any Chromium works, including the one Playwright installs (on Windows that is under %LOCALAPPDATA%\ms-playwright\). The examples below pass it explicitly.

The second trap is easier to miss. Installing pyppeteer downgrades the websockets package to a 2022 release, which can break other libraries living in the same environment. Give Pyppeteer projects their own virtual environment, or expect to re-upgrade websockets after working with it.

Pyppeteer supported browsers

Pyppeteer controls Chromium only. Firefox, WebKit, and everything else stay out of reach, which is one of the concrete reasons the migration section at the end of this guide exists.

Basic Pyppeteer Project

Here’s the smallest complete script, and two details in it matter. The import is lowercase pyppeteer (a capitalized import fails with ModuleNotFoundError), and the entry point is asyncio.run(), since the older get_event_loop() pattern is deprecated in current Python:

import asyncio
from pyppeteer import launch

CHROME = r"C:\path\to\chrome.exe"  # any Chromium build

async def main():
    browser = await launch(headless=True, executablePath=CHROME)
    page = await browser.newPage()
    await page.goto("https://books.toscrape.com/")
    print("Title:", await page.title())
    await browser.close()

asyncio.run(main())

It prints:

Title: All products | Books to Scrape - Sandbox

Every script below keeps this shape and only the body of main() changes.

Advanced Configuration with Pyppeteer

A proxy, a user agent and cookies come up in almost every scraper, and one script covers all three:

import asyncio
from pyppeteer import launch

CHROME = r"C:\path\to\chrome.exe"
UA = ("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
      "(KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36")

async def main():
    browser = await launch(
        headless=True,
        executablePath=CHROME,
        # args=[f"--proxy-server=http://proxy_ip:proxy_port"],
    )
    page = await browser.newPage()
    await page.setUserAgent(UA)
    await page.setCookie({
        "name": "example_cookie",
        "value": "123456789",
        "url": "https://httpbin.org",
    })
    await page.goto("https://httpbin.org/headers")
    print(await page.evaluate("document.body.innerText"))
    await browser.close()

asyncio.run(main())

The run’s output showed the Chrome/152 string echoed back in the User-Agent header, which confirms the override reached the server.

Using proxies

The proxy travels as a Chromium argument in launch(), the commented line in the script above. Uncomment it and fill in the address. For authenticated proxies the format is username:password@ip:port. We covered how proxies work in Python and where rotating ones come from separately.

Modifying user agents

setUserAgent() applies per page, after newPage() and before goto(). Take a current string from our auto-updated user agent table rather than writing one by hand. Rotating through a small list of them per page lowers the profile of a long run.

Managing cookies

setCookie() also applies per page, and the dictionary needs a url (or domain) key in addition to name and value, or Chromium rejects it. The full parameter list is in the API reference the project publishes.

One script demonstrates the whole extraction toolbox, the page’s HTML, a CSS selector, an XPath query, and the wait that makes dynamic pages readable:

import asyncio
from pyppeteer import launch

CHROME = r"C:\path\to\chrome.exe"

async def main():
    browser = await launch(headless=True, executablePath=CHROME)
    page = await browser.newPage()
    await page.goto(
        "https://books.toscrape.com/",
        {"timeout": 15000, "waitUntil": "load"},
    )
    await page.waitForSelector("article.product_pod")

    content = await page.content()
    print("HTML length:", len(content))

    heading = await page.querySelector("h1")
    print("H1:", await page.evaluate("(e) => e.textContent", heading))

    links = await page.xpath("//article[@class='product_pod']//h3/a")
    print("products:", len(links))
    print("first:", await page.evaluate("(e) => e.getAttribute('title')", links[0]))

    await browser.close()

asyncio.run(main())

That returns the full page, around 51,000 characters of HTML, plus the heading and 20 product links with “A Light in the Attic” first.

Get HTML content

page.content() returns the rendered HTML, after JavaScript ran, which is the whole reason to pay for a browser instead of calling requests.

Using XPath and CSS selectors for Data Scraping

querySelector() takes a CSS selector and returns one element handle, querySelectorAll() returns them all, and xpath() always returns a list, so the element comes out by index. A handle isn’t text yet. Reading it goes through page.evaluate() with the element passed in, as both lookups in the script show. Which query language to prefer is its own topic, covered in the CSS selectors cheat sheet and the XPath guide.

Dealing with timeouts

The timeout option on goto() caps how long navigation may take, in milliseconds, and the script above allows 15 seconds before a timeout error is raised.

Implementing waits for page loading

waitUntil decides what counts as “loaded”: load waits for the load event, domcontentloaded settles for the DOM. For scraping, the stronger signal is usually an element you actually need.

Scrape Dynamic Pages

That’s what waitForSelector() does in the script. It continues the moment the element exists, however long that took. On the catalogue used above the products ship in the server’s HTML, so the wait returns immediately and costs nothing. On a page that builds its list in the browser, the same line is what stops the script from reading an empty DOM.

Interacting with the Page

Forms, clicks, scrolling, and screenshots combine into the login flow below, which runs against a sandbox login form that accepts any credentials:

import asyncio
from pyppeteer import launch

CHROME = r"C:\path\to\chrome.exe"

async def main():
    browser = await launch(headless=True, executablePath=CHROME)
    page = await browser.newPage()
    await page.goto("https://quotes.toscrape.com/login")

    await page.type("input#username", "reader")
    await page.type("input#password", "example")

    # a click that triggers navigation must be awaited together with it
    await asyncio.gather(
        page.waitForNavigation(),
        page.click("input[type='submit']"),
    )

    logged_in = await page.querySelector("a[href='/logout']")
    print("logged in:", logged_in is not None)

    await page.evaluate("window.scrollBy(0, 300)")
    await page.screenshot({"path": "screenshot.png"})
    await browser.close()

asyncio.run(main())

The run printed logged in: True and saved the screenshot.

Clicking buttons and elements

click() takes a selector. The nuance the script encodes is navigation. When a click loads a new page, click() and waitForNavigation() must run together through asyncio.gather(), or the script races the page and reads the old one.

Fill an input field

type() fills a field character by character, like a person would. Reading a field back goes through evaluate():

value = await page.evaluate('document.querySelector("input#username").value')

Reading fields back covers the case of values a page generates on its own.

Executing specific actions

page.evaluate() runs any JavaScript in the page’s context, which covers everything without a dedicated method. Two dedicated ones worth knowing are page.keyboard.press("Enter"), which submits most login forms without hunting for the button, and page.hover(".selector") for menus that open on hover.

Login with Pyppeteer

The full flow is the script above. Type into both fields, click together with waitForNavigation(), and confirm the login by looking for an element that only exists in the signed-in state, like the logout link.

Page Scrolling

Scrolling is JavaScript through evaluate(), and window.scrollBy(0, 300) in the script moves the viewport down 300 pixels. Repeating it in a loop until the page height stops changing is the standard infinite-scroll pattern.

Capturing screenshots

screenshot() takes a dictionary with a path, and the extension picks the format (PNG or JPEG). The folder must already exist, or the call fails.

Troubleshooting and Error Handling

The errors people actually hit with Pyppeteer in its unmaintained state cluster around the browser, not the code.

Common issues like “Pyppeteer is not installed”

A capitalized import is the first thing to check, since from Pyppeteer import launch fails with ModuleNotFoundError while the package sits installed right there. The module is lowercase pyppeteer. Beyond that, the message usually means the installation itself ended with an error, and reinstalling with pip install pyppeteer in a clean virtual environment settles it.

Handling unexpected browser closures in Pyppeteer

A browser that dies at launch on a current system is almost always the Chromium download problem from the installation section, and executablePath pointed at an existing build is the fix. pyppeteer-install re-attempts the bundled download, which no longer succeeds on current systems. For crashes mid-run, wrap the session in try/finally so browser.close() runs either way.

Migrating Pyppeteer Code to Playwright

Since the project is dormant, every Pyppeteer script is one dependency freeze away from being legacy code, and the move to Playwright is mostly mechanical. The model is the same, an async browser driven from Python, and the method names translate one to one:

PyppeteerPlaywright (async)
from pyppeteer import launchfrom playwright.async_api import async_playwright
browser = await launch()browser = await p.chromium.launch()
page = await browser.newPage()page = await browser.new_page()
await page.goto(url)await page.goto(url)
await page.content()await page.content()
await page.querySelector("h1")page.locator("h1")
await page.xpath("//h1")page.locator("//h1")
await page.waitForSelector(".x")await page.wait_for_selector(".x")
await page.setUserAgent(ua)await browser.new_context(user_agent=ua)
await page.setCookie(c)await context.add_cookies([c])
await page.type(sel, text)await page.fill(sel, text)
await page.evaluate(js)await page.evaluate(js)
await page.screenshot({"path": p})await page.screenshot(path=p)

Three habits change along the way. Locators replace element handles, so reading text no longer goes through evaluate() and becomes await page.locator("h1").inner_text(). Waiting mostly disappears, because locators wait for their elements on their own. And the browser download works again, since playwright install chromium fetches a current build instead of failing. The basic example above becomes this in Playwright:

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto("https://books.toscrape.com/")
        print("Title:", await page.title())
        await browser.close()

asyncio.run(main())

Both versions of that script print the same title, and the Playwright guide continues from here.

Pyppeteer vs. Other Tools

The comparisons below assume you’re choosing for existing code or constrained environments, since for new projects the maintenance status already decides.

Pyppeteer vs. BeautifulSoup

The two occupy opposite ends of the stack:

FeaturePyppeteerBeautifulSoup
PurposeAutomation of browser interactions, dynamic content execution, and user actions.Parsing HTML and XML documents, extracting information from static markup.
Ease of UseMore involved, a browser instance and asynchronous code to manage.A simpler, declarative approach to HTML parsing.
PerformanceSlower, a browser renders every page.Faster for static page parsing.
Use CasesDynamic web pages, data that appears after JavaScript runs, screenshots.Data extraction from static HTML, navigating and searching through markup.

For static pages, BeautifulSoup with requests is the lighter tool. A browser is worth its memory and startup time only where the content is built by JavaScript.

Pyppeteer vs. Scrapy

Scrapy competes on crawl scale rather than rendering:

FeaturePyppeteerScrapy
PurposeWeb scraping with browser automationGeneral-purpose web crawling and scraping
Browser AutomationYesNo (focused on HTTP requests)
Ease of UseMore complex due to browser integrationEasier for its intended use
ScalabilitySuited to smaller projectsDesigned for scalable and large-scale scraping
PerformanceSlower due to browser launch overheadFaster for traditional HTTP-based scraping

Scrapy wins on crawl scale and scheduling, and reaches for a browser plugin only when a site demands rendering.

Pyppeteer vs. Selenium

The closest comparison, since both drive a real browser:

FeaturePyppeteerSelenium
Browser SupportChromium onlyMultiple browsers (Chrome, Firefox, Edge, etc.)
AsynchronicityAsynchronous (async/await)Synchronous
MaintenanceDormant, last release 2.0.0 in February 2024Active
Use CasesExisting headless Chromium scriptsGeneral-purpose web automation and testing

Selenium remains the multi-browser, actively maintained option of the two, and the Selenium scraping guide covers it end to end.

Conclusion and Takeaways

Pyppeteer still runs, given the two setup adjustments a current system needs. An explicit executablePath, because the bundled Chromium download fails, and an eye on the websockets version it pins. The API itself remains a clean asynchronous browser interface, which is why so many scrapers were built on it.

The library’s dormancy is the real constraint. Existing scripts deserve the two setup fixes above, and their next rewrite deserves the migration table, since every method here has a maintained twin in Playwright. When the browser itself becomes the burden, with proxies and rendering to manage, our Web Scraping API executes the JavaScript on its side and returns the rendered HTML for parsing.

Valentina Skakun
Valentina Skakun
Valentina is a software engineer who builds data extraction tools before writing about them. With a strong background in Python, she also leverages her experience in JavaScript, PHP, R, and Ruby to reverse-engineer complex web architectures.If data renders in a browser, she will find a way to script its extraction.
Articles

Might Be Interesting