Pyppeteer is a Python port of the Puppeteer library from Node.js, and it’s unmaintained. The last release, 2.0.0, shipped in February 2024, and the project has been quiet since. For a new scraper, start with Playwright instead, which covers the same ground with a maintained API. This guide stays useful for the code that already exists, so it covers installing Pyppeteer on a current system (which takes one extra step now), complete runnable examples for every common task, and a migration map for moving the code to Playwright when the time comes.
What is Pyppeteer
Pyppeteer drives a Chromium browser from Python and simulates the actions of a real user. It arrived as a port of Puppeteer, so its method names are Puppeteer’s camelCase (newPage, waitForSelector) and its model is fully asynchronous. The feature set covers page management, selectors and XPath, JavaScript execution in the page’s context, and screenshots.
Getting Started with Pyppeteer
Pyppeteer 2.0.0 declares Python 3.8 or higher and below 4.0, and any code editor will do. The installation has two traps on a current system.
How to install Pyppeteer
The package itself installs normally:
pip install pyppeteerThe first trap is Chromium. Pyppeteer tries to download its own build on first launch, and that download fails on current systems, so point executablePath at a browser you already have. Any Chromium works, including the one Playwright installs (on Windows that is under %LOCALAPPDATA%\ms-playwright\). The examples below pass it explicitly.
The second trap is easier to miss. Installing pyppeteer downgrades the websockets package to a 2022 release, which can break other libraries living in the same environment. Give Pyppeteer projects their own virtual environment, or expect to re-upgrade websockets after working with it.
Pyppeteer supported browsers
Pyppeteer controls Chromium only. Firefox, WebKit, and everything else stay out of reach, which is one of the concrete reasons the migration section at the end of this guide exists.
Basic Pyppeteer Project
Here’s the smallest complete script, and two details in it matter. The import is lowercase pyppeteer (a capitalized import fails with ModuleNotFoundError), and the entry point is asyncio.run(), since the older get_event_loop() pattern is deprecated in current Python:
import asyncio
from pyppeteer import launch
CHROME = r"C:\path\to\chrome.exe" # any Chromium build
async def main():
browser = await launch(headless=True, executablePath=CHROME)
page = await browser.newPage()
await page.goto("https://books.toscrape.com/")
print("Title:", await page.title())
await browser.close()
asyncio.run(main())It prints:
Title: All products | Books to Scrape - SandboxEvery script below keeps this shape and only the body of main() changes.
Advanced Configuration with Pyppeteer
A proxy, a user agent and cookies come up in almost every scraper, and one script covers all three:
import asyncio
from pyppeteer import launch
CHROME = r"C:\path\to\chrome.exe"
UA = ("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36")
async def main():
browser = await launch(
headless=True,
executablePath=CHROME,
# args=[f"--proxy-server=http://proxy_ip:proxy_port"],
)
page = await browser.newPage()
await page.setUserAgent(UA)
await page.setCookie({
"name": "example_cookie",
"value": "123456789",
"url": "https://httpbin.org",
})
await page.goto("https://httpbin.org/headers")
print(await page.evaluate("document.body.innerText"))
await browser.close()
asyncio.run(main())The run’s output showed the Chrome/152 string echoed back in the User-Agent header, which confirms the override reached the server.
Using proxies
The proxy travels as a Chromium argument in launch(), the commented line in the script above. Uncomment it and fill in the address. For authenticated proxies the format is username:password@ip:port. We covered how proxies work in Python and where rotating ones come from separately.
Modifying user agents
setUserAgent() applies per page, after newPage() and before goto(). Take a current string from our auto-updated user agent table rather than writing one by hand. Rotating through a small list of them per page lowers the profile of a long run.
Managing cookies
setCookie() also applies per page, and the dictionary needs a url (or domain) key in addition to name and value, or Chromium rejects it. The full parameter list is in the API reference the project publishes.
Navigating and Extracting Data
One script demonstrates the whole extraction toolbox, the page’s HTML, a CSS selector, an XPath query, and the wait that makes dynamic pages readable:
import asyncio
from pyppeteer import launch
CHROME = r"C:\path\to\chrome.exe"
async def main():
browser = await launch(headless=True, executablePath=CHROME)
page = await browser.newPage()
await page.goto(
"https://books.toscrape.com/",
{"timeout": 15000, "waitUntil": "load"},
)
await page.waitForSelector("article.product_pod")
content = await page.content()
print("HTML length:", len(content))
heading = await page.querySelector("h1")
print("H1:", await page.evaluate("(e) => e.textContent", heading))
links = await page.xpath("//article[@class='product_pod']//h3/a")
print("products:", len(links))
print("first:", await page.evaluate("(e) => e.getAttribute('title')", links[0]))
await browser.close()
asyncio.run(main())That returns the full page, around 51,000 characters of HTML, plus the heading and 20 product links with “A Light in the Attic” first.
Get HTML content
page.content() returns the rendered HTML, after JavaScript ran, which is the whole reason to pay for a browser instead of calling requests.
Using XPath and CSS selectors for Data Scraping
querySelector() takes a CSS selector and returns one element handle, querySelectorAll() returns them all, and xpath() always returns a list, so the element comes out by index. A handle isn’t text yet. Reading it goes through page.evaluate() with the element passed in, as both lookups in the script show. Which query language to prefer is its own topic, covered in the CSS selectors cheat sheet and the XPath guide.
Dealing with timeouts
The timeout option on goto() caps how long navigation may take, in milliseconds, and the script above allows 15 seconds before a timeout error is raised.
Implementing waits for page loading
waitUntil decides what counts as “loaded”: load waits for the load event, domcontentloaded settles for the DOM. For scraping, the stronger signal is usually an element you actually need.
Scrape Dynamic Pages
That’s what waitForSelector() does in the script. It continues the moment the element exists, however long that took. On the catalogue used above the products ship in the server’s HTML, so the wait returns immediately and costs nothing. On a page that builds its list in the browser, the same line is what stops the script from reading an empty DOM.
Interacting with the Page
Forms, clicks, scrolling, and screenshots combine into the login flow below, which runs against a sandbox login form that accepts any credentials:
import asyncio
from pyppeteer import launch
CHROME = r"C:\path\to\chrome.exe"
async def main():
browser = await launch(headless=True, executablePath=CHROME)
page = await browser.newPage()
await page.goto("https://quotes.toscrape.com/login")
await page.type("input#username", "reader")
await page.type("input#password", "example")
# a click that triggers navigation must be awaited together with it
await asyncio.gather(
page.waitForNavigation(),
page.click("input[type='submit']"),
)
logged_in = await page.querySelector("a[href='/logout']")
print("logged in:", logged_in is not None)
await page.evaluate("window.scrollBy(0, 300)")
await page.screenshot({"path": "screenshot.png"})
await browser.close()
asyncio.run(main())The run printed logged in: True and saved the screenshot.
Clicking buttons and elements
click() takes a selector. The nuance the script encodes is navigation. When a click loads a new page, click() and waitForNavigation() must run together through asyncio.gather(), or the script races the page and reads the old one.
Fill an input field
type() fills a field character by character, like a person would. Reading a field back goes through evaluate():
value = await page.evaluate('document.querySelector("input#username").value')Reading fields back covers the case of values a page generates on its own.
Executing specific actions
page.evaluate() runs any JavaScript in the page’s context, which covers everything without a dedicated method. Two dedicated ones worth knowing are page.keyboard.press("Enter"), which submits most login forms without hunting for the button, and page.hover(".selector") for menus that open on hover.
Login with Pyppeteer
The full flow is the script above. Type into both fields, click together with waitForNavigation(), and confirm the login by looking for an element that only exists in the signed-in state, like the logout link.
Page Scrolling
Scrolling is JavaScript through evaluate(), and window.scrollBy(0, 300) in the script moves the viewport down 300 pixels. Repeating it in a loop until the page height stops changing is the standard infinite-scroll pattern.
Capturing screenshots
screenshot() takes a dictionary with a path, and the extension picks the format (PNG or JPEG). The folder must already exist, or the call fails.
Troubleshooting and Error Handling
The errors people actually hit with Pyppeteer in its unmaintained state cluster around the browser, not the code.
Common issues like “Pyppeteer is not installed”
A capitalized import is the first thing to check, since from Pyppeteer import launch fails with ModuleNotFoundError while the package sits installed right there. The module is lowercase pyppeteer. Beyond that, the message usually means the installation itself ended with an error, and reinstalling with pip install pyppeteer in a clean virtual environment settles it.
Handling unexpected browser closures in Pyppeteer
A browser that dies at launch on a current system is almost always the Chromium download problem from the installation section, and executablePath pointed at an existing build is the fix. pyppeteer-install re-attempts the bundled download, which no longer succeeds on current systems. For crashes mid-run, wrap the session in try/finally so browser.close() runs either way.
Migrating Pyppeteer Code to Playwright
Since the project is dormant, every Pyppeteer script is one dependency freeze away from being legacy code, and the move to Playwright is mostly mechanical. The model is the same, an async browser driven from Python, and the method names translate one to one:
| Pyppeteer | Playwright (async) |
|---|---|
from pyppeteer import launch | from playwright.async_api import async_playwright |
browser = await launch() | browser = await p.chromium.launch() |
page = await browser.newPage() | page = await browser.new_page() |
await page.goto(url) | await page.goto(url) |
await page.content() | await page.content() |
await page.querySelector("h1") | page.locator("h1") |
await page.xpath("//h1") | page.locator("//h1") |
await page.waitForSelector(".x") | await page.wait_for_selector(".x") |
await page.setUserAgent(ua) | await browser.new_context(user_agent=ua) |
await page.setCookie(c) | await context.add_cookies([c]) |
await page.type(sel, text) | await page.fill(sel, text) |
await page.evaluate(js) | await page.evaluate(js) |
await page.screenshot({"path": p}) | await page.screenshot(path=p) |
Three habits change along the way. Locators replace element handles, so reading text no longer goes through evaluate() and becomes await page.locator("h1").inner_text(). Waiting mostly disappears, because locators wait for their elements on their own. And the browser download works again, since playwright install chromium fetches a current build instead of failing. The basic example above becomes this in Playwright:
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto("https://books.toscrape.com/")
print("Title:", await page.title())
await browser.close()
asyncio.run(main())Both versions of that script print the same title, and the Playwright guide continues from here.
Pyppeteer vs. Other Tools
The comparisons below assume you’re choosing for existing code or constrained environments, since for new projects the maintenance status already decides.
Pyppeteer vs. BeautifulSoup
The two occupy opposite ends of the stack:
| Feature | Pyppeteer | BeautifulSoup |
|---|---|---|
| Purpose | Automation of browser interactions, dynamic content execution, and user actions. | Parsing HTML and XML documents, extracting information from static markup. |
| Ease of Use | More involved, a browser instance and asynchronous code to manage. | A simpler, declarative approach to HTML parsing. |
| Performance | Slower, a browser renders every page. | Faster for static page parsing. |
| Use Cases | Dynamic web pages, data that appears after JavaScript runs, screenshots. | Data extraction from static HTML, navigating and searching through markup. |
For static pages, BeautifulSoup with requests is the lighter tool. A browser is worth its memory and startup time only where the content is built by JavaScript.
Pyppeteer vs. Scrapy
Scrapy competes on crawl scale rather than rendering:
| Feature | Pyppeteer | Scrapy |
|---|---|---|
| Purpose | Web scraping with browser automation | General-purpose web crawling and scraping |
| Browser Automation | Yes | No (focused on HTTP requests) |
| Ease of Use | More complex due to browser integration | Easier for its intended use |
| Scalability | Suited to smaller projects | Designed for scalable and large-scale scraping |
| Performance | Slower due to browser launch overhead | Faster for traditional HTTP-based scraping |
Scrapy wins on crawl scale and scheduling, and reaches for a browser plugin only when a site demands rendering.
Pyppeteer vs. Selenium
The closest comparison, since both drive a real browser:
| Feature | Pyppeteer | Selenium |
|---|---|---|
| Browser Support | Chromium only | Multiple browsers (Chrome, Firefox, Edge, etc.) |
| Asynchronicity | Asynchronous (async/await) | Synchronous |
| Maintenance | Dormant, last release 2.0.0 in February 2024 | Active |
| Use Cases | Existing headless Chromium scripts | General-purpose web automation and testing |
Selenium remains the multi-browser, actively maintained option of the two, and the Selenium scraping guide covers it end to end.
Conclusion and Takeaways
Pyppeteer still runs, given the two setup adjustments a current system needs. An explicit executablePath, because the bundled Chromium download fails, and an eye on the websockets version it pins. The API itself remains a clean asynchronous browser interface, which is why so many scrapers were built on it.
The library’s dormancy is the real constraint. Existing scripts deserve the two setup fixes above, and their next rewrite deserves the migration table, since every method here has a maintained twin in Playwright. When the browser itself becomes the burden, with proxies and rendering to manage, our Web Scraping API executes the JavaScript on its side and returns the rendered HTML for parsing.


