HasData
Back to all posts

Web Scraping with Playwright in Node.js (2026 Guide)

Web scraping with Playwright in Node.js means driving a real Chromium, Firefox, or WebKit through one API, which is what you want when the data appears only after JavaScript runs. This guide scrapes GitHub Topics, first one page and then pagination, plus request interception, forms, screenshots, and an Excel export, all built on locators rather than raw DOM queries. The proxy example needs an endpoint of your own.

Writing Python? The Playwright with Python guide covers the same ground in that language.

What is Playwright

Playwright is an open-source browser automation framework from Microsoft, built first for Node.js and since ported to Python, .NET, and Java. One API covers all three engines across Windows, Linux and macOS. Our own Node.js library benchmark measured it against Puppeteer on the same task: 251 ms cold start against 1,004 ms, and 167 MB peak memory against 379 MB.

Getting Started with Playwright

The setup is one npm package plus the browser binaries.

Project Setup and Installation

Playwright 1.62 declares Node 20 or newer, so a current Node.js has to be on the machine before anything else. Installing on 18 gets you an EBADENGINE warning and an unsupported runtime.

Create a directory, move into it, and initialize the project with npm init -y, which writes a package.json you can install into.

mkdir playwright-scraping
cd playwright-scraping
npm init -y

Now, you can install Playwright using NPM:

npm install playwright

To use Playwright, you’ll also need to install a compatible browser. Each Playwright version requires specific browser binary versions. Run the following command to install the latest browser versions:

npx playwright install

This will install the latest versions of Chromium, Firefox, and WebKit. You can use any of these browsers in your code, but we’ll use Chromium for this tutorial.

Here’s how the complete process looks:

Terminal showing npm init and npm install playwright finishing in a new project directory

Install npm packages

Open the package.json file and add "type": "module" to support modern JavaScript syntax.

package.json opened in the editor with the type field set to module

Add package.json file

Finally, open the project in your preferred code editor and create a new file named index.js.

Empty index.js created next to package.json in the project folder

Create index.js file

Launching Playwright

The first script opens a page in headless Chromium and closes it again.

import { chromium } from 'playwright';

async function main() {
    // Launch a new instance of a Chromium browser with headless mode
    // enabled for scraping
    const browser = await chromium.launch({
        headless: true
    });

    // Create a new Playwright context to isolate browsing session
    const context = await browser.newContext();
    // Open a new page/tab within the context
    const page = await context.newPage();

    // Navigate to the Node.js topic page on GitHub
    await page.goto('https://github.com/topics/nodejs');

    // Wait for the first repository card instead of a fixed delay
    await page.locator('article.border').first().waitFor();

    // Close the browser instance after task completion
    await browser.close();
}

// Execute the main function
main().catch((err) => {
    console.error(err);
    process.exit(1);
});

The chromium import controls Chromium-based browsers, chromium.launch({ headless: true }) starts one without a visible window, and a context isolates cookies and storage per session. page.goto() navigates to the Node.js topic page, where each repository is an article.border, and the waitFor() on a locator holds until that content is actually there. The topics index at github.com/topics lists topics rather than repositories and carries no such element, which is why the examples here all target a topic page. A fixed waitForTimeout would do the same job worse, since it’s either too short on a slow load or wasted time on a fast one.

That loads the page below, and from here on every example builds on this skeleton.

GitHub Node.js topic page with repository cards, the page the scripts target

Research github page

Basic Web Scraping with Playwright

With the environment in place, the first scraper reads one page of GitHub Topics. Anything a person can do in a browser window is available here, from screenshots to crawling.

Selecting Data to Scrape

We’ll be extracting data from GitHub topics. This will allow you to select the topic and the number of repositories you want to extract. The scraper will then return the information associated with the chosen topic.

GitHub topic page for nodejs with the repository grid the scraper reads

Scrape NodeJS packages

We’ll use Playwright to launch a browser, navigate to the GitHub topics page, and extract the necessary information. This includes details such as the repository owner, repository name, repository URL, the number of stars the repository has, its description, and any associated tags.

One repository card with owner, name, star count, description and topic tags marked out

Extract only useful data

Locating Elements and Extracting Data

A topic page renders 20 repositories, each one an <article> element. Expanding one in DevTools shows every field the scraper needs.

DevTools Elements panel with one article element expanded over the repository card it renders

Find tags using DevTools

The image below shows an expanded <article> element, displaying all the information about the repository.

Expanded article element showing the nested heading links, star counter and topic-tag anchors

Use classes

These are the selectors the current markup answers to:

FieldSelectorNote
Cardarticle.border20 per page
User and repositoryh3 afirst link is the owner, second the repository
Repository URLh3 a second link, hrefrelative, so resolve it against https://github.com
Stars#repo-stars-counter-starthe exact count sits in the title attribute
Descriptionp.color-fg-mutedabsent on some cards
Tagsa.topic-tagzero or more per card

GitHub has renamed the description element. The div.px-3 > p selector that older Playwright tutorials use matches nothing today, while p.color-fg-muted matches on all 60 cards. That’s the argument for locators over one-shot DOM queries, since a locator resolves at the moment of use and can be checked for existence before you read it.

Playwright offers two families of extraction. The page.$$eval(selector, fn) family serializes a function into the page and returns whatever it computes, which is fast but runs against the DOM as it stands at that instant. The page.locator(selector) family returns a handle that re-resolves on every call and waits for the element on its own. For scraping, the locator family removes the whole class of races that fixed delays are usually added to paper over.

Here’s one card’s worth of fields through locators:

const cards = page.locator('article.border');
await cards.first().waitFor();

const card = cards.first();
const links = card.locator('h3 a');
const user = (await links.nth(0).innerText()).trim();
const repoName = (await links.nth(1).innerText()).trim();
const repoUrl = new URL(await links.nth(1).getAttribute('href'), 'https://github.com').href;
const stars = await card.locator('#repo-stars-counter-star').getAttribute('title');
const tags = await card.locator('a.topic-tag').allInnerTexts();

allInnerTexts() returns every match at once, which is why the tags need no loop. Descriptions do need a guard, since a repository without one has no p at all:

const description = card.locator('p.color-fg-muted');
const text = (await description.count()) ? (await description.innerText()).trim() : '';

The complete single-page scraper puts those pieces in a loop over cards.all():

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();

await page.goto('https://github.com/topics/nodejs');

// One locator per field, resolved per card. Locators wait for the element themselves,
// so there is no fixed delay anywhere in this script.
const cards = page.locator('article.border');
await cards.first().waitFor();

const repos = [];
for (const card of await cards.all()) {
  const links = card.locator('h3 a');
  const description = card.locator('p.color-fg-muted');
  repos.push({
    user: (await links.nth(0).innerText()).trim(),
    repoName: (await links.nth(1).innerText()).trim(),
    repoUrl: new URL(await links.nth(1).getAttribute('href'), 'https://github.com').href,
    stars: await card.locator('#repo-stars-counter-star').getAttribute('title'),
    description: (await description.count()) ? (await description.innerText()).trim() : '',
    tags: (await card.locator('a.topic-tag').allInnerTexts()).map((t) => t.trim()),
  });
}

console.log(`Extracted ${repos.length} repositories`);
console.log(repos[0]);

await browser.close();

Running it prints the count and the first record:

Extracted 20 repositories
{
  user: 'freeCodeCamp',
  repoName: 'freeCodeCamp',
  repoUrl: 'https://github.com/freeCodeCamp/freeCodeCamp',
  stars: '454,932',
  description: "freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.",
  tags: [
    'react',         'nodejs',
    'javascript',    'd3',
    'teachers',      'community',
    'education',     'programming',
    'curriculum',    'math',
    'freecodecamp',  'learn-to-code',
    'nonprofits',    'careers',
    'certification'
  ]
}

For the wider Node.js picture beyond Playwright, our Node.js web scraping guide covers the HTTP-and-parser route that skips the browser entirely.

Advanced Scraping Techniques

One page is the easy case. The rest of this guide clicks buttons, paginates, intercepts requests, fills forms, and routes traffic through a proxy.

Clicking Buttons and Waiting for Actions

You can load more repositories by clicking the ‘Load more…’ button at the bottom of the page. Here are the actions to tell Playwright to load more repositories:

  1. Wait for the “Load more…” button to appear.
  2. Click the “Load more…” button.
  3. Wait for the new repositories to load before proceeding.

Bottom of the GitHub topic page with the Load more button that appends the next 20 cards

Add Load more button processing

Handling Dynamic Content and Navigation

GitHub loads the next 20 repositories behind a “Load more” button, so a scraper that wants 60 records clicks it twice. What the scraper needs to know is when the new cards have arrived. A fixed delay guesses at it, while waiting for the card count to grow measures it.

const cards = page.locator('article.border');
const loadMore = page.getByRole('button', { name: 'Load more' });

const before = await cards.count();
await loadMore.click();
await page.waitForFunction(
  (n) => document.querySelectorAll('article.border').length > n,
  before,
  { timeout: 15000 },
);

getByRole targets the button by what a user sees rather than by its class list, which on GitHub changes between deploys. The loop below repeats that until it has the number of records you asked for, skipping the cards it already read:

import { chromium } from 'playwright';

async function scrapeRepos(target) {
  const browser = await chromium.launch({ headless: true });
  const page = await (await browser.newContext()).newPage();
  const rows = [];

  try {
    await page.goto('https://github.com/topics/nodejs');
    const cards = page.locator('article.border');
    const loadMore = page.getByRole('button', { name: 'Load more' });

    while (rows.length < target) {
      await cards.first().waitFor();
      const seen = rows.length;

      for (const card of (await cards.all()).slice(seen)) {
        const links = card.locator('h3 a');
        const description = card.locator('p.color-fg-muted');
        rows.push({
          user: (await links.nth(0).innerText()).trim(),
          repoName: (await links.nth(1).innerText()).trim(),
          repoUrl: new URL(await links.nth(1).getAttribute('href'), 'https://github.com').href,
          stars: await card.locator('#repo-stars-counter-star').getAttribute('title'),
          description: (await description.count()) ? (await description.innerText()).trim() : '',
          tags: (await card.locator('a.topic-tag').allInnerTexts()).join(', '),
        });
      }

      if (rows.length >= target || !(await loadMore.count())) break;

      const before = await cards.count();
      await loadMore.click();
      try {
        await page.waitForFunction(
          (n) => document.querySelectorAll('article.border').length > n,
          before,
          { timeout: 15000 },
        );
      } catch {
        break; // no new cards arrived, keep what we have
      }
    }
  } finally {
    await browser.close();
  }

  return rows.slice(0, target);
}

const repos = await scrapeRepos(60);
console.log(`Collected ${repos.length} repositories`);

Three things in that loop do the work. The slice(seen) skips re-reading cards already collected, since GitHub appends rather than replaces. The catch { break } around the wait is what actually ends the crawl, because GitHub keeps the “Load more” button on the page long past the point where it returns anything new (it was still there after 26 clicks and 520 cards on a narrow topic), so the honest end-of-data signal is the wait timing out. Without that catch the exception escapes the function and every collected row is lost. And the finally closes the browser on both paths, which together with the catch is the error handling this scraper needs.

Using Proxies with Playwright

Sites that rate-limit by address or serve different content per country make the exit IP part of the scraper’s configuration. Playwright takes a proxy per browser launch, or per context when different pages need different exits. Pass the endpoint as server, and add username and password when the provider authenticates that way:

import { chromium } from 'playwright';

const launchOptions = {
    proxy: {
        server: 'http://YOUR_PROXY_HOST:PORT',
        // username: 'user',
        // password: 'pass',
    },
};

(async () => {
    const browser = await chromium.launch(launchOptions);
    const page = await browser.newPage();

    // A liveness check: does anything come back through this endpoint at all
    await page.goto('https://www.scrapethissite.com/pages/simple/');
    console.log(await page.locator('h3.country-name').count(), 'countries loaded');

    await browser.close();
})();

That tells you the proxy is alive, nothing more: the page is static, so the same 250 countries come back with or without a proxy in front. To see the exit address itself, request a page your own infrastructure serves and read the client IP from your logs, which keeps someone else’s echo service out of the loop.

Note: free proxies rarely last more than a few hours, so treat them as a way to check the wiring rather than as infrastructure for a real run.

Intercepting HTTP requests

With Playwright, you can easily monitor and modify network traffic, such as HTTP and HTTPS requests, XMLHttpRequests (XHRs), and fetch requests. Below is a code snippet that shows how to modify a request header.

import { chromium } from 'playwright';

(async () => {
    const browser = await chromium.launch({ headless: true });
    const context = await browser.newContext();
    const page = await context.newPage();

    await page.route('https://httpbin.org/headers', async (route, request) => {
        // Get original headers
        const originalHeaders = request.headers();

        // Modify the Accept-Language and User-Agent headers
        const modifiedHeaders = { ...originalHeaders };
        modifiedHeaders['accept-language'] = 'fr-FR'; // Change to French
        modifiedHeaders['user-agent'] = 'Mozilla/5.0 (Windows Phone 10.0; Android 4.2.1; Microsoft; RM-1127_16056) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.135 Mobile Safari/537.36 Edge/12.10536';

        // Continue the request with modified headers
        await route.continue({
            headers: modifiedHeaders,
        });
    });

    // Make the request with modified headers
    await page.goto('https://httpbin.org/headers');

    // Extract the data to see if the fields are updated
    const response = JSON.parse(await page.locator('pre').innerText());

    console.log('Response:', response);

    await browser.close();
})();

The code sets up a route handler for the URL https://httpbin.org/headers. This handler intercepts requests and modifies specific headers, like “Accept-Language” and “User-Agent”, before sending them.

Inside the handler function, it modifies the Accept-Language header to 'fr-FR' (French) and the User-Agent header as well. After these modifications, the handler continues the request with the updated headers using route.continue().

Opening https://httpbin.org/headers in a normal browser shows the unmodified header set.

httpbin headers response in a normal browser, showing the default Accept-Language and User-Agent

Original headers

Modified header, returned by our code.

httpbin headers response after the route handler rewrote Accept-Language to fr-FR and swapped the User-Agent

Modified headers

Filling Forms

fill() writes into an input and click() submits, and both work off locators. Targeting the field by its visible label or placeholder rather than by class survives a redesign:

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await (await browser.newContext()).newPage();

await page.goto('https://www.scrapethissite.com/pages/forms/');

// getByPlaceholder and getByRole target what the user sees, so they survive class renames
await page.getByPlaceholder('Search for Teams').fill('Toronto');

// The search page already lists 25 teams. Waiting for the result URL before counting
// keeps those rows out of the total.
await Promise.all([
  page.waitForURL(/q=Toronto/),
  page.getByRole('button', { name: 'Search' }).click(),
]);

const rows = page.locator('table.table tr.team');
await rows.first().waitFor();

console.log(`${await rows.count()} rows for "Toronto"`);
console.log((await rows.first().innerText()).replace(/\s+/g, ' ').trim());

await browser.close();

The run returns 21 rows, the first of them the 1990 Maple Leafs:

21 rows for "Toronto"
Toronto Maple Leafs 1990 23 46 0.287 241 318 -77

Login forms work the same way mechanically, with two differences worth planning for. Credentials belong in environment variables rather than in the script, and a site that fronts its login with a bot check will stop the automation there regardless of how the form is filled.

Screenshot Capture of Web Pages

Screenshots are one call, and they’re the fastest way to see what the scraper saw when a selector comes back empty.

import { chromium } from 'playwright';

async function screenShot() {
    const browser = await chromium.launch({
        headless: true
    });

    const context = await browser.newContext();
    const page = await context.newPage();

    await page.setViewportSize({ width: 1280, height: 800 }); // set screenshot dimension
    await page.goto('https://github.com/topics/nodejs')
    await page.screenshot({ path: 'images/screenshot.png' })
    await browser.close()
}
screenShot().catch(console.error);

The saved file:

Full-viewport screenshot saved by the script, the GitHub topic page at 1280 by 800

Find Repositories

To capture the entire page, set the fullPage property to true. You can also change the image format to jpg or jpeg for saving in different formats.

await page.screenshot({ path: 'images/screenshot.jpg', fullPage: true });

To capture a specific area on a webpage, use the clip property. It requires defining four values:

  • x: Horizontal distance from the top-left corner of the capture area.
  • y: Vertical distance from the top-left corner of the capture area.
  • width: Width of the capture area.
  • height: Height of the capture area.
await page.screenshot({
        path: "images/screenshot.png", fullPage: false, clip: {
            x: 5,
            y: 5,
            width: 320,
            height: 160
        }
    });

The clipped region:

Clipped screenshot, a 320 by 160 region cut from the top-left of the same page

Use topics

Save Scraped Data to Excel

Printing to the console is fine while developing. For a deliverable, exceljs writes the rows straight to a spreadsheet:

npm install exceljs

Columns map to the object keys the scraper already produces, so the export is three calls:

import ExcelJS from 'exceljs';

const workbook = new ExcelJS.Workbook();
const sheet = workbook.addWorksheet('GitHub Repositories');
sheet.columns = [
  { header: 'User', key: 'user' },
  { header: 'Repository', key: 'repoName' },
  { header: 'Stars', key: 'stars' },
  { header: 'Description', key: 'description' },
  { header: 'URL', key: 'repoUrl' },
  { header: 'Tags', key: 'tags' },
];
sheet.addRows(rows);
await workbook.xlsx.writeFile('github_repos.xlsx');

addRows takes the whole array at once, and tags needs to be a string by then, which is why the paginating scraper joins it with commas instead of keeping the array. Drop those three calls in place of the return rows.slice(0, target) at the end of scrapeRepos from the previous section, and the run writes a spreadsheet instead of returning an array.

Asking for 60 records clicks “Load more” twice and writes the file:

Saved 60 repositories to github_repos.xlsx

The result:

Excel file with the scraped GitHub repositories, one row per repository

File with Results

Comparison with other tools

Selenium and Puppeteer solve the same problem, and the differences that matter for scraping are startup cost, browser coverage, and how much waiting code you write yourself.

Playwright drives the three major engines through one API, and that API exists in Python, Java and .NET as well. Its documentation covers the whole surface.

Puppeteer is JavaScript only, and since v23 it drives Firefox as well as Chrome, so the browser gap between it and Playwright is narrower than it used to be. Selenium spans the most languages and browsers, at the cost of speed and of the waiting code you write yourself.

The cold-start row in the table above comes from our own Node.js benchmark, and it retires the old rule of thumb that Puppeteer is the faster of the two.

Weekly npm downloads have Playwright ahead of Puppeteer since 2024, and both far ahead of the Selenium WebDriver package.

The three tools line up like this:

ParameterPlaywrightPuppeteerSelenium
Cold start251 ms1,004 msslowest of the three
DocumentationExcellentExcellentFair
Developer ExperienceBestGoodFair
Language SupportJavaScript, Python, C#, JavaJavaScriptJava, Python, C#, Ruby, JavaScript, Kotlin
ByMicrosoftGoogleCommunity and Sponsors
Weekly npm downloads69M8.6M1.4M
Browser SupportChromium, Firefox, and WebKitChrome and FirefoxChrome, Firefox, IE, Edge, Opera, Safari

Selenium’s breadth is real, and so is its cost. It’s the slowest of the three and needs the most waiting code around it.

Conclusion

Playwright is worth its startup cost on pages that need a real browser. The startup and memory figures above say it costs less than Puppeteer to hold open, and both cost far more than an HTTP client and a parser do on a server-rendered page, which is the comparison that decides whether to open a browser at all.

The scripts here scrape GitHub Topics, first one page and then 60 records across pagination, with request interception, forms, screenshots, and an Excel export along the way. Every extraction goes through locators, which is what keeps them working when GitHub renames a class, as it did to the description selector older tutorials still use.

From here, the paths worth taking are resource blocking (dropping images and fonts cuts page weight before it costs you time), persistent contexts for session reuse, and the official documentation for the parts this guide skipped.

Roman Milyushkevich
Roman Milyushkevich
Roman Milyushkevich is the Co-founder and CTO at HasData, a web scraping API handling billions of requests. He designs the distributed systems, proxy infrastructure, and APIs behind large-scale, reliable data extraction. Roman writes on API design, browser automation, and building scraping pipelines that hold up in production.
Articles

Might Be Interesting