Web scraping with Playwright in Node.js means driving a real Chromium, Firefox, or WebKit through one API, which is what you want when the data appears only after JavaScript runs. This guide scrapes GitHub Topics, first one page and then pagination, plus request interception, forms, screenshots, and an Excel export, all built on locators rather than raw DOM queries. The proxy example needs an endpoint of your own.
Writing Python? The Playwright with Python guide covers the same ground in that language.
What is Playwright
Playwright is an open-source browser automation framework from Microsoft, built first for Node.js and since ported to Python, .NET, and Java. One API covers all three engines across Windows, Linux and macOS. Our own Node.js library benchmark measured it against Puppeteer on the same task: 251 ms cold start against 1,004 ms, and 167 MB peak memory against 379 MB.
Getting Started with Playwright
The setup is one npm package plus the browser binaries.
Project Setup and Installation
Playwright 1.62 declares Node 20 or newer, so a current Node.js has to be on the machine before anything else. Installing on 18 gets you an EBADENGINE warning and an unsupported runtime.
Create a directory, move into it, and initialize the project with npm init -y, which writes a package.json you can install into.
mkdir playwright-scraping
cd playwright-scraping
npm init -yNow, you can install Playwright using NPM:
npm install playwrightTo use Playwright, you’ll also need to install a compatible browser. Each Playwright version requires specific browser binary versions. Run the following command to install the latest browser versions:
npx playwright installThis will install the latest versions of Chromium, Firefox, and WebKit. You can use any of these browsers in your code, but we’ll use Chromium for this tutorial.
Here’s how the complete process looks:

Open the package.json file and add "type": "module" to support modern JavaScript syntax.

Finally, open the project in your preferred code editor and create a new file named index.js.

Launching Playwright
The first script opens a page in headless Chromium and closes it again.
import { chromium } from 'playwright';
async function main() {
// Launch a new instance of a Chromium browser with headless mode
// enabled for scraping
const browser = await chromium.launch({
headless: true
});
// Create a new Playwright context to isolate browsing session
const context = await browser.newContext();
// Open a new page/tab within the context
const page = await context.newPage();
// Navigate to the Node.js topic page on GitHub
await page.goto('https://github.com/topics/nodejs');
// Wait for the first repository card instead of a fixed delay
await page.locator('article.border').first().waitFor();
// Close the browser instance after task completion
await browser.close();
}
// Execute the main function
main().catch((err) => {
console.error(err);
process.exit(1);
});The chromium import controls Chromium-based browsers, chromium.launch({ headless: true }) starts one without a visible window, and a context isolates cookies and storage per session. page.goto() navigates to the Node.js topic page, where each repository is an article.border, and the waitFor() on a locator holds until that content is actually there. The topics index at github.com/topics lists topics rather than repositories and carries no such element, which is why the examples here all target a topic page. A fixed waitForTimeout would do the same job worse, since it’s either too short on a slow load or wasted time on a fast one.
That loads the page below, and from here on every example builds on this skeleton.

Basic Web Scraping with Playwright
With the environment in place, the first scraper reads one page of GitHub Topics. Anything a person can do in a browser window is available here, from screenshots to crawling.
Selecting Data to Scrape
We’ll be extracting data from GitHub topics. This will allow you to select the topic and the number of repositories you want to extract. The scraper will then return the information associated with the chosen topic.

We’ll use Playwright to launch a browser, navigate to the GitHub topics page, and extract the necessary information. This includes details such as the repository owner, repository name, repository URL, the number of stars the repository has, its description, and any associated tags.

Locating Elements and Extracting Data
A topic page renders 20 repositories, each one an <article> element. Expanding one in DevTools shows every field the scraper needs.

The image below shows an expanded <article> element, displaying all the information about the repository.

These are the selectors the current markup answers to:
| Field | Selector | Note |
|---|---|---|
| Card | article.border | 20 per page |
| User and repository | h3 a | first link is the owner, second the repository |
| Repository URL | h3 a second link, href | relative, so resolve it against https://github.com |
| Stars | #repo-stars-counter-star | the exact count sits in the title attribute |
| Description | p.color-fg-muted | absent on some cards |
| Tags | a.topic-tag | zero or more per card |
GitHub has renamed the description element. The div.px-3 > p selector that older Playwright tutorials use matches nothing today, while p.color-fg-muted matches on all 60 cards. That’s the argument for locators over one-shot DOM queries, since a locator resolves at the moment of use and can be checked for existence before you read it.
Playwright offers two families of extraction. The page.$$eval(selector, fn) family serializes a function into the page and returns whatever it computes, which is fast but runs against the DOM as it stands at that instant. The page.locator(selector) family returns a handle that re-resolves on every call and waits for the element on its own. For scraping, the locator family removes the whole class of races that fixed delays are usually added to paper over.
Here’s one card’s worth of fields through locators:
const cards = page.locator('article.border');
await cards.first().waitFor();
const card = cards.first();
const links = card.locator('h3 a');
const user = (await links.nth(0).innerText()).trim();
const repoName = (await links.nth(1).innerText()).trim();
const repoUrl = new URL(await links.nth(1).getAttribute('href'), 'https://github.com').href;
const stars = await card.locator('#repo-stars-counter-star').getAttribute('title');
const tags = await card.locator('a.topic-tag').allInnerTexts();allInnerTexts() returns every match at once, which is why the tags need no loop. Descriptions do need a guard, since a repository without one has no p at all:
const description = card.locator('p.color-fg-muted');
const text = (await description.count()) ? (await description.innerText()).trim() : '';The complete single-page scraper puts those pieces in a loop over cards.all():
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();
await page.goto('https://github.com/topics/nodejs');
// One locator per field, resolved per card. Locators wait for the element themselves,
// so there is no fixed delay anywhere in this script.
const cards = page.locator('article.border');
await cards.first().waitFor();
const repos = [];
for (const card of await cards.all()) {
const links = card.locator('h3 a');
const description = card.locator('p.color-fg-muted');
repos.push({
user: (await links.nth(0).innerText()).trim(),
repoName: (await links.nth(1).innerText()).trim(),
repoUrl: new URL(await links.nth(1).getAttribute('href'), 'https://github.com').href,
stars: await card.locator('#repo-stars-counter-star').getAttribute('title'),
description: (await description.count()) ? (await description.innerText()).trim() : '',
tags: (await card.locator('a.topic-tag').allInnerTexts()).map((t) => t.trim()),
});
}
console.log(`Extracted ${repos.length} repositories`);
console.log(repos[0]);
await browser.close();Running it prints the count and the first record:
Extracted 20 repositories
{
user: 'freeCodeCamp',
repoName: 'freeCodeCamp',
repoUrl: 'https://github.com/freeCodeCamp/freeCodeCamp',
stars: '454,932',
description: "freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.",
tags: [
'react', 'nodejs',
'javascript', 'd3',
'teachers', 'community',
'education', 'programming',
'curriculum', 'math',
'freecodecamp', 'learn-to-code',
'nonprofits', 'careers',
'certification'
]
}For the wider Node.js picture beyond Playwright, our Node.js web scraping guide covers the HTTP-and-parser route that skips the browser entirely.
Advanced Scraping Techniques
One page is the easy case. The rest of this guide clicks buttons, paginates, intercepts requests, fills forms, and routes traffic through a proxy.
Clicking Buttons and Waiting for Actions
You can load more repositories by clicking the ‘Load more…’ button at the bottom of the page. Here are the actions to tell Playwright to load more repositories:
- Wait for the “Load more…” button to appear.
- Click the “Load more…” button.
- Wait for the new repositories to load before proceeding.

Handling Dynamic Content and Navigation
GitHub loads the next 20 repositories behind a “Load more” button, so a scraper that wants 60 records clicks it twice. What the scraper needs to know is when the new cards have arrived. A fixed delay guesses at it, while waiting for the card count to grow measures it.
const cards = page.locator('article.border');
const loadMore = page.getByRole('button', { name: 'Load more' });
const before = await cards.count();
await loadMore.click();
await page.waitForFunction(
(n) => document.querySelectorAll('article.border').length > n,
before,
{ timeout: 15000 },
);getByRole targets the button by what a user sees rather than by its class list, which on GitHub changes between deploys. The loop below repeats that until it has the number of records you asked for, skipping the cards it already read:
import { chromium } from 'playwright';
async function scrapeRepos(target) {
const browser = await chromium.launch({ headless: true });
const page = await (await browser.newContext()).newPage();
const rows = [];
try {
await page.goto('https://github.com/topics/nodejs');
const cards = page.locator('article.border');
const loadMore = page.getByRole('button', { name: 'Load more' });
while (rows.length < target) {
await cards.first().waitFor();
const seen = rows.length;
for (const card of (await cards.all()).slice(seen)) {
const links = card.locator('h3 a');
const description = card.locator('p.color-fg-muted');
rows.push({
user: (await links.nth(0).innerText()).trim(),
repoName: (await links.nth(1).innerText()).trim(),
repoUrl: new URL(await links.nth(1).getAttribute('href'), 'https://github.com').href,
stars: await card.locator('#repo-stars-counter-star').getAttribute('title'),
description: (await description.count()) ? (await description.innerText()).trim() : '',
tags: (await card.locator('a.topic-tag').allInnerTexts()).join(', '),
});
}
if (rows.length >= target || !(await loadMore.count())) break;
const before = await cards.count();
await loadMore.click();
try {
await page.waitForFunction(
(n) => document.querySelectorAll('article.border').length > n,
before,
{ timeout: 15000 },
);
} catch {
break; // no new cards arrived, keep what we have
}
}
} finally {
await browser.close();
}
return rows.slice(0, target);
}
const repos = await scrapeRepos(60);
console.log(`Collected ${repos.length} repositories`);Three things in that loop do the work. The slice(seen) skips re-reading cards already collected, since GitHub appends rather than replaces. The catch { break } around the wait is what actually ends the crawl, because GitHub keeps the “Load more” button on the page long past the point where it returns anything new (it was still there after 26 clicks and 520 cards on a narrow topic), so the honest end-of-data signal is the wait timing out. Without that catch the exception escapes the function and every collected row is lost. And the finally closes the browser on both paths, which together with the catch is the error handling this scraper needs.
Using Proxies with Playwright
Sites that rate-limit by address or serve different content per country make the exit IP part of the scraper’s configuration. Playwright takes a proxy per browser launch, or per context when different pages need different exits. Pass the endpoint as server, and add username and password when the provider authenticates that way:
import { chromium } from 'playwright';
const launchOptions = {
proxy: {
server: 'http://YOUR_PROXY_HOST:PORT',
// username: 'user',
// password: 'pass',
},
};
(async () => {
const browser = await chromium.launch(launchOptions);
const page = await browser.newPage();
// A liveness check: does anything come back through this endpoint at all
await page.goto('https://www.scrapethissite.com/pages/simple/');
console.log(await page.locator('h3.country-name').count(), 'countries loaded');
await browser.close();
})();That tells you the proxy is alive, nothing more: the page is static, so the same 250 countries come back with or without a proxy in front. To see the exit address itself, request a page your own infrastructure serves and read the client IP from your logs, which keeps someone else’s echo service out of the loop.
Note: free proxies rarely last more than a few hours, so treat them as a way to check the wiring rather than as infrastructure for a real run.
Intercepting HTTP requests
With Playwright, you can easily monitor and modify network traffic, such as HTTP and HTTPS requests, XMLHttpRequests (XHRs), and fetch requests. Below is a code snippet that shows how to modify a request header.
import { chromium } from 'playwright';
(async () => {
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext();
const page = await context.newPage();
await page.route('https://httpbin.org/headers', async (route, request) => {
// Get original headers
const originalHeaders = request.headers();
// Modify the Accept-Language and User-Agent headers
const modifiedHeaders = { ...originalHeaders };
modifiedHeaders['accept-language'] = 'fr-FR'; // Change to French
modifiedHeaders['user-agent'] = 'Mozilla/5.0 (Windows Phone 10.0; Android 4.2.1; Microsoft; RM-1127_16056) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.135 Mobile Safari/537.36 Edge/12.10536';
// Continue the request with modified headers
await route.continue({
headers: modifiedHeaders,
});
});
// Make the request with modified headers
await page.goto('https://httpbin.org/headers');
// Extract the data to see if the fields are updated
const response = JSON.parse(await page.locator('pre').innerText());
console.log('Response:', response);
await browser.close();
})();The code sets up a route handler for the URL https://httpbin.org/headers. This handler intercepts requests and modifies specific headers, like “Accept-Language” and “User-Agent”, before sending them.
Inside the handler function, it modifies the Accept-Language header to 'fr-FR' (French) and the User-Agent header as well. After these modifications, the handler continues the request with the updated headers using route.continue().
Opening https://httpbin.org/headers in a normal browser shows the unmodified header set.

Modified header, returned by our code.

Filling Forms
fill() writes into an input and click() submits, and both work off locators. Targeting the field by its visible label or placeholder rather than by class survives a redesign:
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await (await browser.newContext()).newPage();
await page.goto('https://www.scrapethissite.com/pages/forms/');
// getByPlaceholder and getByRole target what the user sees, so they survive class renames
await page.getByPlaceholder('Search for Teams').fill('Toronto');
// The search page already lists 25 teams. Waiting for the result URL before counting
// keeps those rows out of the total.
await Promise.all([
page.waitForURL(/q=Toronto/),
page.getByRole('button', { name: 'Search' }).click(),
]);
const rows = page.locator('table.table tr.team');
await rows.first().waitFor();
console.log(`${await rows.count()} rows for "Toronto"`);
console.log((await rows.first().innerText()).replace(/\s+/g, ' ').trim());
await browser.close();The run returns 21 rows, the first of them the 1990 Maple Leafs:
21 rows for "Toronto"
Toronto Maple Leafs 1990 23 46 0.287 241 318 -77Login forms work the same way mechanically, with two differences worth planning for. Credentials belong in environment variables rather than in the script, and a site that fronts its login with a bot check will stop the automation there regardless of how the form is filled.
Screenshot Capture of Web Pages
Screenshots are one call, and they’re the fastest way to see what the scraper saw when a selector comes back empty.
import { chromium } from 'playwright';
async function screenShot() {
const browser = await chromium.launch({
headless: true
});
const context = await browser.newContext();
const page = await context.newPage();
await page.setViewportSize({ width: 1280, height: 800 }); // set screenshot dimension
await page.goto('https://github.com/topics/nodejs')
await page.screenshot({ path: 'images/screenshot.png' })
await browser.close()
}
screenShot().catch(console.error);The saved file:

To capture the entire page, set the fullPage property to true. You can also change the image format to jpg or jpeg for saving in different formats.
await page.screenshot({ path: 'images/screenshot.jpg', fullPage: true });To capture a specific area on a webpage, use the clip property. It requires defining four values:
x: Horizontal distance from the top-left corner of the capture area.y: Vertical distance from the top-left corner of the capture area.width: Width of the capture area.height: Height of the capture area.
await page.screenshot({
path: "images/screenshot.png", fullPage: false, clip: {
x: 5,
y: 5,
width: 320,
height: 160
}
});The clipped region:

Save Scraped Data to Excel
Printing to the console is fine while developing. For a deliverable, exceljs writes the rows straight to a spreadsheet:
npm install exceljsColumns map to the object keys the scraper already produces, so the export is three calls:
import ExcelJS from 'exceljs';
const workbook = new ExcelJS.Workbook();
const sheet = workbook.addWorksheet('GitHub Repositories');
sheet.columns = [
{ header: 'User', key: 'user' },
{ header: 'Repository', key: 'repoName' },
{ header: 'Stars', key: 'stars' },
{ header: 'Description', key: 'description' },
{ header: 'URL', key: 'repoUrl' },
{ header: 'Tags', key: 'tags' },
];
sheet.addRows(rows);
await workbook.xlsx.writeFile('github_repos.xlsx');addRows takes the whole array at once, and tags needs to be a string by then, which is why the paginating scraper joins it with commas instead of keeping the array. Drop those three calls in place of the return rows.slice(0, target) at the end of scrapeRepos from the previous section, and the run writes a spreadsheet instead of returning an array.
Asking for 60 records clicks “Load more” twice and writes the file:
Saved 60 repositories to github_repos.xlsxThe result:

Comparison with other tools
Selenium and Puppeteer solve the same problem, and the differences that matter for scraping are startup cost, browser coverage, and how much waiting code you write yourself.
Playwright drives the three major engines through one API, and that API exists in Python, Java and .NET as well. Its documentation covers the whole surface.
Puppeteer is JavaScript only, and since v23 it drives Firefox as well as Chrome, so the browser gap between it and Playwright is narrower than it used to be. Selenium spans the most languages and browsers, at the cost of speed and of the waiting code you write yourself.
The cold-start row in the table above comes from our own Node.js benchmark, and it retires the old rule of thumb that Puppeteer is the faster of the two.
Weekly npm downloads have Playwright ahead of Puppeteer since 2024, and both far ahead of the Selenium WebDriver package.
The three tools line up like this:
| Parameter | Playwright | Puppeteer | Selenium |
|---|---|---|---|
| Cold start | 251 ms | 1,004 ms | slowest of the three |
| Documentation | Excellent | Excellent | Fair |
| Developer Experience | Best | Good | Fair |
| Language Support | JavaScript, Python, C#, Java | JavaScript | Java, Python, C#, Ruby, JavaScript, Kotlin |
| By | Microsoft | Community and Sponsors | |
| Weekly npm downloads | 69M | 8.6M | 1.4M |
| Browser Support | Chromium, Firefox, and WebKit | Chrome and Firefox | Chrome, Firefox, IE, Edge, Opera, Safari |
Selenium’s breadth is real, and so is its cost. It’s the slowest of the three and needs the most waiting code around it.
Conclusion
Playwright is worth its startup cost on pages that need a real browser. The startup and memory figures above say it costs less than Puppeteer to hold open, and both cost far more than an HTTP client and a parser do on a server-rendered page, which is the comparison that decides whether to open a browser at all.
The scripts here scrape GitHub Topics, first one page and then 60 records across pagination, with request interception, forms, screenshots, and an Excel export along the way. Every extraction goes through locators, which is what keeps them working when GitHub renames a class, as it did to the description selector older tutorials still use.
From here, the paths worth taking are resource blocking (dropping images and fonts cuts page weight before it costs you time), persistent contexts for session reuse, and the official documentation for the parts this guide skipped.


