HasData
Back to all posts

5 Chrome DevTools Tips & Tricks for Scraping

DevTools is the fastest web scraping tool you already have. Before any Python runs, the browser will tell you whether the selectors hold, whether the data is in the markup or arrives by a later request, and what that request looks like as code.

1. Copy a Selector, Then Prove It

Right-click the element on the page, pick Inspect, and the Elements panel opens on its node. From the node’s right-click menu, Copy offers Copy selector and Copy XPath:

The Elements panel on books.toscrape.com with a product node's context menu open and the Copy submenu showing Copy selector, Copy XPath, and Copy full XPath

Copy Element's XPath

Copied selectors are a starting point rather than an answer. Chrome builds them long and brittle, pinned to nth-child positions that break on the next listing. The habit that pays is proving a selector before it goes anywhere near code, in the Console:

// how many elements the CSS selector matches
$$('article.product_pod').length

// the same check for an XPath
$x("//article[@class='product_pod']//p[@class='price_color']").length

$$ is DevTools shorthand for querySelectorAll, and $x evaluates XPath. On the books.toscrape.com demo catalogue both return 20, one per product card, which is the number the page shows. A count of 0 means the selector is wrong, and a count far above the visible items means it is too loose. The same check works in the Elements panel search (Ctrl+F inside the panel), which highlights each match in the markup:

The Elements panel search box on the demo catalogue with a selector typed in and the matching nodes highlighted in the markup

Searching Element by Selector or XPath

Our CSS selectors cheat sheet covers how to write the short, stable kind by hand.

2. Extract Data in the Console Before Writing a Scraper

The Console runs the same DOM code a scraper will, so a whole extraction can be rehearsed on the open page. This returns the demo catalogue as objects, and copy() puts the result on the clipboard as JSON:

copy([...document.querySelectorAll('article.product_pod')].map(b => ({
  title: b.querySelector('h3 a').getAttribute('title'),
  price: b.querySelector('.price_color').textContent,
  inStock: b.querySelector('.availability').textContent.includes('In stock')
})))

That returns 20 rows with the title, price, and stock flag, pasted straight into a file. When a page carries structured data, the shortcut is better still. Many product pages embed their whole record as ld+json, and one line parses it:

JSON.parse(document.querySelector('script[type="application/ld+json"]').textContent)

That takes the first block on the page, which is the Product on some sites and a breadcrumb trail or an FAQ on others. When a page carries several, pick by type instead:

[...document.querySelectorAll('script[type="application/ld+json"]')]
  .map(s => JSON.parse(s.textContent))
  .find(d => d['@type'] === 'Product')

On a live product listing that one line returns the whole record:

Chrome Console on an Etsy listing with the parsed ld+json Product object expanded, showing rating, brand, price and sku fields

A product page's ld+json parsed in the Console

The expanded object carries the rating, brand, price, and SKU with no selectors at all, which is the technique our Etsy guide builds a whole scraper on.

3. Find the JSON the Page Is Actually Loading

When content appears on scroll or after a click, the markup is the wrong place to look. Open the Network tab, filter to Fetch/XHR, and act on the page. Whatever appears in the list is the request that carries the data:

Network tab filtered to Fetch/XHR on an infinite-scroll demo, with the quotes request selected and the JSON response shown

The JSON endpoint behind an infinite-scroll page

On the scrolling demo above, every scroll fires /api/quotes?page=N, and the Response tab shows clean JSON with the author, tags, and text. That endpoint answers a plain HTTP request, so a scraper skips the browser entirely and pages through the API the site built for itself.

Two checkboxes make this workable. Preserve log keeps the list across page loads, so a request fired during navigation survives to be read:

The Network tab toolbar with the Preserve log checkbox ticked, which keeps requests in the list across page loads

Preserve Log

And the Headers and Payload tabs on each request show what the server expects, the method, the parameters, and any tokens, which is everything a replay needs:

The Headers tab of a selected request, showing the method, the status, and the request and response header sets

Response & Request Headers

Reading those two tabs is where a dynamic-content scraper in Python starts.

4. Copy a Request as cURL, Then as Code

A request that works in the Network tab can leave DevTools as a complete command. Right-click it, Copy, Copy as cURL:

The Network tab on the quotes demo with a quotes request selected and Copy as cURL highlighted in the context menu

Convert a query from Network to code

The copied command carries the full header set, cookies included, so it reproduces the browser’s exact request. Pasting it into our curl to Python converter returns the same request as requests code, which is the shortest path from “the page does it” to “my script does it”. Strip the headers the request works without, since every kept header is one more thing that goes stale.

5. Debug with $0 and Live Expressions

$0 is the element currently selected in the Elements panel, so after clicking a node there, the Console can interrogate it directly:

$0.textContent          // what a scraper would read from it
getEventListeners($0)   // which JS is attached, e.g. the scroll handler that loads more

A live expression, the eye icon in the Console toolbar, re-evaluates an expression continuously. Pinning document.querySelectorAll('article.product_pod').length while scrolling shows the exact moment new items land in the DOM, which answers whether a wait in the scraper should key on time or on element count.

What This Replaces

A session with these techniques settles the plan before any code exists. The selectors are proven, the extraction is rehearsed, the markup-or-API question is decided, and the working request already exists as code. The scraper that gets written afterwards is shorter, because every guess it would have encoded is already answered.

Roman Milyushkevich
Roman Milyushkevich
Roman Milyushkevich is the Co-founder and CTO at HasData, a web scraping API handling billions of requests. He designs the distributed systems, proxy infrastructure, and APIs behind large-scale, reliable data extraction. Roman writes on API design, browser automation, and building scraping pipelines that hold up in production.
Articles

Might Be Interesting