HasData
Back to all posts

Web Scraping vs API: The Best Way to Extract Data

Take the official API when it exists and carries your fields, scrape the page when it doesn’t, and before writing either, check the two routes in between, the internal API the page itself calls and a provider’s data API for the target. We measured the trade-offs on live targets for this comparison. On GitHub, the REST API returned 86 fields per repository with exact counts while the page showed rounded ones, at a fiftieth of the traffic. On Stack Overflow, the page answered a plain request with a 403 while the official API served the same question data anonymously, 300 calls a day.

Four Ways to Get the Data

Every “web scraping vs API” decision actually has four candidates. Scraping reads the rendered page, the official API returns documented JSON under a quota, the internal API is the undocumented endpoint the page itself fetches data from (often callable directly), and the third-party data API is a provider’s maintained scraper for one target, sold as documented JSON per request.

CriterionScraping the pageOfficial APIInternal API of the pageThird-party data API
AvailabilityAny public pageOnly where the site ships oneWherever the page loads data over XHRWhere a provider covers the target
FormatHTML you parse yourselfDocumented JSONUndocumented JSONDocumented JSON
ExactnessWhat the page shows, often roundedExact valuesExact values, usually the page’s own sourceThe page’s values, parsed by the provider
CoverageEverything a visitor seesThe fields the vendor chose to exposeWhat this page needs, sometimes more than the public APIThe provider’s schema, built from the page
Auth and quotasNone, but anti-bot insteadKeys, rate limits, paid tiersSession cookies or signed requests, no documented quotaKey, per-request billing
Blocking riskFull anti-bot surfaceContractual (throttling, key revocation)Between the two, and endpoints move without noticeOn the provider’s side
BreakageLayout changesVersioned, deprecations announcedSilent changes any dayLayout changes, fixed by the provider
TermsSite’s terms applyExplicit API termsUsually the same site terms, read themProvider terms, and the site’s terms still govern the data
MaintenanceSelector fixesVersion migrationsRe-discovery when the endpoint changesSchema updates on the provider’s timeline
Best forData no API exposesLong-lived integrationsJS-rendered pages without a public APICovered targets past the official quotas

On the live targets, the rows that decided the choice were quotas, exactness, and blocking, and the measured examples below put numbers on all three.

When the Official API Is Enough

An official API wins on stability and exactness, and it bills those advantages in quotas. The two we measured live make the range concrete. GitHub’s REST API allows 60 unauthenticated requests per hour (we read the live counter from its rate_limit endpoint mid-run) and 5,000 per hour with a token per its documentation. The Stack Exchange API answered anonymously with a quota_remaining field right in the response body, counting down from 300 requests per day.

One request against a documented endpoint replaces a page fetch and the parser behind it:

import requests

r = requests.get("https://api.github.com/repos/psf/requests", timeout=30)
repo = r.json()
print(repo["stargazers_count"], repo["open_issues_count"], repo["license"]["spdx_id"])

The response carries 86 top-level fields, including exact created_at and pushed_at timestamps the page shows only as relative dates, and an exact subscribers_count where the page rounds to “1.3k watching”. Pagination is documented (per_page caps at 100 on most list endpoints, with Link headers pointing at the next page), errors are status codes with JSON bodies, and none of it changes when the site redesigns.

The quota ladder is where planning happens, and it differs per vendor more than any other property:

TargetAnonymousWith a free key or tokenAbove that
GitHub REST API60 requests/hour (measured live)5,000/hour per its docsGitHub Apps quotas
Stack Exchange API300 requests/day (measured live)10,000/day per its docsnegotiated access
A page’s internal APIno documented quotanot applicableanti-bot decides
Scraping the pageno quotanot applicableanti-bot decides
A third-party data APIkey requiredtrial tiers vary by providerper-request billing

GitHub also attaches the live counter to every response as the X-RateLimit-Remaining header, so a client can check its budget without a separate call. Build that check in from day one.

The costs are the mirror image. The vendor picks the fields, so an API can lag the page or omit what you actually came for. Commercial catalogs in particular tend to expose inventory but not prices, reviews, or ranking positions. Quotas turn into paid tiers exactly at the volumes where scraping starts to look tempting. And the key ties every request to your account, which makes the terms of service enforceable in a way an anonymous page fetch is not.

Pros and cons of API scraping

Pros and cons of API scraping

Whether those trade-offs survive contact with a real target is what we measured next.

What the API Returns vs What the Page Shows

GitHub’s page rounded the star count on 10 of 10 repositories we checked, and the API returned the exact integer every time, in 69 KB against 3.6 MB of page weight. The method was one request per side per repository, the page parsed with BeautifulSoup against the API record for the same entity, and here is the full comparison:

RepositoryPage showsAPI returns
psf/requests54.3k stars54,278
pallets/flask72.2k stars72,169
django/django89.6k stars89,566
numpy/numpy32.6k stars32,642
pandas-dev/pandas49.6k stars49,622
scrapy/scrapy64.2k stars64,167
python/cpython75.6k stars75,593
microsoft/playwright95.5k stars95,524
puppeteer/puppeteer95.5k stars95,536
SeleniumHQ/selenium34.5k stars34,456

GitHub’s markup happens to keep the exact number in a title attribute, so a scraper can dig it out, and that is the kind of markup trivia an API user never needs to know:

# pip install requests beautifulsoup4
import requests
from bs4 import BeautifulSoup

page = requests.get("https://github.com/psf/requests",
                    headers={"User-Agent": "Mozilla/5.0"}, timeout=30)
star_el = BeautifulSoup(page.text, "html.parser").select_one("#repo-stars-counter-star")
api = requests.get("https://api.github.com/repos/psf/requests", timeout=30).json()

print("page shows:", star_el.get_text(strip=True))   # 54.3k
print("page title attr:", star_el.get("title"))      # 54,278
print("api returns:", api["stargazers_count"])       # 54278

Three lines of parsing, one selector to maintain, and the answer still depends on GitHub keeping that attribute, against one documented field on the API side.

The Stack Overflow side of the measurement made the blocking column concrete. The question page returned 403 to a plain requests call with a browser User-Agent, while the official API served the same question’s data anonymously. One response carried 15 fields, an exact view count of 1,998,112 on a question whose page our client never got to see, and the quota counter. For that target, the API is the only route that works without anti-bot tooling.

Both of our live targets are developer platforms with generous APIs. The pattern inverts on many commercial targets, where the public API exposes a trimmed subset and the page carries prices, reviews, and availability the API omits. That gap is what the internal endpoint often covers.

Finding the Internal API

Pages that render data with JavaScript get that data from somewhere, and that somewhere is an internal API you can usually call directly. Open DevTools, switch to the Network tab, filter by Fetch/XHR, reload the page, and look for responses whose JSON contains the values you see rendered. The request URL, its parameters, and its headers are all copyable from the same panel.

We measured the payoff on a sandbox page built for exactly this pattern. The initial HTML of scrapethissite.com’s AJAX page contains none of the film data it displays (we checked the raw response for the first rendered title, zero matches), so scraping it head-on requires a browser. The endpoint the page calls returns everything at once:

import requests

r = requests.get(
    "https://www.scrapethissite.com/pages/ajax-javascript/",
    params={"ajax": "true", "year": 2015},
    timeout=30,
)
films = r.json()
print(len(films), films[0])

The endpoint answered in 0.67 seconds with 1.6 KB of JSON, 16 films across a 5-field schema (best_picture appears only on the winner), against a 13 KB HTML shell that carries no data at all. The year parameter is also the endpoint’s pagination, so walking the archive is a loop over years, not over rendered pages. No browser, no rendering wait, no selectors, and the response is the page’s own source of truth, the same payload the visitors’ page renders.

Internal endpoints sit between the other two options on risk. They are undocumented, so they change without notice and carry no quota you can plan around, and the site’s terms of service apply to them the same way they apply to the page. Some sites version them, sign the requests, or check session cookies, and how much of that you meet is itself a signal of how much the operator minds. Treat a found endpoint as a fact about today’s page, not a contract.

Third-Party Data APIs

A fourth route runs through a provider. Third-party data APIs are maintained scrapers behind an API contract. The provider keeps the parsers, proxies, and rendering current for one target and sells the output as documented JSON, priced per request. The category exists for the two cases the first-party routes leave open.

The first case is a target with no official API at all. Zillow exposes no public listings API, so for an integration the choice is your own parser against a Zillow data API that returns listings and property records as JSON. The second case is an official API that answers a different question than the page. Google’s Custom Search JSON API allows 100 free queries per day with a hard ceiling of 10,000 per day, and it queries a configured search engine rather than returning the live SERP with its features. Rank tracking and SERP research need the page Google actually serves, so SERP APIs exist as their own category, and HasData’s Google SERP API returns that live SERP as structured JSON.

The trade-offs combine the columns on either side of it in the table above. Exactness and format behave like an official API, since the provider delivers parsed fields. Coverage and breakage behave like scraping, because underneath it is scraping: the schema follows the page, and a layout change breaks the provider’s parser until the provider’s fix restores it. The blocking surface moves off your infrastructure, and the terms question stays put, since the target site’s terms govern the data no matter who fetches it. What you give up is control of the schedule, because a changed schema or a deprecated field arrives on the provider’s timeline rather than yours.

When Scraping Is the Answer

Scraping remains the route for everything the previous three options exclude, plus every target no provider covers. Sites with no API, fields the API omits, and pages where the data exists only in rendered HTML all end here. It reads what a visitor sees, so its coverage is complete by definition, and everything else about it is maintenance. Selectors follow the layout, pagination means walking the next-page links yourself, dynamic content needs rendering or the internal endpoint above, and the anti-bot surface is all yours to handle. Our Stack Overflow probe is what that looks like in practice, a 403 before any of the question’s HTML arrives.

Pros and cons of web scraping

Pros and cons of web scraping

The build-vs-buy line inside scraping is where a web scraping API comes in. It takes a URL and returns the rendered HTML or extracted JSON, with proxies and JS rendering on the service side, billed per successful request. That keeps the coverage of scraping while outsourcing the two costs that grow with scale, infrastructure and blocking. Which service fits which workload is a separate comparison, and we keep one maintained in the web scraping API roundup.

Benefits of Web Scraping API

Benefits of Web Scraping API

With the three routes compared and measured, the terms themselves take one section to pin down.

What Web Scraping and API Scraping Mean

Web scraping is the automated extraction of data from web pages. Fetch the HTML, parse it, keep the fields. An API (Application Programming Interface) is a documented request-response contract a service exposes for programs. You send a request to an endpoint URL and receive structured data, usually JSON. API scraping is simply collecting data through such endpoints at scale, official or internal, instead of parsing pages.

In the everyday sense, using an API does not count as web scraping. An API returns data meant for programmatic consumption, while scraping extracts data from markup meant for people. The distinction matters legally and operationally, because API access is governed by explicit terms and keys, while page access is governed by the site’s general terms, and where scraping stands legally is its own topic.

Conclusion

The measured picture matches the folk wisdom only halfway. On the two developer platforms we measured, the official API won every axis, but the vendor decides what it carries and the quota decides how fast you can take it, and both are checkable before writing any code. Compare the API’s field list with the page for one entity you care about, divide your daily volume by the free quota, and pick per target. The internal endpoint and the parser cover the targets that fail either check, a third-party data API covers them without the maintenance wherever a provider already supports the target, and a scraping API absorbs the blocking work when you keep the parser yourself.

Sergey Ermakovich
Sergey Ermakovich
Sergey is the Co-founder and CMO at HasData, a web scraping API handling billions of requests. He specializes in web data extraction infrastructure, large-scale scraping reliability, and technical SEO. Sergey writes extensively on headless browser orchestration, API development, and scaling data pipelines for enterprise applications.
Articles

Might Be Interesting