Take the official API when it exists and carries your fields, scrape the page when it doesn’t, and before writing either, check the two routes in between, the internal API the page itself calls and a provider’s data API for the target. We measured the trade-offs on live targets for this comparison. On GitHub, the REST API returned 86 fields per repository with exact counts while the page showed rounded ones, at a fiftieth of the traffic. On Stack Overflow, the page answered a plain request with a 403 while the official API served the same question data anonymously, 300 calls a day.
Four Ways to Get the Data
Every “web scraping vs API” decision actually has four candidates. Scraping reads the rendered page, the official API returns documented JSON under a quota, the internal API is the undocumented endpoint the page itself fetches data from (often callable directly), and the third-party data API is a provider’s maintained scraper for one target, sold as documented JSON per request.
| Criterion | Scraping the page | Official API | Internal API of the page | Third-party data API |
|---|---|---|---|---|
| Availability | Any public page | Only where the site ships one | Wherever the page loads data over XHR | Where a provider covers the target |
| Format | HTML you parse yourself | Documented JSON | Undocumented JSON | Documented JSON |
| Exactness | What the page shows, often rounded | Exact values | Exact values, usually the page’s own source | The page’s values, parsed by the provider |
| Coverage | Everything a visitor sees | The fields the vendor chose to expose | What this page needs, sometimes more than the public API | The provider’s schema, built from the page |
| Auth and quotas | None, but anti-bot instead | Keys, rate limits, paid tiers | Session cookies or signed requests, no documented quota | Key, per-request billing |
| Blocking risk | Full anti-bot surface | Contractual (throttling, key revocation) | Between the two, and endpoints move without notice | On the provider’s side |
| Breakage | Layout changes | Versioned, deprecations announced | Silent changes any day | Layout changes, fixed by the provider |
| Terms | Site’s terms apply | Explicit API terms | Usually the same site terms, read them | Provider terms, and the site’s terms still govern the data |
| Maintenance | Selector fixes | Version migrations | Re-discovery when the endpoint changes | Schema updates on the provider’s timeline |
| Best for | Data no API exposes | Long-lived integrations | JS-rendered pages without a public API | Covered targets past the official quotas |
On the live targets, the rows that decided the choice were quotas, exactness, and blocking, and the measured examples below put numbers on all three.
When the Official API Is Enough
An official API wins on stability and exactness, and it bills those advantages in quotas. The two we measured live make the range concrete. GitHub’s REST API allows 60 unauthenticated requests per hour (we read the live counter from its rate_limit endpoint mid-run) and 5,000 per hour with a token per its documentation. The Stack Exchange API answered anonymously with a quota_remaining field right in the response body, counting down from 300 requests per day.
One request against a documented endpoint replaces a page fetch and the parser behind it:
import requests
r = requests.get("https://api.github.com/repos/psf/requests", timeout=30)
repo = r.json()
print(repo["stargazers_count"], repo["open_issues_count"], repo["license"]["spdx_id"])The response carries 86 top-level fields, including exact created_at and pushed_at timestamps the page shows only as relative dates, and an exact subscribers_count where the page rounds to “1.3k watching”. Pagination is documented (per_page caps at 100 on most list endpoints, with Link headers pointing at the next page), errors are status codes with JSON bodies, and none of it changes when the site redesigns.
The quota ladder is where planning happens, and it differs per vendor more than any other property:
| Target | Anonymous | With a free key or token | Above that |
|---|---|---|---|
| GitHub REST API | 60 requests/hour (measured live) | 5,000/hour per its docs | GitHub Apps quotas |
| Stack Exchange API | 300 requests/day (measured live) | 10,000/day per its docs | negotiated access |
| A page’s internal API | no documented quota | not applicable | anti-bot decides |
| Scraping the page | no quota | not applicable | anti-bot decides |
| A third-party data API | key required | trial tiers vary by provider | per-request billing |
GitHub also attaches the live counter to every response as the X-RateLimit-Remaining header, so a client can check its budget without a separate call. Build that check in from day one.
The costs are the mirror image. The vendor picks the fields, so an API can lag the page or omit what you actually came for. Commercial catalogs in particular tend to expose inventory but not prices, reviews, or ranking positions. Quotas turn into paid tiers exactly at the volumes where scraping starts to look tempting. And the key ties every request to your account, which makes the terms of service enforceable in a way an anonymous page fetch is not.

Whether those trade-offs survive contact with a real target is what we measured next.
What the API Returns vs What the Page Shows
GitHub’s page rounded the star count on 10 of 10 repositories we checked, and the API returned the exact integer every time, in 69 KB against 3.6 MB of page weight. The method was one request per side per repository, the page parsed with BeautifulSoup against the API record for the same entity, and here is the full comparison:
| Repository | Page shows | API returns |
|---|---|---|
| psf/requests | 54.3k stars | 54,278 |
| pallets/flask | 72.2k stars | 72,169 |
| django/django | 89.6k stars | 89,566 |
| numpy/numpy | 32.6k stars | 32,642 |
| pandas-dev/pandas | 49.6k stars | 49,622 |
| scrapy/scrapy | 64.2k stars | 64,167 |
| python/cpython | 75.6k stars | 75,593 |
| microsoft/playwright | 95.5k stars | 95,524 |
| puppeteer/puppeteer | 95.5k stars | 95,536 |
| SeleniumHQ/selenium | 34.5k stars | 34,456 |
GitHub’s markup happens to keep the exact number in a title attribute, so a scraper can dig it out, and that is the kind of markup trivia an API user never needs to know:
# pip install requests beautifulsoup4
import requests
from bs4 import BeautifulSoup
page = requests.get("https://github.com/psf/requests",
headers={"User-Agent": "Mozilla/5.0"}, timeout=30)
star_el = BeautifulSoup(page.text, "html.parser").select_one("#repo-stars-counter-star")
api = requests.get("https://api.github.com/repos/psf/requests", timeout=30).json()
print("page shows:", star_el.get_text(strip=True)) # 54.3k
print("page title attr:", star_el.get("title")) # 54,278
print("api returns:", api["stargazers_count"]) # 54278Three lines of parsing, one selector to maintain, and the answer still depends on GitHub keeping that attribute, against one documented field on the API side.
The Stack Overflow side of the measurement made the blocking column concrete. The question page returned 403 to a plain requests call with a browser User-Agent, while the official API served the same question’s data anonymously. One response carried 15 fields, an exact view count of 1,998,112 on a question whose page our client never got to see, and the quota counter. For that target, the API is the only route that works without anti-bot tooling.
Both of our live targets are developer platforms with generous APIs. The pattern inverts on many commercial targets, where the public API exposes a trimmed subset and the page carries prices, reviews, and availability the API omits. That gap is what the internal endpoint often covers.
Finding the Internal API
Pages that render data with JavaScript get that data from somewhere, and that somewhere is an internal API you can usually call directly. Open DevTools, switch to the Network tab, filter by Fetch/XHR, reload the page, and look for responses whose JSON contains the values you see rendered. The request URL, its parameters, and its headers are all copyable from the same panel.
We measured the payoff on a sandbox page built for exactly this pattern. The initial HTML of scrapethissite.com’s AJAX page contains none of the film data it displays (we checked the raw response for the first rendered title, zero matches), so scraping it head-on requires a browser. The endpoint the page calls returns everything at once:
import requests
r = requests.get(
"https://www.scrapethissite.com/pages/ajax-javascript/",
params={"ajax": "true", "year": 2015},
timeout=30,
)
films = r.json()
print(len(films), films[0])The endpoint answered in 0.67 seconds with 1.6 KB of JSON, 16 films across a 5-field schema (best_picture appears only on the winner), against a 13 KB HTML shell that carries no data at all. The year parameter is also the endpoint’s pagination, so walking the archive is a loop over years, not over rendered pages. No browser, no rendering wait, no selectors, and the response is the page’s own source of truth, the same payload the visitors’ page renders.
Internal endpoints sit between the other two options on risk. They are undocumented, so they change without notice and carry no quota you can plan around, and the site’s terms of service apply to them the same way they apply to the page. Some sites version them, sign the requests, or check session cookies, and how much of that you meet is itself a signal of how much the operator minds. Treat a found endpoint as a fact about today’s page, not a contract.
Third-Party Data APIs
A fourth route runs through a provider. Third-party data APIs are maintained scrapers behind an API contract. The provider keeps the parsers, proxies, and rendering current for one target and sells the output as documented JSON, priced per request. The category exists for the two cases the first-party routes leave open.
The first case is a target with no official API at all. Zillow exposes no public listings API, so for an integration the choice is your own parser against a Zillow data API that returns listings and property records as JSON. The second case is an official API that answers a different question than the page. Google’s Custom Search JSON API allows 100 free queries per day with a hard ceiling of 10,000 per day, and it queries a configured search engine rather than returning the live SERP with its features. Rank tracking and SERP research need the page Google actually serves, so SERP APIs exist as their own category, and HasData’s Google SERP API returns that live SERP as structured JSON.
The trade-offs combine the columns on either side of it in the table above. Exactness and format behave like an official API, since the provider delivers parsed fields. Coverage and breakage behave like scraping, because underneath it is scraping: the schema follows the page, and a layout change breaks the provider’s parser until the provider’s fix restores it. The blocking surface moves off your infrastructure, and the terms question stays put, since the target site’s terms govern the data no matter who fetches it. What you give up is control of the schedule, because a changed schema or a deprecated field arrives on the provider’s timeline rather than yours.
When Scraping Is the Answer
Scraping remains the route for everything the previous three options exclude, plus every target no provider covers. Sites with no API, fields the API omits, and pages where the data exists only in rendered HTML all end here. It reads what a visitor sees, so its coverage is complete by definition, and everything else about it is maintenance. Selectors follow the layout, pagination means walking the next-page links yourself, dynamic content needs rendering or the internal endpoint above, and the anti-bot surface is all yours to handle. Our Stack Overflow probe is what that looks like in practice, a 403 before any of the question’s HTML arrives.

The build-vs-buy line inside scraping is where a web scraping API comes in. It takes a URL and returns the rendered HTML or extracted JSON, with proxies and JS rendering on the service side, billed per successful request. That keeps the coverage of scraping while outsourcing the two costs that grow with scale, infrastructure and blocking. Which service fits which workload is a separate comparison, and we keep one maintained in the web scraping API roundup.

With the three routes compared and measured, the terms themselves take one section to pin down.
What Web Scraping and API Scraping Mean
Web scraping is the automated extraction of data from web pages. Fetch the HTML, parse it, keep the fields. An API (Application Programming Interface) is a documented request-response contract a service exposes for programs. You send a request to an endpoint URL and receive structured data, usually JSON. API scraping is simply collecting data through such endpoints at scale, official or internal, instead of parsing pages.
In the everyday sense, using an API does not count as web scraping. An API returns data meant for programmatic consumption, while scraping extracts data from markup meant for people. The distinction matters legally and operationally, because API access is governed by explicit terms and keys, while page access is governed by the site’s general terms, and where scraping stands legally is its own topic.
Conclusion
The measured picture matches the folk wisdom only halfway. On the two developer platforms we measured, the official API won every axis, but the vendor decides what it carries and the quota decides how fast you can take it, and both are checkable before writing any code. Compare the API’s field list with the page for one entity you care about, divide your daily volume by the free quota, and pick per target. The internal endpoint and the parser cover the targets that fail either check, a third-party data API covers them without the maintenance wherever a provider already supports the target, and a scraping API absorbs the blocking work when you keep the parser yourself.


