HasData
Back to all posts

How HasData Scrapers Achieve 99.9% Uptime at Millions of Requests

At HasData, reliability is something we engineer, monitor, and validate every single day (and we mean it). In this post, we’ll show you the systems behind our 99.9% uptime: synthetic testing, monitoring dashboards, proxy health checks, and infrastructure choices.

Synthetic Tests Validate the APIs Every Day

We continuously run synthetic tests across all our APIs. Several times per day, each API is exercised with at least 10 parameter variations.

Take our Google SERP API for example. A q=coffee query with a location set should come back with at least 7 organic results, a knowledge graph, a local pack, related questions and pagination. The same query with no location carries neither the local pack nor the related questions, so each variation has its own expected shape rather than one template for all of them.

We validate each block individually. For organic results, for instance, we check that every entry includes a link, title, and snippet. If anything falls short, we know immediately.

All results flow straight to Slack, so the team is alerted before customers ever feel the impact.

Synthetic test failure for Google SERP Ad Results check showing a missing required field error

The second kind of alert is the same check on a mobile layout, where several fields went missing at once:

Synthetic test failure for Google SERP Mobile Recipes Results check showing three missing required fields

Each message names the exact check and the fields it expected, so nobody starts the morning by searching for what broke.

Monitoring Dashboards for Success Rates and Latency

Synthetic tests catch regressions, but real-time visibility into production traffic is just as important. Every API has two key dashboards:

  • The success and failed requests chart tracks the ratio of successful responses to failures.
  • The latency chart measures p50, p80, p90 and p99.

If failures or p99 latency increases, alerts go to our monitoring channel. From there, engineers can drill into the exact request ID, with full logs and cross-service traces.

Google SERP API Grafana dashboard showing request volume, success and failed ratio, latency percentiles, and trace logs

The dashboard above is the Google SERP API on an ordinary day, and these are the numbers it holds:

What the charts trackThe day above
p50 latencyabout 2.4 s, flat across the window
p80 latencyabout 4.3 s
p90 latencyabout 4.9 s, with one rise to 5.8 s
p99 latency8.9 to 10 s, with one spike to 11.4 s and a second to 10.8 s
Failed traces logged2, at 11.5 s and 11.6 s
Successful trace durations1.2 to 2.5 s

Averages would hide most of that. The p99 line is the one that moves first when something degrades, which is why alerts key on it.

Proxy Health, the Hidden Layer

Much of our success rate comes down to the proxy networks powering our APIs. We monitor them just as closely as the APIs themselves:

  • Success rate per API per retry
  • Traffic volume per proxy
  • Median response size of scraped pages

If a proxy underperforms, we isolate and replace it before it affects users.

Proxy health dashboard showing total data processed, API attempts funnel by proxy type with success rates, and top proxies by GBs

First-attempt success runs between 76% and 91% depending on the proxy type, and that spread is where the routing decisions come from.

Infrastructure and Observability

Our APIs run on a self-hosted Kubernetes cluster. Managing our own infrastructure gives us the control we need for performance and scaling.

  • Dedicated servers run our database instances
  • Dedicated servers power the monitoring stack
  • Grafana + Prometheus track metrics across the system
  • ClickHouse stores traces and logs for high-volume analysis

This setup scales to millions of requests while keeping costs predictable and visibility high.

Kubernetes cluster dashboard showing 25.3 CPU cores in use, 141 GiB of RAM against 704 GiB total, 26 nodes and 370 running pods

The same Grafana that watches the APIs watches the cluster underneath them.

How It Works in Practice

Here’s how these systems connect when an issue occurs:

  1. A synthetic test fails and alerts the team in Slack
  2. Latency charts confirm a spike in p99 latency
  3. Proxy monitoring shows a drop in success rate for one of the proxies
  4. Engineers reroute traffic, replace the failing proxy, and confirm resolution

The loop from detection to fix is fast and transparent because we monitor and trace every layer, and the request ID carries through all four steps.

Why This Matters to Our Customers

Everyone claims “99.9% uptime.” For us, it’s a 24/7 engineering process.

  • Failures are caught early, often before they reach production scale
  • Latency is continuously tracked across percentiles, not just averages
  • Infrastructure is built for scale, instead of just a minimum viable setup.

With HasData, you don’t have to wonder if your requests will succeed. That is what the monitoring, testing, and infrastructure above are for.

Roman Milyushkevich
Roman Milyushkevich
Roman Milyushkevich is the Co-founder and CTO at HasData, a web scraping API handling billions of requests. He designs the distributed systems, proxy infrastructure, and APIs behind large-scale, reliable data extraction. Roman writes on API design, browser automation, and building scraping pipelines that hold up in production.
Articles

Might Be Interesting