At HasData, reliability is something we engineer, monitor, and validate every single day (and we mean it). In this post, we’ll show you the systems behind our 99.9% uptime: synthetic testing, monitoring dashboards, proxy health checks, and infrastructure choices.
Synthetic Tests Validate the APIs Every Day
We continuously run synthetic tests across all our APIs. Several times per day, each API is exercised with at least 10 parameter variations.
Take our Google SERP API for example. A q=coffee query with a location set should come back with at least 7 organic results, a knowledge graph, a local pack, related questions and pagination. The same query with no location carries neither the local pack nor the related questions, so each variation has its own expected shape rather than one template for all of them.
We validate each block individually. For organic results, for instance, we check that every entry includes a link, title, and snippet. If anything falls short, we know immediately.
All results flow straight to Slack, so the team is alerted before customers ever feel the impact.

The second kind of alert is the same check on a mobile layout, where several fields went missing at once:

Each message names the exact check and the fields it expected, so nobody starts the morning by searching for what broke.
Monitoring Dashboards for Success Rates and Latency
Synthetic tests catch regressions, but real-time visibility into production traffic is just as important. Every API has two key dashboards:
- The success and failed requests chart tracks the ratio of successful responses to failures.
- The latency chart measures p50, p80, p90 and p99.
If failures or p99 latency increases, alerts go to our monitoring channel. From there, engineers can drill into the exact request ID, with full logs and cross-service traces.

The dashboard above is the Google SERP API on an ordinary day, and these are the numbers it holds:
| What the charts track | The day above |
|---|---|
| p50 latency | about 2.4 s, flat across the window |
| p80 latency | about 4.3 s |
| p90 latency | about 4.9 s, with one rise to 5.8 s |
| p99 latency | 8.9 to 10 s, with one spike to 11.4 s and a second to 10.8 s |
| Failed traces logged | 2, at 11.5 s and 11.6 s |
| Successful trace durations | 1.2 to 2.5 s |
Averages would hide most of that. The p99 line is the one that moves first when something degrades, which is why alerts key on it.
Proxy Health, the Hidden Layer
Much of our success rate comes down to the proxy networks powering our APIs. We monitor them just as closely as the APIs themselves:
- Success rate per API per retry
- Traffic volume per proxy
- Median response size of scraped pages
If a proxy underperforms, we isolate and replace it before it affects users.

First-attempt success runs between 76% and 91% depending on the proxy type, and that spread is where the routing decisions come from.
Infrastructure and Observability
Our APIs run on a self-hosted Kubernetes cluster. Managing our own infrastructure gives us the control we need for performance and scaling.
- Dedicated servers run our database instances
- Dedicated servers power the monitoring stack
- Grafana + Prometheus track metrics across the system
- ClickHouse stores traces and logs for high-volume analysis
This setup scales to millions of requests while keeping costs predictable and visibility high.
.2r0BFlvR_Pg0o5.png)
The same Grafana that watches the APIs watches the cluster underneath them.
How It Works in Practice
Here’s how these systems connect when an issue occurs:
- A synthetic test fails and alerts the team in Slack
- Latency charts confirm a spike in p99 latency
- Proxy monitoring shows a drop in success rate for one of the proxies
- Engineers reroute traffic, replace the failing proxy, and confirm resolution
The loop from detection to fix is fast and transparent because we monitor and trace every layer, and the request ID carries through all four steps.
Why This Matters to Our Customers
Everyone claims “99.9% uptime.” For us, it’s a 24/7 engineering process.
- Failures are caught early, often before they reach production scale
- Latency is continuously tracked across percentiles, not just averages
- Infrastructure is built for scale, instead of just a minimum viable setup.
With HasData, you don’t have to wonder if your requests will succeed. That is what the monitoring, testing, and infrastructure above are for.


