RPC Latency Benchmarking: Measure the Workload, Not Just the Endpoint
Benchmark RPC latency with realistic methods, load profiles, p95 and p99, error rates and freshness checks. Use reproducible tests to choose node capacity.
Start with the requests that matter
A fast eth_blockNumber response shows that an endpoint can answer one lightweight request. It does not predict the performance of contract calls, historical-state reads or log queries. A useful RPC benchmark measures your application's methods from the regions where that application runs.
Define the acceptance criteria before testing: the method mix, request rate, concurrency, required history and tolerable errors. Include a freshness check. A response from a lagging node may be fast but unsuitable for the workload. Our blockchain RPC guide explains the different requests an endpoint may need to serve.
Keep the comparison reproducible
Use the same requests, block references and load profile for each candidate. Record the client configuration, test location, connection-reuse behavior and timeout settings. For repeatable historical tests, choose finalized blocks and verify that all candidates support the required data.
Measure warm and cold behavior separately when both matter. Repeatedly requesting the same historical result may benefit from caching. Opening a new connection for every request measures connection establishment as well as RPC execution. Neither test is inherently wrong, but neither should be presented as the other.
Start with a small functional check, then increase load within an agreed test plan. Test only infrastructure you own or have permission to benchmark. Avoid using public community endpoints as unrestricted load-test targets.
Report errors beside latency
Separate successful-request latency from failed requests and timeouts. Otherwise, a service that rejects difficult requests quickly can appear faster than one that completes them. Report attempted traffic, completed useful work, error rate and latency together, broken down by method.
Percentiles answer different questions. The median describes a typical observation; p95 and p99 describe the slower end of the measured distribution. Include sample counts and test duration. With only 100 successful observations, a p99 estimate is driven by roughly the slowest one or two measurements and should not be treated as a stable production characteristic.
State how the load was generated. A closed-loop test waits for work to finish before sending more, so offered traffic can fall when the endpoint slows down. An arrival-rate test schedules work independently of individual response times, subject to the load generator's capacity. The Grafana k6 documentation explains why these models can produce different results. Record achieved traffic and missed scheduling targets, not just the configured rate.
Use the results to choose capacity
Look for the point where more concurrency stops increasing completed work and starts increasing latency or errors. Compare that point with your expected peak workload. Test mixed traffic too: a backfill may meet its own throughput target while making live reads too slow.
A high p99 does not identify the cause by itself. Investigate client-side pauses, network variability, queueing, query complexity and storage behavior before concluding that dedicated infrastructure will solve the problem. Repeat important tests across relevant operating periods rather than selecting the best short run. After choosing capacity, use RPC endpoint monitoring to track ongoing health and test the same workload against your failover endpoint.
DTEAM dedicated nodes have no artificial request-rate limits or usage quotas, but their sustainable throughput still depends on hardware and workload. Review our dedicated RPC offering, then send us your network, application region and representative method mix. We can scope a node configuration against the workload you intend to measure, including networks not currently listed on our site.