§1 · The gate and the dimensions
We do not roll everything into one number. A provider first clears an eligibility gate: the basic plumbing, pass or fail. It then earns a grade on the five scored dimensions. Price is shown as context and never graded. Only the five weighted dimensions enter the Proxy Index.
The gate: benchmarkable, or not
Six commodity checks decide whether a proxy is worth grading at all: can it connect, does traffic leave through the proxy, does HTTP work, is the exit ASN sane, does the country match, and can it move a real response body. These are pass/fail, not scored. Against unprotected targets almost every serious provider clears all six — which is exactly why scoring them told buyers nothing. So we gate on them and move on.
The five scored dimensions, and price as context
Dimensions are graded 0–100 from inputs we already collect, and each carries its inputs in the open so the grade is legible rather than asserted. The Proxy Index weights workload success 3× and speed, pool quality, trust & sourcing, and session 1× each. Price is shown as context and carries no weight. If any positively weighted dimension is unavailable, no Proxy Index score exists; we never drop it, renormalize the rest, or invent a number.
- Workload successworkload_success
FED BY fixture success rates against real target classes, blended with browser-engine success where present.
SCORE ROLE 3× Proxy Index weight
WHY NOT 100 real and protected targets fail some fraction of the time for everyone; a residential pool medians around the low 90s, datacenter far lower.
- Speedspeed
FED BY median and p95 latency to three targets: our own test pages, a CDN edge, and a fixed European origin. Each target is read off one absolute log curve (100 at 100ms or faster, 0 at 10s or slower), its median and p95 blended evenly, and the targets measured this window carry equal weight. Averaged with throughput where measured.
SCORE ROLE 1× Proxy Index weight
WHY NOT 100 the curve is absolute, so a 100 needs a fast median and a fast tail. Latency is bounded by physics and upstream device quality, and the tail is where jobs actually wait: a multi-second p95 pulls the grade down however good the typical request looks, because that is the request the queue is stuck behind.
- Pool qualitypool_quality
FED BY unique exits per attempt, repeat rate, top-exit-IP share, and exit-IP entropy from the observed exits.
SCORE ROLE 1× Proxy Index weight
WHY NOT 100 advertised pool size is marketing; realized uniqueness and subnet spread are what a target sees, and recycling drags them down.
- Trust & sourcingtrust_sourcing
FED BY source-labelled IP reputation risk share, usage-type honesty (claimed vs observed), and city-level geo precision once enrichment coverage is sufficient.
SCORE ROLE 1× Proxy Index weight
WHY NOT 100 reputation is vendor-calibrated and can be contaminated by other customers; we grade only when at least five unique exits have 80% coverage from one labelled source.
- Sessionsession
FED BY sticky-session hold rate and duration through the advertised window.
SCORE ROLE 1× Proxy Index weight
WHY NOT 100 sessions degrade as upstream residential devices go offline; a long hold depends on hardware the provider does not own.
- Valuevalue
FED BY the list price as the provider states it, shown as context and never graded.
SCORE ROLE context only · 0× weight
Say a proxy runs 432 workload attempts and 192 succeed, and its browser-engine pass rate is 0%. Workload success is the blend of those two measured rates — here roughly a 22, not a 100 — because the grade is built from what actually completed against real targets, not from whether the socket opened.
request success 192/432 · browser-engine success 0% → workload success ≈ 22 / 100
How one dimension grade is actually built. The inputs travel with the grade, so a reader can see why the number landed where it did — and why a clean gate does not buy a clean dimension.
Latency is scored on one fixed log curve: 100ms or faster is a 100, 10s or slower is a 0, and everything between falls along a straight line in log space. Log space because latency is multiplicative: 100ms to 200ms is the same felt step as 2s to 4s. Each target is read twice, at the median and at the p95, and the two are blended evenly.
Say a proxy answers our own test pages at a 42ms median and a 780ms p95. The median is faster than the 100ms top of the curve, so it scores 100; the p95 sits mid-curve at 55. Blended, that target is a 78. The CDN edge and the European origin were not measured this window, so neither takes a share and latency stays 78. Averaged with a throughput sub-score of 100, speed is 89: a proxy that moves plenty of bytes and still keeps some jobs waiting.
requests to our own pages: median 42ms → 100 · p95 780ms → 55 · request latency = (100 + 55) / 2 = 78 · edge and origin not measured · latency = 78 · throughput 150 Mbps → 100 · speed = (78 + 100) / 2 = 89 / 100
One percentile would not be enough. Read alone, the median calls this proxy a 100 and the p95 calls it a 55; the grade is neither, and both numbers ship with it so you can see which half you care about. A single averaged latency figure would hide the gap entirely.
Consistency is part of the grade
We run rolling 30-day windows, so a dimension is scored across daily buckets rather than in one lucky afternoon. A provider that holds steady reads differently from one that spikes and decays, and that stability is a property of the grade — not a footnote to it.