Every proxy benchmark eventually faces the same temptation: roll everything into one impressive-looking number. We went the other way. A proxy first has to clear six pass/fail plumbing checks. Each check needs at least 30 comparable observations and a raw pass rate of at least 99%. Only then can the five positively weighted buyer dimensions produce a Proxy Index score.
The gate: commodity plumbing, pass or fail
Connectivity, exit identity, HTTP reachability, ASN accuracy, country accuracy, and throughput reachability are commoditized across the industry. A provider that fails any of them has an operations problem, not a quality tradeoff — so these checks gate eligibility instead of contributing to a score. Passing the gate is table stakes, and we say so.
The dimensions: where providers actually differ
- Workload success — did real-target requests complete? This is the closest thing to the answer a buyer wants, so the default Proxy Index weights it 3×.
- Speed — median and p95 latency on one absolute log curve (100 at 100ms, 0 at 10s), blended 50/50, then averaged with observed throughput when present. A fast median cannot hide a stalling tail.
- Pool quality — breadth and churn of observed exits, not the number printed on a pricing page.
- Trust & sourcing — do the exits look like what the provider claims they are?
- Session — do sticky sessions actually hold the same exit across follow-up requests?
- Value — graded only when a normalized pricing snapshot exists, then reported as context with zero Proxy Index weight. No snapshot, no grade, and we say why.
A positively weighted dimension with missing evidence renders as "not graded" with the honest reason — never as zero, and never as an invented number. The default Proxy Index weights workload success 3× and speed, pool quality, trust, and session 1× each. A custom scoring lens can re-weight only ranking-eligible proxies.
What we refuse to do
No retroactive scoring edits, and no score when the gate or a required dimension is missing. A strict score may remain labeled diagnostic evidence when the sample is too thin to rank. The methodology page carries the full definition of every check and threshold, and current reports expose the aggregate counts behind each gate verdict.