Proxy Index
Scoring lensApplied across rankings and provider pages
Preset

How we score proxy providers

The complete definition of the baseline gate, five scored dimensions, price context, ranking thresholds, and missing-evidence rules.

6
Gate checks
5
Scored dimensions
Per report
Published vantage detail
Declared
Window per report
Canonical proxy scorecard v1
Scorecard version

§1 · The gate and the dimensions

We do not roll everything into one number. A provider first clears an eligibility gate: the basic plumbing, pass or fail. It then earns a grade on the five scored dimensions. Price is shown as context and never graded. Only the five weighted dimensions enter the Proxy Index.

The gate: benchmarkable, or not

Six commodity checks decide whether a proxy is worth grading at all: can it connect, does traffic leave through the proxy, does HTTP work, is the exit ASN sane, does the country match, and can it move a real response body. These are pass/fail, not scored. Against unprotected targets almost every serious provider clears all six — which is exactly why scoring them told buyers nothing. So we gate on them and move on.

Gate
benchmarkable = connection ∧ exit identity ∧ http delivery ∧ asn lookup ∧ geo lookup ∧ throughput
Reported as one row: benchmarkable, passed 6/6 baseline checks. It is not a number. A proxy that fails the gate is marked not benchmarkable and its dimensions are suppressed.
Each check needs at least 30 comparable observations and a raw pass rate of at least 99%. The verdict is derived from numerator and denominator before the displayed rate is rounded.

The five scored dimensions, and price as context

Dimensions are graded 0–100 from inputs we already collect, and each carries its inputs in the open so the grade is legible rather than asserted. The Proxy Index weights workload success 3× and speed, pool quality, trust & sourcing, and session 1× each. Price is shown as context and carries no weight. If any positively weighted dimension is unavailable, no Proxy Index score exists; we never drop it, renormalize the rest, or invent a number.

  • Workload successworkload_success

    FED BY fixture success rates against real target classes, blended with browser-engine success where present.

    SCORE ROLE 3× Proxy Index weight

    WHY NOT 100 real and protected targets fail some fraction of the time for everyone; a residential pool medians around the low 90s, datacenter far lower.

  • Speedspeed

    FED BY median and p95 latency to three targets: our own test pages, a CDN edge, and a fixed European origin. Each target is read off one absolute log curve (100 at 100ms or faster, 0 at 10s or slower), its median and p95 blended evenly, and the targets measured this window carry equal weight. Averaged with throughput where measured.

    SCORE ROLE 1× Proxy Index weight

    WHY NOT 100 the curve is absolute, so a 100 needs a fast median and a fast tail. Latency is bounded by physics and upstream device quality, and the tail is where jobs actually wait: a multi-second p95 pulls the grade down however good the typical request looks, because that is the request the queue is stuck behind.

  • Pool qualitypool_quality

    FED BY unique exits per attempt, repeat rate, top-exit-IP share, and exit-IP entropy from the observed exits.

    SCORE ROLE 1× Proxy Index weight

    WHY NOT 100 advertised pool size is marketing; realized uniqueness and subnet spread are what a target sees, and recycling drags them down.

  • Trust & sourcingtrust_sourcing

    FED BY source-labelled IP reputation risk share, usage-type honesty (claimed vs observed), and city-level geo precision once enrichment coverage is sufficient.

    SCORE ROLE 1× Proxy Index weight

    WHY NOT 100 reputation is vendor-calibrated and can be contaminated by other customers; we grade only when at least five unique exits have 80% coverage from one labelled source.

  • Sessionsession

    FED BY sticky-session hold rate and duration through the advertised window.

    SCORE ROLE 1× Proxy Index weight

    WHY NOT 100 sessions degrade as upstream residential devices go offline; a long hold depends on hardware the provider does not own.

  • Valuevalue

    FED BY the list price as the provider states it, shown as context and never graded.

    SCORE ROLE context only · 0× weight

Worked example · grading one dimensionworkload_success

Say a proxy runs 432 workload attempts and 192 succeed, and its browser-engine pass rate is 0%. Workload success is the blend of those two measured rates — here roughly a 22, not a 100 — because the grade is built from what actually completed against real targets, not from whether the socket opened.

request success 192/432 · browser-engine success 0% → workload success ≈ 22 / 100

How one dimension grade is actually built. The inputs travel with the grade, so a reader can see why the number landed where it did — and why a clean gate does not buy a clean dimension.

Worked example · why speed is not one numberspeed

Latency is scored on one fixed log curve: 100ms or faster is a 100, 10s or slower is a 0, and everything between falls along a straight line in log space. Log space because latency is multiplicative: 100ms to 200ms is the same felt step as 2s to 4s. Each target is read twice, at the median and at the p95, and the two are blended evenly.

Say a proxy answers our own test pages at a 42ms median and a 780ms p95. The median is faster than the 100ms top of the curve, so it scores 100; the p95 sits mid-curve at 55. Blended, that target is a 78. The CDN edge and the European origin were not measured this window, so neither takes a share and latency stays 78. Averaged with a throughput sub-score of 100, speed is 89: a proxy that moves plenty of bytes and still keeps some jobs waiting.

requests to our own pages: median 42ms → 100 · p95 780ms → 55 · request latency = (100 + 55) / 2 = 78 · edge and origin not measured · latency = 78 · throughput 150 Mbps → 100 · speed = (78 + 100) / 2 = 89 / 100

One percentile would not be enough. Read alone, the median calls this proxy a 100 and the p95 calls it a 55; the grade is neither, and both numbers ship with it so you can see which half you care about. A single averaged latency figure would hide the gap entirely.

Consistency is part of the grade

We run rolling 30-day windows, so a dimension is scored across daily buckets rather than in one lucky afternoon. A provider that holds steady reads differently from one that spikes and decays, and that stability is a property of the grade — not a footnote to it.

On the legacy compositeThe old single composite score is being retired. It graded only the commodity gate checks, so it clustered near 100 and hid every measurement that mattered. We still retain it as explicitly labeled diagnostic history, but it cannot supply the Proxy Index, rank, grades, structured data, or leadership claims.
§2 · MEASUREMENT

Declared windows and explicit sample evidence.

Every report names its own evidence window, and rank calculations never mix window kinds. Current payloads expose aggregate gate numerators and denominators, confidence, contributing-run counts, endpoint counts, and critical coverage warnings. Failures caused by our own benchmark infrastructure, not the provider, are excluded from published pass rates and reported on each test record. Missing facts stay missing.

Published vantage latencynot exposed
The current public feed does not expose a per-vantage latency breakdown. Report pages show only the aggregate latency evidence actually published.

No city, node count, or per-vantage sample size is inferred from aggregate data.

  • 30comparable observations required before one baseline check can be evaluated.
  • ≥2eligible benchmark runs must contribute before a strict score may receive a rank.
  • ≥80%confidence is required for ranking; lower-confidence strict scores are labeled diagnostic.
  • 5 / 8minimum endpoints: five for datacenter, ISP, and unknown; eight for residential, mobile, and mixed — with zero critical coverage warnings.

§3 · What each fixture is

Workload success is measured against pages we build and serve ourselves, in two kinds. Plain pages hold difficulty constant: the same set is configured for every provider under the same pass checks, so a result there says the provider fetched the page and brought it back intact, and the target is not what separates providers. Defended pages are hostile on purpose, pages we host that refuse the wrong kind of traffic, so we can measure how often a provider gets blocked. Their results are shown on a provider page but not folded into the grade in this phase. Per-fixture results are published as aggregates, success rate and latency.

Plain pages

Each plain fixture is run both as a raw HTTP request and through three browser engines (Playwright Chromium, Camoufox, and Patchright), and the two kinds of run are checked differently. The browser runs check the expected status, the required elements, the required text, and a body at least as long as that fixture's floor, and scan the same response for refusal: block status codes, block and CAPTCHA selectors and URL patterns, and challenge text. The raw HTTP run checks the expected status only. Those browser runs record a blocked or challenged response as that, not as a generic failure, and the split shows whether a provider carries browser traffic as well as raw requests. A browser run stops its clock when the page reports its content ready, the domcontentloaded wait, rather than when the last image arrives. That page-load time is published on its own and never folded into the speed grade, which reads raw HTTP requests only.

  • Product pagefixture-ecommerce-product

    One product detail page with a breadcrumb, a product card, a price, and a stock status. It stands in for reading a single item's price and availability.

    PASSES WHEN status 200, title contains Fixture Storefront, product card and price selectors present, the text In stock present, body text at least 250 characters.

  • JSON APIfixture-json-workload

    A JSON endpoint with a fixed schema and six item records, registered as a plain HTTPS target. It stands in for API scraping, and it is the one fixture with no markup to validate against.

    PASSES WHEN status 200, the payload's fixture identifier present, body at least 100 characters, no selector checks because there is no DOM.

  • Heavy HTML pagefixture-large-html

    One page built from 160 repeated text sections, registered as a download target rather than a page target. It stands in for work whose cost is bytes rather than structure, so it is where transfer behavior shows up.

    PASSES WHEN status 200, title contains Fixture Large HTML, the body container present, body text at least 60,000 characters.

  • Category listingfixture-listing

    A grid of twelve listing cards with identical shape. It stands in for pulling the same fields out of many repeated rows.

    PASSES WHEN status 200, title contains Fixture Listing Page, grid and card selectors present, the page heading Fixture Marketplace Listings present, body text at least 900 characters.

  • Search resultsfixture-serp

    A search page with a query box and eight repeated result cards, each linking on to another fixture. It stands in for collecting a ranked result list and following it onward.

    PASSES WHEN status 200, title contains Fixture Search Results, the result list container present, the first result's text present, body text at least 500 characters.

Defended pages

Each defended fixture refuses one thing a raw request cannot get past but a real browser can: a per-client request budget that answers 429 once you cross it, an inline compute challenge that only script execution clears, or a redirect chain that only a browser follows to the end. We host every one of them, so the difficulty is the same for every provider and a block is a real block, not a third party having a bad day. A 429 or an unsolved challenge is recorded as blocked or challenged, which is the signal we want. These results are shown on a provider page and left ungraded in this phase.

  • Rate limit disciplinedefended-rate-limit

    A page that counts each request against a fixed per-window budget, keyed on the exit IP the request arrives from. Requests from one exit IP cross the budget and come back 429, which scores as blocked; spreading the same requests across rotating exit IPs stays under budget and keeps the 200 page. It stands in for a site that rations requests per client and refuses the ones over its limit.

    PASSES WHEN status 200, title contains Fixture Rate Limit Budget, the defended-rate-limit selector present, the text Request served within the rate-limit budget present.

  • Compute checkdefended-compute

    A page that hides its real content behind an inline compute challenge. A browser engine runs the challenge script, which clears the running marker, reveals the protected body, and stamps a solved flag on it; a raw fetch cannot run the script, so it keeps the unsolved marker and fails. It stands in for a compute-gated page that only real browser execution can open.

    PASSES WHEN status 200, title contains Fixture Compute Check, the challenge-solved marker and the protected content and body selectors present, the text Access granted after the compute challenge present.

  • Redirect trapdefended-redirect

    An entry page that answers with a 302 and bounces through a short chain of redirects before a 200 landing page. A browser engine follows the chain to the end; a raw fetch does not follow redirects, so it stops at the first hop and fails, and this fixture runs through browser engines only for that reason. It stands in for a site that serves its real content only after a redirect a real browser would follow.

    PASSES WHEN status 200 on the landing page, title contains Fixture Redirect Landing, the defended-redirect selector present, the text You followed the redirect chain to the end present.

What a fixture cannot tell youHolding the target constant is what makes providers comparable, and it is also the limit. Defended pages measure blocking against pages we run, so they show how a provider handles a rationed budget, a compute challenge, or a redirect chain, but not how it fares against the specific site you care about, which sets its own rules. Run your own targets before you scale on these numbers.
REFERENCE · CHECK FAMILIES

What each measured family checks — and what a single number hides.

Each family below is one input the eligibility gate and the graded dimensions are built from — not a score of its own. For each: the exact ratio we compute, why it earns its place, and the detail a single headline number can't show — drawn live from the current cohort.

§4

Connection

Core · since v1.0

A proxy that cannot establish a connection cannot be used by any higher-level workflow.

weight0.18
weight share18%
Legacy family diagnostic
successful attemptseligible attempts
Weight 0.18 · retained as diagnostic history, never used as a Proxy Index fallback.

Checks whether the proxy accepts a basic TCP connection.

Scores are normalized per provider and compared within the issue cohort.

WHAT THE SINGLE NUMBER HIDES

Connection failure blocks every other signal.

Connection pass rate is the base layer. If a proxy cannot establish a TCP session, higher-level HTTP, geo, ASN, and session checks never get a chance to run.

§5

Exit identity

Core · since v1.0

Exit identity proves traffic is actually routed through a proxy and gives downstream checks a stable observed IP.

weight0.22
weight share22%
Legacy family diagnostic
successful attemptseligible attempts
Weight 0.22 · retained as diagnostic history, never used as a Proxy Index fallback.

Checks whether the proxy returns a usable exit IP through an IP echo target.

Scores are normalized per provider and compared within the issue cohort.

WHAT THE SINGLE NUMBER HIDES

No observed exit means no trustworthy proxy evidence.

Exit identity proves traffic actually leaves through a proxy. Without a usable observed IP, downstream classification and targeting metrics become unreliable.

§6

HTTP reachability

Core · since v1.0

HTTP reachability is the closest low-cost signal for whether the endpoint can carry normal web traffic.

weight0.18
weight share18%
Legacy family diagnostic
successful attemptseligible attempts
Weight 0.18 · retained as diagnostic history, never used as a Proxy Index fallback.

Checks whether a proxied HTTP request reaches the benchmark target successfully.

Scores are normalized per provider and compared within the issue cohort.

WHAT THE SINGLE NUMBER HIDES

TCP success does not guarantee usable web traffic.

A reachable socket is not enough. HTTP reachability catches endpoints that connect but cannot carry normal web requests through benchmark targets.

§7

Throughput

Core · since v1.0

Large pages, media assets, and browser sessions need proxies that can move more than tiny responses.

weight0.10
weight share10%
Legacy family diagnostic
successful attemptseligible attempts
Weight 0.10 · retained as diagnostic history, never used as a Proxy Index fallback.

Checks whether the proxy can complete bounded download samples against controlled targets.

Scores are normalized per provider and compared within the issue cohort.

WHAT THE SINGLE NUMBER HIDES

Tiny responses hide slow transfer paths.

Small health checks can pass while heavier pages stall. Throughput checks whether the proxy can move realistic response bodies, not just connect.

§8

Session persistence

Added in v2.0

We added session stability once it became clear that geo and ASN scores said nothing about whether a multi-step flow - login, cart, checkout - would survive.

weight0.07
weight share7%
Legacy family diagnostic
successful attemptseligible attempts
Weight 0.07 · retained as diagnostic history, never used as a Proxy Index fallback.

Checks whether sticky-session follow-up requests keep the same observed exit identity.

Sticky sessions must hold through the advertised threshold without rotation or re-auth.

WHAT THE SINGLE NUMBER HIDES

Two pools, same pass-rate, different deaths.

A single pass-rate cannot show the shape of failure. Some pools hold perfectly and then cliff the instant the advertised window closes; others bleed sessions steadily from the first minute.

§9

Geo accuracy

Core · since v1.0

Geo is the most-requested targeting dimension and the easiest to fake - a pool can sell US residential while quietly announcing from a Frankfurt datacenter. This area catches the gap between the label on the tin and where the IP actually sits.

weight0.14
weight share14%
Legacy family diagnostic
matching geo lookup attemptscomparable executable provider-quality geo lookup attempts
Weight 0.14 · retained as diagnostic history, never used as a Proxy Index fallback.

Checks whether observed country matches the declared country when geo metadata is available.

Country-level disagreement is flagged when comparable lookups fall below the issue threshold.

WHAT THE SINGLE NUMBER HIDES

Country is easy. Cities are where pools split.

The headline percentage is country-weighted, so it reads high almost everywhere. The moment you target a region or a city the field spreads out - and that precision drop never shows up in the single number.

§10

ASN accuracy

Core · since v1.0

Modern anti-bot systems score the announcing ASN before they ever look at behavior. A residential IP that announces from a hosting network is flagged on arrival, so a clean ASN profile is doing quiet work on every request.

weight0.11
weight share11%
Legacy family diagnostic
matching ASN lookup attemptscomparable executable provider-quality ASN lookup attempts
Weight 0.11 · retained as diagnostic history, never used as a Proxy Index fallback.

Checks whether the proxy exits from the declared autonomous system when ASN metadata is available.

Mixed hosting/residential ranges trigger caveat ASN-1 when cleanliness falls below threshold.

WHAT THE SINGLE NUMBER HIDES

A letter grade compresses the whole pool.

Two pools can both grade A while one is 99% genuine residential ASNs and the other blends in a hosting slice it would rather you did not see. The composition is what actually trips detection.

§11 · CAVEAT CODES

Editorial caveats and eligibility warnings are different.

Editorial caveat codes add context without changing rank order. A critical coverage warning is stricter: it makes the publication diagnostic-only and removes its rank until a valid immutable revision clears the warning.

CodeMeansEffect on score
GEO-3Country match high; sub-region match below threshold.None — informational only.
ASN-1ASN occasionally announces as hosting; mixed pool.None — informational only.
CAP-2Elevated challenge rate against one of three WAF vendors.None — informational only.
§12 · CHANGELOG

Versioned, dated, and never applied backwards.

Methodology changes are versioned and dated. Past results are not rescored under new versions; each dated evidence snapshot preserves the model used.

VersionEffectiveChange
Five representative workloads2026-08-26Expanded Workload success from two controlled browser jobs to five. Product and search extraction now run alongside JSON API, large-page, and listing extraction. Each job must return the expected structured output, and the five completion rates carry equal weight.
vNext dimension grades2026-08-25Activated provider-level vNext grades after 25 eligible observations. Availability, Workload success, IP quality, and Session control use observed rates. Speed scores median and p95 job completion on a fixed one-to-ten-second log curve. Existing rankings keep their recorded comparison model until another provider has eligible vNext grades.
Five-dimension vNext contract2026-08-24Defined the additive vNext evidence contract for Availability, Workload success, Speed, IP quality, and Session control. Existing publications and rankings keep their recorded model. The new model remains inactive while its measurements run in shadow mode.
Graded vantage region2026-08-22Added the graded vantage region to series vantage policy evidence, naming the observed worker region that grades are measured from.
Latency measured honestly2026-08-17Latency is now measured as three named populations, not one blended number: requests to our own pages, a CDN edge, and a fixed European origin. Only successful requests count toward each figure, and an all-attempts check publishes alongside it so scheduler waits and download time no longer get counted as latency. Speed grades moved up to reflect the corrected measurement, and defended pages now report real block data instead of a placeholder.
Speed latency log curve2026-08-07Introduced proxy_index_speed_latency_log_v1 for the speed dimension's latency sub-score. Median and p95 latency are each scored on one absolute log curve, 100 at 100ms or faster and 0 at 10000ms or slower, and blended 50/50; the throughput sub-score and the 50/50 split between latency and throughput are unchanged. The superseded banding read p95 alone and hit 0 at 2000ms, which put every measured provider at exactly 0 latency and 50 speed, so the dimension could not separate anyone. Immutable publications keep their recorded latency banding and are never rescored; each one names the policy that produced it in the speed dimension's gradedBasis. Until each line publishes again on its normal daily cycle, current rows may therefore mix banded and log-curve speed grades, and cross-provider speed comparisons should be read against each row's stated gradedBasis. The comparison model proxy_index_dimension_score_v1 and its 3·1·1·1·1 weights are unchanged.
Provider-quality failure accounting2026-07-30Versioned provider-quality failure accounting as provider_quality_excludes_benchmark_infrastructure_v1. Public success excludes only browser_service_error, config_error, and worker_error failures, reports a null success rate when no provider-quality attempts exist, and exposes raw failure totals plus a fixed-key exclusion receipt and infrastructure failure rate. Ambiguous browser_error and benchmark_target_error failures remain provider-quality evidence.
Universal baseline gate v22026-07-30Introduced proxy_index_baseline_gate_v2. Its six stable checks now measure universally available execution reachability, including successful ASN and country lookup enrichment. Expected-versus-observed ASN and country accuracy remain conditional metrics when a benchmark supplies an expected claim. Immutable v1 publications retain their recorded accuracy-gate semantics and are never reinterpreted as v2.
Strict public comparison contract v12026-07-29Unified public comparison on proxy_index_dimension_score_v1 with proxy_index_baseline_gate_v1. A line receives a comparison score and rank only when all six baseline checks pass at 99% or better with at least 30 comparable observations each, the public sample thresholds pass, and every positively weighted 3·1·1·1·1 dimension is graded. Country and ASN checks use expected-versus-observed accuracy; missing evidence is not renormalized. The former family-health composite remains labeled diagnostic history.
Gate and graded dimensions2026-07-10Added an additive reporting layer: an eligibility gate over the commoditized plumbing checks (connection, exit identity, HTTP, ASN, country geo, throughput reachability) and six graded dimensions (workload success, speed, pool quality, trust and sourcing, session, value) derived from existing metric snapshots. Dimensions grade only where inputs exist and report an honest not-graded status otherwise. The composite score and its weighting are unchanged.
Region and city composition2026-05-28Added observed region and city composition metrics for public rolling report geography breadth.
Series vantage policy2026-05-28Added structured series-level vantage policy evidence comparing declared worker regions against observed report regions.
Reputation risk metric2026-05-28Added high-risk IP share as a measured metric when reputation enrichment includes risk bands or risk scores.
Usage-type accuracy2026-05-28Added usage-type accuracy as a measured metric when exit enrichment includes residential, mobile, ISP, or datacenter classification.
Vantage and concurrency slices2026-05-28Added public time-series coverage for declared worker vantage-region and concurrency-tier success rates.
Browser-engine probe metrics2026-05-21Added browser-engine probe metric definitions for target-class success, blocks, challenges, semantic validation, and navigation latency.
Session persistence scoring2026-05-21Added session persistence as a scored family for sticky-session scraper workflows.
Throughput metrics2026-05-19Added public throughput metric definitions and chart-ready rolling time series from download sample checks.
Trust metadata2026-05-18Added public trust metadata: confidence tier, coverage warnings, sample design, methodology notes, and scraper-facing interpretation.
Public rolling report v22026-05-17Added public-safe scorecard labels, threshold meanings, metric definitions, and time-series explanations.
A note on rescoringWe do not rescore dated results under new methodology. This is intentional: a benchmark must be a record of what was measured at the time, under the rules of the time.

See the methodology applied

Method matters most where it meets data. See these rules applied to the current field, or browse the dated evidence they produced.