← Back to the graph01 / 10 · tool

Free proxies that actually work.

Proxy Scraper: free proxies that actually work – 700+ sources, five checks, a fresh list every hour.
Year
2026
Role
Author & maintainer
Stack
Python, asyncio, GitHub Actions, Prometheus
Links
Live site ↗Source on GitHub ↗

Overview

Scrapes HTTP, SOCKS4 and SOCKS5 proxies from 700+ sources and keeps only the ones that pass real checks – re-verified every hour by GitHub Actions.

Free proxy lists are mostly noise: dead hosts, honeypots and proxies that quietly inject scripts into the pages they serve – about one in five does. Proxy Scraper collects around a million candidates per run and puts every hit through the same gauntlet: a honeypot filter, content-integrity checks, HTTPS with verified TLS, anonymity, country and provider lookups, and spam blocklists.

It learns with every run which sources are worth scraping, ranks proxies by how likely they are to still be up, and ships the result as a live list, a static website with a page per country, a rotating proxy server, an MCP server for AI agents and a Discord bot.

700+
sources
~1M
candidates per run
25 s
to check them
750+
tests on 3 OSes
maximilianfeix.github.io/proxy-scraper/
Proxy Scraper: free proxies that actually work – 700+ sources, five checks, a fresh list every hour.

Implementation

  1. 01

    An hourly pipeline on free CI

    A scheduled GitHub Actions workflow scrapes, checks and publishes the list, mirrors it to its own repository, keeps a daily snapshot as a release asset and a Parquet copy on Hugging Face.

  2. 02

    Checks that catch liars

    Each proxy fetches a known page; any byte that differs means injected content. TLS is verified end to end, and honeypots are detected by the way they answer requests no real proxy would accept.

  3. 03

    A server, not just a list

    A rotating SOCKS5/HTTP proxy server with sticky sessions and Prometheus metrics, plus an MCP server so AI agents can ask for a working proxy by country.

Technical challenges

  • A million sockets in 25 seconds

    The checker is fully asynchronous with separate connect and read timeouts, bounded concurrency and early exits, so the run fits comfortably into a CI job.

  • Sources that rot

    Source statistics are cached between runs; sources that stop delivering fall down the ranking instead of being removed by hand.

Architecture

Collect → deduplicate → check in stages, cheapest first → rank → publish. Every stage is a plain function over an async stream, which keeps each one testable against fake proxies on localhost.

textproxy-scraper
sources (700+) ─▶ scrape ─▶ dedupe ─▶ connect ─▶ integrity ─▶ TLS ─▶ anonymity ─▶ geo/ASN
                                                                              │
        website · JSON/CSV · rotating server · MCP · Discord ◀── rank ◀────────┘