Free proxies that actually work.

- Year
- 2026
- Role
- Author & maintainer
- Stack
- Python, asyncio, GitHub Actions, Prometheus
Overview
Scrapes HTTP, SOCKS4 and SOCKS5 proxies from 700+ sources and keeps only the ones that pass real checks – re-verified every hour by GitHub Actions.
Free proxy lists are mostly noise: dead hosts, honeypots and proxies that quietly inject scripts into the pages they serve – about one in five does. Proxy Scraper collects around a million candidates per run and puts every hit through the same gauntlet: a honeypot filter, content-integrity checks, HTTPS with verified TLS, anonymity, country and provider lookups, and spam blocklists.
It learns with every run which sources are worth scraping, ranks proxies by how likely they are to still be up, and ships the result as a live list, a static website with a page per country, a rotating proxy server, an MCP server for AI agents and a Discord bot.
- 700+
- sources
- ~1M
- candidates per run
- 25 s
- to check them
- 750+
- tests on 3 OSes

Implementation
01
An hourly pipeline on free CI
A scheduled GitHub Actions workflow scrapes, checks and publishes the list, mirrors it to its own repository, keeps a daily snapshot as a release asset and a Parquet copy on Hugging Face.
02
Checks that catch liars
Each proxy fetches a known page; any byte that differs means injected content. TLS is verified end to end, and honeypots are detected by the way they answer requests no real proxy would accept.
03
A server, not just a list
A rotating SOCKS5/HTTP proxy server with sticky sessions and Prometheus metrics, plus an MCP server so AI agents can ask for a working proxy by country.
Technical challenges
A million sockets in 25 seconds
The checker is fully asynchronous with separate connect and read timeouts, bounded concurrency and early exits, so the run fits comfortably into a CI job.
Sources that rot
Source statistics are cached between runs; sources that stop delivering fall down the ranking instead of being removed by hand.
Architecture
Collect → deduplicate → check in stages, cheapest first → rank → publish. Every stage is a plain function over an async stream, which keeps each one testable against fake proxies on localhost.
sources (700+) ─▶ scrape ─▶ dedupe ─▶ connect ─▶ integrity ─▶ TLS ─▶ anonymity ─▶ geo/ASN
│
website · JSON/CSV · rotating server · MCP · Discord ◀── rank ◀────────┘