Every website bot sooner or later runs into one of three problems:
- a captcha
- anti-bot protection
- heavy JS that has to be rendered to submit a form
A spider that walks a site with raw http requests is helpless here. There’s only one solution: launch a browser and let it handle the hard part. But the browser has to be hidden too. Under the hood it sends the same http requests as a Python spider, it just does so from the context of running JS. And that context is full of markers that tell the site someone is controlling the browser.
For this task I wrote a microservice, Robotex. Below I’ll explain how it works and why it’s needed when Scrapling and FlareSolverr already exist.
What was used before and now#
I used to go with Selenium. It runs Firefox through geckodriver over the Marionette protocol and gives itself away quite noticeably: navigator.webdriver, driver traces in the page environment and so on. Almost all of that can be hidden with an add-on that adjusts the page context before the site loads. That’s what I did, but it’s a constant race against every new check.
Now there are tools that solve this at the level of the browser itself:
- Camoufox - a Firefox build where fingerprints (screen, fonts, WebGL, core count, etc.) are spoofed in the engine code rather than through JS.
- Patchright - a patched Playwright that removes the CDP leaks anti-bots use to detect Chromium automation.
Scrapling works on top of them. It has a ready-made mode for getting through anti-bot pages, and I borrowed its algorithm. It turned out that passing Cloudflare isn’t about solving a puzzle. If the browser environment looks human and the IP is clean, it’s enough to wait for the widget and click the checkbox with a random offset and delay. The whole difficulty is correctly detecting that you’re facing a check, and making sure the browser no longer looks suspicious by the time it clicks.
For small volumes without accounts, Scrapling is enough. Problems start when accounts come in. Then the IP, User-Agent, cookies and browser fingerprint have to be tied into one profile and live together. Change the proxy but keep the old cookies, and the account goes into verification or gets banned. The protection starts nagging you with captchas, and accounts die.
What Robotex is#
Robotex takes over all the browser work: it gets past the anti-bot, runs a scenario on the page and hands the bot cookies it can use to keep browsing the site with plain requests.
It differs from FlareSolverr and Byparr in that those can only open a page and get cf_clearance. Robotex runs a scenario: fill a form, click, solve a captcha, wait for an element. And it differs from Scrapling in that it’s not a library inside your process but a separate service. The bot can be written in anything, all it needs is an http client.
Under the hood are FastAPI and a real Chrome driven by Patchright. I ported the browser wrapper from Scrapling. I started with Camoufox but eventually settled on Chrome.
How it works#
The bot sends a scenario:
POST /v1/solve
Content-Type: application/json
X-API-Key: <key>
{
"url": "https://example.com/login",
"actions": [
{"type": "fill", "selector": "#email", "value": "user@example.com"},
{"type": "fill", "selector": "#password", "value": "..."},
{"type": "mouse_move", "selector": "#submit"},
{"type": "click", "selector": "#submit"},
{"type": "wait_for", "selector": ".dashboard", "timeout": 30}
],
"proxy": "socks5://user:pass@host:port",
"html": true,
"timeout": 120
}The scenario fills in the login form and waits for the dashboard to load. Besides these actions there are submit (a click that waits for navigation to a new page), delay, inner_html (grab a piece of the page) and captcha, more on that below.
Since the request has no session_id cookie, the service creates a new browser profile and returns it in the Set-Cookie header. The profile is stored on disk: cookies, local storage and fingerprint. The browser locale and timezone are matched to the proxy’s geo, so the IP and the environment don’t contradict each other. When the bot needs a browser again, it sends this cookie, and Robotex continues in the same profile. To the site it’s still the same user.
The response:
{
"status": "ok",
"user_agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) ...",
"cookies": [
{"name": "sessionid", "value": "...", "domain": ".example.com", "path": "/"}
],
"actions": [
{"type": "fill", "ok": true},
{"type": "fill", "ok": true},
{"type": "mouse_move", "ok": true},
{"type": "click", "ok": true},
{"type": "wait_for", "ok": true}
],
"html": "<head>...</head><body>...</body>"
}From there the bot browses the site on its own, with these cookies, the same User-Agent and through the same proxy. If an action fails, status will be err, and actions shows at which step and why.
Captcha#
If a captcha is expected on the form, add a step for it and a recognition service key (2captcha is supported right now):
{
"url": "https://example.com/login",
"actions": [
{"type": "fill", "selector": "#email", "value": "user@example.com"},
{"type": "fill", "selector": "#password", "value": "..."},
{"type": "captcha", "selector": "#captcha", "value": "hcaptcha"},
{"type": "click", "selector": "#submit"},
{"type": "wait_for", "selector": ".dashboard", "timeout": 30}
],
"captcha": "<2captcha key>",
"proxy": "socks5://user:pass@host:port"
}The captcha will be solved before the submit click.
Anti-bot page#
An anti-bot page can appear at any moment, or not at all. So it’s not a scenario step but a setting for the whole request:
{
"url": "https://example.com/login",
"actions": [ ... ],
"captcha_type": "cloudflare",
"captcha_detect_locator": "#anti-bot-page-only-id",
"captcha_success_locator": "#site-only-id",
"captcha": "<2captcha key>",
"proxy": "socks5://user:pass@host:port"
}captcha_detect_locator is a selector that exists only on the check page, captcha_success_locator exists only on the site itself. Right after the page loads and before running the scenario, Robotex checks whether there’s a barrier in front of it. If there is, it gets through on its own: waits, clicks the checkbox. If the check escalated to a full image captcha, it sends it to the recognition service. If the check pops up in the middle of a scenario, there’s a separate pass_challenge action for that.
For heavy sites there are a couple more useful options: disable_resources turns off loading of images and fonts, blocked_domains cuts requests to unneeded domains like analytics, and cookies lets you pass in ready-made cookies.
Limitations#
There’s no silver bullet here:
- This is an early MVP. API key checks, a task queue and billing are still planned
- a browser eats hundreds of megabytes of memory, so one instance currently handles 10 concurrent sessions
- proxy quality matters a lot, no fingerprint will help on a dirty IP
- Turnstile doesn’t pass on every site, I’m still working on stability. Anti-bots get updated, and getting through a specific site has to be tested and maintained
On the upside, the bot stays lightweight: it needs a browser only where there’s no way around it, and the rest of the time it works with fast http requests.
Code#
Sources: git.sokolab.xyz/s0k0l/robotex. The docs folder also has the spec for the anti-bot bypass module, if you’re curious how it works inside.
Use it only on sites where you have the right to do so.
