[{"content":"","date":"11 October 2026","externalUrl":null,"permalink":"/en/","section":"Alexander Sokolov","summary":"","title":"Alexander Sokolov","type":"page"},{"content":"","date":"11 October 2026","externalUrl":null,"permalink":"/en/tags/anti-bot/","section":"Tags","summary":"","title":"Anti-Bot","type":"tags"},{"content":"","date":"11 October 2026","externalUrl":null,"permalink":"/en/tags/api/","section":"Tags","summary":"","title":"Api","type":"tags"},{"content":"","date":"11 October 2026","externalUrl":null,"permalink":"/en/blog/","section":"Blog","summary":"","title":"Blog","type":"blog"},{"content":"","date":"11 October 2026","externalUrl":null,"permalink":"/en/tags/parsing/","section":"Tags","summary":"","title":"Parsing","type":"tags"},{"content":"I recently wrote about how Robotex works. It\u0026rsquo;s my microservice that takes over all the browser work: it gets past anti-bot protection, runs a scenario on the page and hands your bot the cookies. Now it has a home of its own: robotex.sokolab.xyz.\nWhy a separate page # A blog post is good for explaining how something works. But a project needs a place where you can quickly see what it is and why, get access and ask a question. So each of my projects will get its own subdomain with a landing page, and the blog will cover what\u0026rsquo;s new.\nFree # Robotex will be a free service. You only pay for your own proxies and, if needed, captcha recognition in 2captcha.\nIt\u0026rsquo;s an early MVP for now, so I grant access manually. Register in my Redmine, I\u0026rsquo;ll approve the account and send you the API address.\nWhere to discuss # The project has its own section in Redmine:\na forum for questions, ideas and scenario discussions issues for bugs and feature requests news and plans If it\u0026rsquo;s more convenient, message me on Telegram.\nThe source code is still open: git.sokolab.xyz/s0k0l/robotex.\n","date":"11 October 2026","externalUrl":null,"permalink":"/en/blog/robotex-free-saas/","section":"Blog","summary":"Robotex has its own page now: robotex.sokolab.xyz. It’s free to use, access on request","title":"Robotex is now a free service","type":"blog"},{"content":"","date":"11 October 2026","externalUrl":null,"permalink":"/en/tags/saas/","section":"Tags","summary":"","title":"Saas","type":"tags"},{"content":"","date":"11 October 2026","externalUrl":null,"permalink":"/en/tags/","section":"Tags","summary":"","title":"Tags","type":"tags"},{"content":"While going through old accounts, I came across my trail in Social Mapper. It\u0026rsquo;s an OSINT tool by Jacob Wilkin (Greenwolf): it searches for people across social networks and uses facial recognition to stitch one person\u0026rsquo;s profiles from different sites into a single report. Pentesters and red teams use it to build a list of a company\u0026rsquo;s employees and their pages. The repository has over 4000 stars and 800 forks.\nIn September 2020 I sent it 20 commits in two pull requests, and I\u0026rsquo;m still listed as the third contributor by commit count, after the author\u0026rsquo;s own two accounts. Here\u0026rsquo;s what happened.\nDouban # It all started with issue #166. The Chinese social network Douban stopped working: the module failed with HTTP 418, and it was unclear whether it had logged in or not. The issue had been open since April. The author replied that the site seemed to have changed, but he couldn\u0026rsquo;t check because he couldn\u0026rsquo;t find where to sign up.\nSocial Mapper works through Selenium and Firefox, so this was familiar territory for me. Douban had changed its page title and login form, and the module was looking for the old ones. In PR #193 I:\nfixed the login for the new form and the title check added a clear message when the login page doesn\u0026rsquo;t load as expected, instead of failing silently moved the Firefox and geckodriver paths into the FIREFOX_BINARY and GECKODRIVER environment variables, so they don\u0026rsquo;t have to be on PATH fixed several Python 3 bytes vs str errors in file reading and writing left over from the Python 2 days The funny part is that I didn\u0026rsquo;t have a Douban account either, so I couldn\u0026rsquo;t fully test the login. I asked the issue author to clone my fork and switch to the branch with the fix.\nWhile the PR was waiting for review, the author pushed a batch of his own fixes and missed mine. He apologized and asked me to resubmit. I resolved the conflicts, pinged him a day later, and the PR was merged on September 9.\nTabs and spaces # While digging through the code, I noticed that social_mapper.py used 4-space indentation while the other modules used tabs. In Python that\u0026rsquo;s not cosmetic: mixed indentation breaks code or, worse, makes it misleading. I opened issue #195 and right away sent PR #196:\nthe whole project moved to 4 spaces, as PEP 8 recommends string literals used as comments were replaced with real comments the code was formatted to PEP 8 and unused imports were removed added .gitignore and .editorconfig In total: 11 files, +1657 and −1430 lines. The author merged it the same day, said I seemed to be on a coding spree, and asked whether I had any ideas on getting around Facebook\u0026rsquo;s new rate limit. He also added me to the thanks list in the README, where I still am.\nWhat I took away from it # Someone else\u0026rsquo;s open source project turned out to be a great place to practice reading unfamiliar code and working with a maintainer. A small targeted fix that solves someone\u0026rsquo;s problem gets accepted gladly. After that you can propose something bigger, like refactoring the whole project.\nAnd yes, my Development Standards now say tabs instead of spaces. Six years later I changed my mind, but the main rule stayed the same: one indentation style per project, and a tool enforces it, not a person.\nThe project is no longer actively maintained, but the author still accepts pull requests. If you use it, do so only within legal penetration tests and with the client\u0026rsquo;s consent.\n","date":"10 October 2026","externalUrl":null,"permalink":"/en/blog/social-mapper-contribution/","section":"Blog","summary":"Two pull requests to a 4000-star OSINT tool: I fixed Douban and moved the whole project to spaces","title":"How I contributed to Social Mapper","type":"blog"},{"content":"","date":"10 October 2026","externalUrl":null,"permalink":"/en/tags/open-source/","section":"Tags","summary":"","title":"Open-Source","type":"tags"},{"content":"","date":"10 October 2026","externalUrl":null,"permalink":"/en/tags/osint/","section":"Tags","summary":"","title":"Osint","type":"tags"},{"content":"","date":"10 October 2026","externalUrl":null,"permalink":"/en/tags/python/","section":"Tags","summary":"","title":"Python","type":"tags"},{"content":"","date":"10 October 2026","externalUrl":null,"permalink":"/en/tags/selenium/","section":"Tags","summary":"","title":"Selenium","type":"tags"},{"content":"Every website bot sooner or later runs into one of three problems:\na captcha anti-bot protection heavy JS that has to be rendered to submit a form A spider that walks a site with raw http requests is helpless here. There\u0026rsquo;s only one solution: launch a browser and let it handle the hard part. But the browser has to be hidden too. Under the hood it sends the same http requests as a Python spider, it just does so from the context of running JS. And that context is full of markers that tell the site someone is controlling the browser.\nFor this task I wrote a microservice, Robotex. Below I\u0026rsquo;ll explain how it works and why it\u0026rsquo;s needed when Scrapling and FlareSolverr already exist.\nWhat was used before and now # I used to go with Selenium. It runs Firefox through geckodriver over the Marionette protocol and gives itself away quite noticeably: navigator.webdriver, driver traces in the page environment and so on. Almost all of that can be hidden with an add-on that adjusts the page context before the site loads. That\u0026rsquo;s what I did, but it\u0026rsquo;s a constant race against every new check.\nNow there are tools that solve this at the level of the browser itself:\nCamoufox - a Firefox build where fingerprints (screen, fonts, WebGL, core count, etc.) are spoofed in the engine code rather than through JS. Patchright - a patched Playwright that removes the CDP leaks anti-bots use to detect Chromium automation. Scrapling works on top of them. It has a ready-made mode for getting through anti-bot pages, and I borrowed its algorithm. It turned out that passing Cloudflare isn\u0026rsquo;t about solving a puzzle. If the browser environment looks human and the IP is clean, it\u0026rsquo;s enough to wait for the widget and click the checkbox with a random offset and delay. The whole difficulty is correctly detecting that you\u0026rsquo;re facing a check, and making sure the browser no longer looks suspicious by the time it clicks.\nFor small volumes without accounts, Scrapling is enough. Problems start when accounts come in. Then the IP, User-Agent, cookies and browser fingerprint have to be tied into one profile and live together. Change the proxy but keep the old cookies, and the account goes into verification or gets banned. The protection starts nagging you with captchas, and accounts die.\nWhat Robotex is # Robotex takes over all the browser work: it gets past the anti-bot, runs a scenario on the page and hands the bot cookies it can use to keep browsing the site with plain requests.\nIt differs from FlareSolverr and Byparr in that those can only open a page and get cf_clearance. Robotex runs a scenario: fill a form, click, solve a captcha, wait for an element. And it differs from Scrapling in that it\u0026rsquo;s not a library inside your process but a separate service. The bot can be written in anything, all it needs is an http client.\nUnder the hood are FastAPI and a real Chrome driven by Patchright. I ported the browser wrapper from Scrapling. I started with Camoufox but eventually settled on Chrome.\nHow it works # The bot sends a scenario:\nPOST /v1/solve Content-Type: application/json X-API-Key: \u0026lt;key\u0026gt; { \u0026#34;url\u0026#34;: \u0026#34;https://example.com/login\u0026#34;, \u0026#34;actions\u0026#34;: [ {\u0026#34;type\u0026#34;: \u0026#34;fill\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#email\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;user@example.com\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;fill\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#password\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;...\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;mouse_move\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#submit\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;click\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#submit\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;wait_for\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;.dashboard\u0026#34;, \u0026#34;timeout\u0026#34;: 30} ], \u0026#34;proxy\u0026#34;: \u0026#34;socks5://user:pass@host:port\u0026#34;, \u0026#34;html\u0026#34;: true, \u0026#34;timeout\u0026#34;: 120 } The scenario fills in the login form and waits for the dashboard to load. Besides these actions there are submit (a click that waits for navigation to a new page), delay, inner_html (grab a piece of the page) and captcha, more on that below.\nSince the request has no session_id cookie, the service creates a new browser profile and returns it in the Set-Cookie header. The profile is stored on disk: cookies, local storage and fingerprint. The browser locale and timezone are matched to the proxy\u0026rsquo;s geo, so the IP and the environment don\u0026rsquo;t contradict each other. When the bot needs a browser again, it sends this cookie, and Robotex continues in the same profile. To the site it\u0026rsquo;s still the same user.\nThe response:\n{ \u0026#34;status\u0026#34;: \u0026#34;ok\u0026#34;, \u0026#34;user_agent\u0026#34;: \u0026#34;Mozilla/5.0 (Windows NT 10.0; Win64; x64) ...\u0026#34;, \u0026#34;cookies\u0026#34;: [ {\u0026#34;name\u0026#34;: \u0026#34;sessionid\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;...\u0026#34;, \u0026#34;domain\u0026#34;: \u0026#34;.example.com\u0026#34;, \u0026#34;path\u0026#34;: \u0026#34;/\u0026#34;} ], \u0026#34;actions\u0026#34;: [ {\u0026#34;type\u0026#34;: \u0026#34;fill\u0026#34;, \u0026#34;ok\u0026#34;: true}, {\u0026#34;type\u0026#34;: \u0026#34;fill\u0026#34;, \u0026#34;ok\u0026#34;: true}, {\u0026#34;type\u0026#34;: \u0026#34;mouse_move\u0026#34;, \u0026#34;ok\u0026#34;: true}, {\u0026#34;type\u0026#34;: \u0026#34;click\u0026#34;, \u0026#34;ok\u0026#34;: true}, {\u0026#34;type\u0026#34;: \u0026#34;wait_for\u0026#34;, \u0026#34;ok\u0026#34;: true} ], \u0026#34;html\u0026#34;: \u0026#34;\u0026lt;head\u0026gt;...\u0026lt;/head\u0026gt;\u0026lt;body\u0026gt;...\u0026lt;/body\u0026gt;\u0026#34; } From there the bot browses the site on its own, with these cookies, the same User-Agent and through the same proxy. If an action fails, status will be err, and actions shows at which step and why.\nCaptcha # If a captcha is expected on the form, add a step for it and a recognition service key (2captcha is supported right now):\n{ \u0026#34;url\u0026#34;: \u0026#34;https://example.com/login\u0026#34;, \u0026#34;actions\u0026#34;: [ {\u0026#34;type\u0026#34;: \u0026#34;fill\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#email\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;user@example.com\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;fill\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#password\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;...\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;captcha\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#captcha\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;hcaptcha\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;click\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;#submit\u0026#34;}, {\u0026#34;type\u0026#34;: \u0026#34;wait_for\u0026#34;, \u0026#34;selector\u0026#34;: \u0026#34;.dashboard\u0026#34;, \u0026#34;timeout\u0026#34;: 30} ], \u0026#34;captcha\u0026#34;: \u0026#34;\u0026lt;2captcha key\u0026gt;\u0026#34;, \u0026#34;proxy\u0026#34;: \u0026#34;socks5://user:pass@host:port\u0026#34; } The captcha will be solved before the submit click.\nAnti-bot page # An anti-bot page can appear at any moment, or not at all. So it\u0026rsquo;s not a scenario step but a setting for the whole request:\n{ \u0026#34;url\u0026#34;: \u0026#34;https://example.com/login\u0026#34;, \u0026#34;actions\u0026#34;: [ ... ], \u0026#34;captcha_type\u0026#34;: \u0026#34;cloudflare\u0026#34;, \u0026#34;captcha_detect_locator\u0026#34;: \u0026#34;#anti-bot-page-only-id\u0026#34;, \u0026#34;captcha_success_locator\u0026#34;: \u0026#34;#site-only-id\u0026#34;, \u0026#34;captcha\u0026#34;: \u0026#34;\u0026lt;2captcha key\u0026gt;\u0026#34;, \u0026#34;proxy\u0026#34;: \u0026#34;socks5://user:pass@host:port\u0026#34; } captcha_detect_locator is a selector that exists only on the check page, captcha_success_locator exists only on the site itself. Right after the page loads and before running the scenario, Robotex checks whether there\u0026rsquo;s a barrier in front of it. If there is, it gets through on its own: waits, clicks the checkbox. If the check escalated to a full image captcha, it sends it to the recognition service. If the check pops up in the middle of a scenario, there\u0026rsquo;s a separate pass_challenge action for that.\nFor heavy sites there are a couple more useful options: disable_resources turns off loading of images and fonts, blocked_domains cuts requests to unneeded domains like analytics, and cookies lets you pass in ready-made cookies.\nLimitations # There\u0026rsquo;s no silver bullet here:\nThis is an early MVP. API key checks, a task queue and billing are still planned a browser eats hundreds of megabytes of memory, so one instance currently handles 10 concurrent sessions proxy quality matters a lot, no fingerprint will help on a dirty IP Turnstile doesn\u0026rsquo;t pass on every site, I\u0026rsquo;m still working on stability. Anti-bots get updated, and getting through a specific site has to be tested and maintained On the upside, the bot stays lightweight: it needs a browser only where there\u0026rsquo;s no way around it, and the rest of the time it works with fast http requests.\nCode # Sources: git.sokolab.xyz/s0k0l/robotex. The docs folder also has the spec for the anti-bot bypass module, if you\u0026rsquo;re curious how it works inside.\nUse it only on sites where you have the right to do so.\n","date":"9 October 2026","externalUrl":null,"permalink":"/en/blog/robotex-browser-automation-api/","section":"Blog","summary":"How a bot can scrape a site behind anti-bot protection or with heavy JS","title":"Robotex, a Browser Automation API","type":"blog"},{"content":"Usually, after you buy a server, you get an email with an IP address and the root password. That\u0026rsquo;s enough to connect to the server over ssh and start setting it up.\nJust 3 steps:\nSystem update ssh setup nftables setup Step zero is preparing ssh keys and a config for the server. To generate keys, run:\ns0k0l:~$ ssh-keygen Generating public/private ed25519 key pair. Enter file in which to save the key (/~/.ssh/id_ed25519): It asks for a path for the key. You can just type the key name you want, and the pair will be created in the directory you ran the command from. The first time you can skip this and press Enter.\nThe second question is about a passphrase for the key. For ssh keys to production infrastructure it\u0026rsquo;s better to set one. If you don\u0026rsquo;t want to, just hit Enter.\nOnce the hosting provider brings the server up, copy the key over with\ns0k0l:~$ ssh-copy-id root@\u0026lt;ip\u0026gt; /usr/bin/ssh-copy-id: INFO: attempting to log in with the new key(s), to filter out any that are already installed /usr/bin/ssh-copy-id: INFO: 3 key(s) remain to be installed -- if you are prompted now it is to install the new keys root@\u0026lt;ip\u0026gt;\u0026#39;s password: Enter the password and you should see a success message suggesting you connect to the server, which is exactly what we\u0026rsquo;ll do. But first add a new entry to ~/.ssh/config:\nHost \u0026lt;name\u0026gt; HostName \u0026lt;ip\u0026gt; Port 22 IdentityFile ~/.ssh/id_ed25519 Now you can connect by name\nssh \u0026lt;name\u0026gt; On the server, the first thing to do is run an update:\napt update \u0026amp;\u0026amp; apt dist-upgrade -y \u0026amp;\u0026amp; reboot This updates the whole system and reboots it. Wait about 5 minutes and try to connect again.\nNow we need to configure ssh so that nobody else can connect to the server on the default port or log in with a password. This is the most important part.\nOpen /etc/ssh/sshd_config for editing and find the parameters for the listen address and port.\nPort 5872 AddressFamily inet ListenAddress \u0026lt;ip\u0026gt; #ListenAddress :: Set them like this to strictly define how the server can be reached. We have an ipv4 address, so we can put it in the config explicitly. And the port should be changed to a random one in the range from 1024 to 65535. To check which ports are already taken, run:\nss -tlnup If the port you want isn\u0026rsquo;t in the output, it\u0026rsquo;s free.\nNext, find the following in the ssh config:\nPermitRootLogin prohibit-password PubkeyAuthentication yes PasswordAuthentication no These settings disable ssh password authentication for all users, including root (PermitRootLogin).\nAfter saving the config, restart the ssh service. Stay in your session until you\u0026rsquo;ve connected in parallel.\nsystemctl restart ssh Open another terminal, fix the port in ~/.ssh/config and connect to the server. If the connection works, you can close the first terminal. SSH is done.\nThe last step is setting up nftables. It\u0026rsquo;s the standard firewall in Debian, iptables is no longer needed.\nThe rules live in a single file, /etc/nftables.conf. The principle of a strict config is simple: all incoming traffic is denied, we allow only what is actually needed. Outgoing traffic stays open, otherwise updates and DNS will break.\nOpen /etc/nftables.conf and replace its contents entirely:\n#!/usr/sbin/nft -f flush ruleset table inet filter { chain input { type filter hook input priority 0; policy drop; # let established connections through, drop garbage ct state established,related accept ct state invalid drop # loopback interface iif lo accept # ping, but no more than 5 per second ip protocol icmp icmp type echo-request limit rate 5/second accept ip protocol icmp accept meta l4proto ipv6-icmp accept # ssh on our port, no more than 10 new connections per minute from one ip tcp dport 5872 ct state new meter ssh_limit { ip saddr limit rate 10/minute } accept # uncomment if the server will host a website # tcp dport { 80, 443 } accept } chain forward { type filter hook forward priority 0; policy drop; } chain output { type filter hook output priority 0; policy accept; } } Change port 5872 to the one you set in sshd_config. This is the most important line in the config, a mistake here will lock you out of the server.\nFirst, check the config for errors without applying anything:\nnft -c -f /etc/nftables.conf If the command printed nothing, the syntax is fine. Now a safety net in case we do lock ourselves out. Start a timer that will flush all rules in 2 minutes:\n(sleep 120 \u0026amp;\u0026amp; nft flush ruleset) \u0026amp; And apply the config:\nnft -f /etc/nftables.conf Then, as with ssh, open another terminal and connect. If you got in, everything is fine, cancel the timer:\nkill %1 If you didn\u0026rsquo;t get in, just wait 2 minutes, the rules will be flushed automatically and you can calmly look for the mistake.\nTo see what is actually applied right now, run:\nnft list ruleset All that\u0026rsquo;s left is to enable loading the rules at boot, otherwise after a reboot the server will be left without a firewall:\nsystemctl enable --now nftables If the server will run Docker, keep in mind that it writes its own rules and can open container ports bypassing your config. That\u0026rsquo;s a topic for a separate article.\nThat\u0026rsquo;s it. The server is updated, ssh only lets you in with a key and on a non-standard port, and the firewall blocks everything we haven\u0026rsquo;t explicitly allowed. From here you can start installing your services.\nDisclaimer: the timer trick is needed because a firewall can drop even an already established connection. The example config has the line ct state established,related accept, which means \u0026ldquo;allow connections that are already established\u0026rdquo;. So in theory there shouldn\u0026rsquo;t be a disconnect.\n","date":"8 October 2026","externalUrl":null,"permalink":"/en/blog/zero-setup-debian-server/","section":"Blog","summary":"What to do with a Debian server after you buy it. An article for Linux users.","title":"Initial Debian server setup","type":"blog"},{"content":"","date":"8 October 2026","externalUrl":null,"permalink":"/en/tags/linux/","section":"Tags","summary":"","title":"Linux","type":"tags"},{"content":"","date":"8 October 2026","externalUrl":null,"permalink":"/en/tags/server/","section":"Tags","summary":"","title":"Server","type":"tags"},{"content":" Intro # The first thing you need is a server. Just google \u0026ldquo;Debian VPS\u0026rdquo; and pick from what comes up.\nImportant: if you\u0026rsquo;re in Russia, it\u0026rsquo;s better to choose a hosting provider and a server in your own country. Then search Yandex for \u0026ldquo;аренда виртуального сервера Debian\u0026rdquo;.\nWhile the server is being set up, you have about an hour to buy a domain name and register a Cloudflare account. The first one is obvious, and CF is needed to hide your server\u0026rsquo;s IP address from site visitors. This definitely improves server security, but it has to be done before the domain is first pointed at the server. So after buying the domain, go to the control panel (where you bought it) and change the NS servers to the ones Cloudflare gives you after you add the project in your account. After that it takes some time (up to 3 days) for the domain to be re-parked, since the whole internet has to learn about it. But after that no visitor will be able to find out your site\u0026rsquo;s IP, unless you leak it, of course. We\u0026rsquo;ll talk about that some other time.\nWhat you have now:\nthe IP and root password a domain name parked at CF records set up in CF: yourdomain.com - IP - the root record, the domain name opens the blog home page git.yourdomain.com - IP - a subdomain for Gitea Let\u0026rsquo;s go # Initial Debian server setup -\u0026gt;\nInstall Docker with\ncurl -fsSL https://get.docker.com -o install-docker.sh sh install-docker.sh Log out of the server, because now we need to prepare everything locally.\nOn your computer, install Hugo from the GitHub repository, because the Debian package repository has an old version.\nhttps://github.com/gohugoio/hugo/releases\nInstall Git on your computer to work with the repository.\nOn the server we\u0026rsquo;ll run 2 applications in Docker:\ngit - Gitea, a git service, something like GitHub hugo - the site builder and Caddy. Caddy serves the finished pages and is also the only entry point from outside, Gitea traffic goes through it too Hugo builds blog pages from Markdown files, converting everything to HTML. The workflow looks like this:\ncreate an md file in the repository fill it in, commit and push the builder in Docker checks the repository once a minute and sees a new commit it builds the site and swaps the old version for the new one Okay, now pick a place on your computer and create an apps folder, where we\u0026rsquo;ll keep the configuration of the Docker applications on the server. The final structure will look like this:\napps/ ├── git/ │ └── docker-compose.yaml └── hugo/ ├── docker-compose.yaml ├── caddy.conf ├── dot_env ├── .gitignore └── data/ └── builder/ └── build.sh There are only configs here. All data (repositories, certificates, the built site) lives in Docker named volumes, not in folders next to the configs. This way apps can safely be kept in git and copied to the server without fear of overwriting data. I\u0026rsquo;ll cover backups below.\nBoth applications talk over a shared web network. Only Caddy exposes ports to the outside.\ngit # apps/git/docker-compose.yaml:\n# Git application for docker-server services: gitea: image: gitea/gitea:28.0-rootless restart: unless-stopped environment: GITEA__server__DOMAIN: git.yourdomain.com GITEA__server__ROOT_URL: https://git.yourdomain.com/ GITEA__server__DISABLE_SSH: \u0026#34;true\u0026#34; GITEA__database__DB_TYPE: sqlite3 GITEA__service__DISABLE_REGISTRATION: \u0026#34;true\u0026#34; volumes: - data:/var/lib/gitea - config:/etc/gitea networks: [ web ] networks: web: external: true volumes: data: config: We use the rootless image, so inside the container Gitea doesn\u0026rsquo;t run as root. This is where named volumes really fit: when creating a volume, Docker sets its owner from the image, and you don\u0026rsquo;t have to chown anything by hand.\nWe don\u0026rsquo;t expose any ports, only Caddy reaches Gitea over the web network. SSH is disabled because Cloudflare doesn\u0026rsquo;t proxy it, and we don\u0026rsquo;t want to expose the server\u0026rsquo;s IP. We\u0026rsquo;ll push over https. Registration is closed so nobody but you can create accounts there.\nhugo # apps/hugo/docker-compose.yaml:\n# Landing page and blog site on Hugo services: builder: image: ghcr.io/gohugoio/hugo:v0.167.0 restart: unless-stopped user: root entrypoint: [ \u0026#34;sh\u0026#34;, \u0026#34;/build.sh\u0026#34; ] environment: REPO_URL: ${REPO_URL} volumes: - ./data/builder/build.sh:/build.sh:ro - src:/src - public:/public networks: [ web ] # to reach git-gitea-1 web: image: caddy:2.10.2-alpine restart: unless-stopped ports: - \u0026#34;80:80\u0026#34; - \u0026#34;443:443\u0026#34; volumes: - ./caddy.conf:/etc/caddy/Caddyfile:ro - public:/srv:ro - caddy_data:/data - caddy_config:/config networks: [ web ] # to proxy to git-gitea-1 networks: web: external: true volumes: src: public: caddy_data: caddy_config: You don\u0026rsquo;t need to build your own image. builder is the official Hugo image, it already includes git. Use the same version as on your computer, you can check it with hugo version. user: root is needed because the image runs as a regular user by default, while Docker creates volumes as root.\nweb is Caddy. It faces the outside on 80 and 443, serves the site from the public volume shared with builder, and proxies the git subdomain to Gitea. It keeps its certificates in caddy_data.\napps/hugo/caddy.conf:\nyourdomain.com { tls internal encode gzip root * /srv/live file_server handle_errors { rewrite * /404.html file_server } } www.yourdomain.com { tls internal redir https://yourdomain.com{uri} permanent } git.yourdomain.com { tls internal reverse_proxy git-gitea-1:3000 } The site is served from /srv/live, why exactly from there will become clear from the build script. handle_errors makes non-existent pages show the theme\u0026rsquo;s nice 404 instead of an empty response. www simply redirects to the main domain so the site has a single address.\nDocker Compose builds the name git-gitea-1 itself from the folder name and the service name: folder git, service gitea, first instance. That\u0026rsquo;s why it matters that the folder is named exactly like that.\ntls internal means Caddy issues a certificate for itself. The browser never sees it, because visitors connect to Cloudflare, and Cloudflare connects to us. So in the Cloudflare panel, under SSL/TLS, set the mode to Full. Not Flexible, otherwise traffic from Cloudflare to the server goes unencrypted, and not Full (strict), which won\u0026rsquo;t accept such a certificate.\napps/hugo/dot_env is a template for the variables:\nREPO_URL=http://git-gitea-1:3000/\u0026lt;name\u0026gt;/site.git On the server you copy it to .env and fill in your username and repository. The .env itself doesn\u0026rsquo;t go into git, that\u0026rsquo;s what the .gitignore next to it with a single .env line is for. The builder clones the repository straight from the Gitea container over the internal network, without going out to the internet. If the repository is private, add a token: http://\u0026lt;name\u0026gt;:\u0026lt;token\u0026gt;@git-gitea-1:3000/\u0026lt;name\u0026gt;/site.git. The token is issued in the Gitea user settings, under Applications. In that case hide the file from prying eyes with chmod 600 .env.\napps/hugo/data/builder/build.sh:\n#!/bin/sh # Once a minute: new commit -\u0026gt; build into /public/next -\u0026gt; swap /public/live. # Build error - the previous site stays. # The repository address is taken from REPO_URL on every start: change in .env + redeploy = new source. cd /src if [ -d .git ]; then git remote set-url origin \u0026#34;$REPO_URL\u0026#34; else git clone --recurse-submodules \u0026#34;$REPO_URL\u0026#34; . || exit 1 fi last=\u0026#39;\u0026#39; while true; do if git fetch -q origin HEAD \u0026amp;\u0026amp; git reset -q --hard FETCH_HEAD \\ \u0026amp;\u0026amp; git submodule sync -q --recursive \u0026amp;\u0026amp; git submodule update -q --init --recursive; then rev=$(git rev-parse HEAD) if [ \u0026#34;$rev\u0026#34; != \u0026#34;$last\u0026#34; ]; then rm -rf /public/next if hugo --minify -d /public/next; then rm -rf /public/old [ -d /public/live ] \u0026amp;\u0026amp; mv /public/live /public/old mv /public/next /public/live last=$rev echo \u0026#34;published $rev\u0026#34; else echo \u0026#34;build of $rev failed, site unchanged\u0026#34; fi fi fi sleep 60 done The main trick of the script is that the site is built into a separate next folder, and only if the build succeeds does it replace the working live one. If you push a broken commit, the site simply stays as it was, and the log shows what broke. The previous version is kept in old in case you need to roll back quickly.\ngit reset --hard is used here on purpose instead of git pull. If you force push or rewrite history, a regular pull will break, while reset just takes whatever is in the repository. And set-url at startup lets you change the source: edit REPO_URL in .env, restart the container, and the builder pulls from the new place.\nUploading to the server # The configs are ready, let\u0026rsquo;s send the folder to the server. \u0026lt;name\u0026gt; is the host name from ~/.ssh/config that we set up in the previous article:\ns0k0l:~$ scp -r apps \u0026lt;name\u0026gt;:/opt/ Before starting, we need to open the firewall for the site. In the previous article we closed all incoming traffic, and I warned there that Docker lives by its own rules. Let\u0026rsquo;s sort it out now.\nOpen /etc/nftables.conf on the server. In the input chain, uncomment the line for the site:\ntcp dport { 80, 443 } accept And change the forward chain to this:\nchain forward { type filter hook forward priority 0; policy drop; ct state established,related accept ct status dnat accept iifname \u0026#34;docker0\u0026#34; accept iifname \u0026#34;br-*\u0026#34; accept } } The thing is, traffic to containers goes not through input but through forward. With our strict policy drop, containers can neither accept connections nor reach the internet. These rules let through traffic to the ports Docker published (in our case only Caddy\u0026rsquo;s 80 and 443) and outgoing traffic from the containers themselves. Everything else is still closed.\nCheck and apply the same way as last time, with a timer just in case:\nnft -c -f /etc/nftables.conf (sleep 120 \u0026amp;\u0026amp; nft flush ruleset) \u0026amp; nft -f /etc/nftables.conf Check that ssh is alive, cancel the timer with kill %1 and restart Docker:\nsystemctl restart docker This is mandatory. Our config starts with flush ruleset, which also wipes the rules Docker created for itself. After a restart Docker creates them again. Remember this: every time you restart nftables, restart Docker too.\nLaunch # Create the shared network through which the containers will see each other:\ndocker network create web Bring up Gitea and, for now, only Caddy without the builder. The builder has nothing to clone until there\u0026rsquo;s a repository in Gitea:\ncd /opt/apps/git \u0026amp;\u0026amp; docker compose up -d cd /opt/apps/hugo \u0026amp;\u0026amp; cp dot_env .env \u0026amp;\u0026amp; docker compose up -d web Open https://git.yourdomain.com in the browser. Gitea will show the initial setup page, the database is already set in the config. At the bottom of the page, in the administrator account section, create your user. We closed registration, so this is the only way to get an account.\nCreate a site repository in Gitea. Now push the site there from your computer:\ns0k0l:~$ cd my-blog s0k0l:~$ git remote add origin https://git.yourdomain.com/\u0026lt;name\u0026gt;/site.git s0k0l:~$ git push -u origin main If the theme is included as a git submodule, the theme repository must also be reachable by the builder at the URL in .gitmodules. The easiest way is to make a mirror of the theme in your Gitea and put its address in .gitmodules.\nPut your real username into /opt/apps/hugo/.env and start the builder:\ncd /opt/apps/hugo \u0026amp;\u0026amp; docker compose up -d Check that it built the site:\ndocker logs -f hugo-builder-1 If published \u0026lt;commit hash\u0026gt; shows up in the log, open https://yourdomain.com, the blog is working.\nBackups # Since the data lives in named volumes, you can\u0026rsquo;t see it in apps. You can list them with docker volume ls. Two of them matter here: git_data and git_config, they hold all the repositories and Gitea settings. The site doesn\u0026rsquo;t need backing up, it\u0026rsquo;s built from the repository, and Caddy will issue certificates again.\nBack up Gitea:\ndocker compose -f /opt/apps/git/docker-compose.yaml stop docker run --rm -v git_data:/data -v git_config:/config -v /root:/backup alpine \\ tar czf /backup/gitea-$(date +%F).tgz /data /config docker compose -f /opt/apps/git/docker-compose.yaml start We stop Gitea during the backup so the SQLite database isn\u0026rsquo;t written halfway. Pull the finished archive to your machine with scp.\nHow to write now # The whole process looks like this:\ncreate a post with hugo new content blog/my-post/index.md write it and preview it locally with hugo server commit and push a minute later the post is on the site For convenience I added a Makefile to the repository so I don\u0026rsquo;t have to type the commands by hand:\nall: hugo new content $(path) publish: git add . git commit -m \u0026#39;chore: $(msg)\u0026#39; git push A new post is make path=blog/my-post/index.md, publishing is make publish msg=\u0026quot;new post\u0026quot;.\nThat\u0026rsquo;s it. A server, your own git, a blog and auto-deploy, all on one inexpensive VPS with no third-party services except Cloudflare.\n","date":"8 October 2026","externalUrl":null,"permalink":"/en/blog/how-to-make-blog/","section":"Blog","summary":"How to set up your own blog","title":"How a programmer can start their own blog","type":"blog"},{"content":"","date":"8 October 2026","externalUrl":null,"permalink":"/en/tags/hugo/","section":"Tags","summary":"","title":"Hugo","type":"tags"},{"content":" Who I am # My name is Alexander Sokolov, online I go by s0k0l. I\u0026rsquo;m a generalist IT specialist: user behavior emulation, data extraction and information security. I write scrapers and bots that get past anti-bot protection, build websites and network services, and secure infrastructure and development.\nI prefer to work from a spec and take a project from idea to MVP or release. If there is no spec, I\u0026rsquo;ll write it myself: do the research, pick the technologies, estimate complexity and timelines. I work by my own Software Development Standards and always start with requirements and architecture, not with code.\nMy result is a solved problem. Source code is the proof that the solution works. If we agreed on something, I will do everything to keep my promise.\nWhat I do # I create software and hardware products for people and businesses that help them live better and work more comfortably.\nSoftware: bots, services, automation # I\u0026rsquo;m drawn to tasks around the web and Linux:\nwebsite scraping and anti-bot bypass website bots and chatbots high-load systems built on queues and microservices websites, web applications and services accepting crypto payments directly, on-chain, with no intermediaries information security for servers and development My detailed track record is in the CV.\nHardware: devices, electronics, prototypes # I learned microcontroller programming because I dream of building robots. In the meantime I built a hardware password manager, ProtoKey. I plan to keep growing in this direction, I have plenty of ideas for useful devices.\nWhere to find me # Everything alive is here now:\nBlog - articles on development, servers and automation git.sokolab.xyz - my code and examples I used to be active on other platforms too. I barely visit them now, but I keep the links, they show where I started.\nStack Overflow (in Russian). For many years I answered questions there and asked colleagues myself. The site is almost empty now, AI answers these questions instead.\nMedium. My first blog. I wrote about development there, but someone else\u0026rsquo;s platform got old fast, so now I write here.\nGitHub. Code for the Medium articles and old projects. New code lives on my own git.\nMozilla Add-ons. When I was building data collection tools, I needed to adapt Firefox to my tasks. I wrote a couple of add-ons for that and published them in the public catalog. I made them for myself, but other people started using them, and they still leave comments.\n","date":"8 October 2026","externalUrl":null,"permalink":"/en/about/","section":"Alexander Sokolov","summary":"","title":"About me","type":"page"},{"content":" Download PDF PDF на русском Alexander Sokolov # Senior Python Developer (Backend, Data Collection)\n33 years old. Full time, remote. Not open to relocation or business travel.\nEmail: xalex.sokolov@yahoo.com Telegram: @xs0k0lx (preferred) About me # Python developer with over 8 years of commercial experience. Main specialization: large scale data collection and anti-bot bypass. I also have backend experience with enterprise systems on Django.\nI work with TDD. I can take a task end to end: design the architecture, implement it, deploy it and support it. I try to solve not only my own ticket but also the team\u0026rsquo;s problems when I see they can be removed, like I did with environment deployment at Tsifra.\nHow I can help:\ndata collection systems that handle large volumes and adapt easily to changes on the sources bypassing Cloudflare, captchas and browser fingerprint detection, browser automation backend and integrations on Django and Flask, queues, Redis, PostgreSQL development infrastructure: deployment, environments, CI/CD Besides Python I build embedded devices in C++, so I have a good understanding of security and how systems work at a low level.\nWork experience # ProtoKey, own project # Developer · 2025 - present\nHardware password manager. Took the project from idea to preorders on Planeta.ru (Russian crowdfunding platform).\nDeveloped C++ firmware for ESP32-S3: 3.5\u0026quot; touchscreen, password input over USB HID (the device works as a keyboard), web interface over Wi-Fi. Implemented data protection: AES-256 encryption, Secure Boot and flash encryption, fully autonomous operation without any cloud. Learned circuit design, soldering and hardware debugging on my own. Tsifra # Python Developer · February 2022 - August 2024\nCore team of an international enterprise system for monitoring mining equipment. The Django web panel receives sensor readings, manages their state and visualizes data, the system generates analytical reports for mines.\nMoved the KVS storage from Django ORM to Redis. Wrote a custom database manager that redirects ORM queries to Redis through configuration, with no changes to application code. On my own initiative built a project deployment tool. Because of the many layout options, the local environment used to be assembled slowly and by hand. After rollout the project deploys with one command on a clean Linux, and switching between production environment variants works the same way. Frontend developers became able to keep their test backends up to date on their own. New features, bug fixes, refactoring, test coverage, code review. KRIPTON # Python Developer · June 2019 - January 2022\nBreachReport project, a forum monitoring startup. Responsible for the data collection system, from design to operation.\nDesigned and built a system collecting data from 100+ forums as a set of microservices. Created a Scrapy based framework: a universal forum crawling algorithm, plugins for user behavior emulation and captcha bypass, a convenient API for developers. Adding a new forum came down to configuring ready made tools, which noticeably sped up scaling. Developed stateless microservices for data processing, media downloading and browser management on Flask and Django with Dramatiq and RabbitMQ queues. Load up to 10 million messages per day. Compared an async API against queues and chose queues as the more efficient solution for this load. Took part in setting up and supporting servers and CI/CD. Freelance # Python Developer · January 2016 - February 2019\nTurnkey development: discussing the task with the client, writing the spec, estimating time and cost, implementation and delivery.\nScrapers and bots for websites, Telegram bots, websites, system utilities for Linux. Skills # Backend: Python, Django, Flask, PostgreSQL, Redis, RabbitMQ, Dramatiq, Celery Data collection: Scrapy, Playwright, Selenium, anti-bot and captcha bypass Infrastructure: Linux (Debian, Ubuntu), Docker, Git, CI/CD Practices: TDD, microservice architecture, code review Embedded: C++, ESP32, USB HID, cryptography Languages: Russian (native), English (B2, conversational) Professional activity # ru.stackoverflow: 140+ answers, 2700+ reputation, questions on Python, Django and scraping GitHub: tools for Selenium and hCaptcha bypass Medium: 9 technical articles on Selenium, captcha bypass, caching in Python and Linux Twitch: development streams, from Python backend to C++ firmware Education # Samara State Transport University\nApplied Mathematics and Computer Science, incomplete higher education\n","date":"8 October 2026","externalUrl":null,"permalink":"/en/cv/","section":"Alexander Sokolov","summary":"","title":"CV: Senior Python Developer \u0026 Web Scraping","type":"page"},{"content":"These are the rules I write code by. They are based on Robert Martin\u0026rsquo;s \u0026ldquo;Clean Code\u0026rdquo; and PEP 8 for Python. It\u0026rsquo;s a living document, I extend it as I learn things the hard way.\nThe main idea is simple: code is read far more often than it is written. So I write it for the person who opens the file six months from now. Often that person is me.\n0. No spec, no result # Any work starts with requirements, not with code.\nFirst I pin down what problem we\u0026rsquo;re solving and how we\u0026rsquo;ll know it\u0026rsquo;s solved. Then the architecture: components, data, the boundaries between them. Only after that, code. If requirements change along the way, I update the spec first, then the code. 1. Code style: PEP 8 + my rules # PEP 8 is the baseline standard. My rules work on top of it, like the cascade in CSS: everything I haven\u0026rsquo;t overridden is inherited from PEP 8 as is. There\u0026rsquo;s one override.\nTabs instead of spaces. One level of nesting is one character. Everyone sees indentation at the width they set in their editor, and nothing changes in the file. I never mix spaces and tabs in one file, Python 3 won\u0026rsquo;t allow it anyway.\nEverything else follows PEP 8: line length, blank lines between functions and classes, import order, spaces around operators.\nSo I don\u0026rsquo;t have to keep this in my head, a tool checks and fixes the style. I use ruff:\n# pyproject.toml [tool.ruff] line-length = 100 [tool.ruff.format] indent-style = \u0026#34;tab\u0026#34; [tool.ruff.lint] select = [\u0026#34;E\u0026#34;, \u0026#34;F\u0026#34;, \u0026#34;I\u0026#34;, \u0026#34;N\u0026#34;, \u0026#34;B\u0026#34;, \u0026#34;UP\u0026#34;] ignore = [\u0026#34;W191\u0026#34;] # W191 complains about tabs, for us it\u0026#39;s a deliberate choice IDE: JetBrains (PyCharm). The project settings have tabs enabled and ruff runs on save.\n2. Names # A name should answer three questions: why it exists, what it does and how it\u0026rsquo;s used. If a name needs a comment, it\u0026rsquo;s a bad name.\nVariables and functions in snake_case, classes in PascalCase, constants in UPPER_CASE. Functions are named with a verb: fetch_page, parse_price, send_report. Classes and variables with a noun: Browser, proxy_pool. No abbreviations or single-letter names. Exceptions: i in a short loop and e for an exception. One concept, one word. If the project has fetch, then get, load and retrieve don\u0026rsquo;t show up next to it for the same thing. No type encoding in names: users, not users_list or lst_users. Boolean variables read like a question: is_ready, has_proxy, can_retry. 3. Functions # Small. A function should fit on the screen without scrolling. If it doesn\u0026rsquo;t, it can be split. Do one thing. If a function can\u0026rsquo;t be described in one sentence without the word \u0026ldquo;and\u0026rdquo;, it\u0026rsquo;s two functions. One level of abstraction. A function either orchestrates a process and calls other functions, or works with details. Not both at once. Few arguments. Ideally zero to two. Three or more is a reason to group them into a dataclass. No flag arguments. render(page, True) is hard to read. Better two functions: render_full and render_preview. No hidden side effects. If a function is called check_session, it shouldn\u0026rsquo;t create a new session along the way. Command or query. A function either changes state or returns data. Not both. 4. Comments # The best comment is the one that wasn\u0026rsquo;t needed because the code is clear as it is.\nI write a comment when I need to explain why, not what:\n# Cloudflare returns 403 on the first request without a cookie, so we always make two response = session.get(url) I don\u0026rsquo;t write:\ncomments that retell the code commented-out code. That\u0026rsquo;s what git history is for a change log and authorship in the file header. That\u0026rsquo;s git\u0026rsquo;s job too A docstring is required for a module\u0026rsquo;s public functions and classes. For internal ones, it depends.\n5. Error handling # Exceptions instead of return codes. A function doesn\u0026rsquo;t return None or -1 when something goes wrong, it raises an exception. I catch specific exceptions, never a bare except: or except Exception: pass. My own exceptions for my domain: ProxyBannedError, CaptchaError. They tell you what happened without reading the stack trace. I don\u0026rsquo;t return None where a collection is expected. An empty list is better than a None check at every call site. I handle an error where I know what to do with it. If I don\u0026rsquo;t know, I let it go up. 6. Classes and modules # Small classes with a single responsibility. A class should have one reason to change. If a class talks to the network, parses HTML and writes to the database, that\u0026rsquo;s three classes. High cohesion. A class\u0026rsquo;s methods work with its fields. If a method doesn\u0026rsquo;t touch self, it most likely doesn\u0026rsquo;t belong in this class. Dependencies from outside. A class receives the browser, session or database client in its constructor instead of creating them itself. That makes it easy to test and swap. Law of Demeter. I talk only to immediate neighbors. order.customer.address.city is a sign that a method is needed. Boundaries with third-party code. I wrap third-party libraries in my own thin layer. If the library has to be replaced, the change happens in one place. 7. Tests # Tests are code too, and the same cleanliness requirements apply. A test checks one thing, and its name shows it: test_returns_empty_list_when_page_has_no_items. Every test has the same structure: arrange, act, assert. Tests are fast, independent of each other and give the same result on every run. No trips to the real internet in unit tests. I reproduce a bug with a test first, then fix it. 8. The Boy Scout rule # I leave code cleaner than I found it. There\u0026rsquo;s no need to rewrite everything at once. It\u0026rsquo;s enough to rename one unclear variable or extract one piece into a function every time I touch a file.\n9. Git # One commit, one logical change. Commit messages in the conventional commits format: feat:, fix:, refactor:, docs:, chore:. Code that doesn\u0026rsquo;t pass the linter and tests doesn\u0026rsquo;t get into the main branch. Secrets, keys and .env never get into the repository. ","date":"8 October 2026","externalUrl":null,"permalink":"/en/code-style/","section":"Alexander Sokolov","summary":"","title":"Software Development Standards","type":"page"},{"content":"Today I launched sokolab.xyz, a landing page and a blog. Everything runs on my own server in Docker. Here I\u0026rsquo;ll publish reports on my work, useful things about computers and internet technologies, maybe some reviews and lecture courses.\nThe home page is a landing page with a summary about me and my statuses. The blog page has all posts in chronological order. And if a course appears, or my work turns into some kind of product, it will get its own subdomain where it can live its own life.\nHow it works # The site is built with Hugo, the sources live in my Gitea.\nI write a post in Markdown; I do git push; a minute later it\u0026rsquo;s on the site. ","date":"5 October 2026","externalUrl":null,"permalink":"/en/blog/hello/","section":"Blog","summary":"How I set up a website and Gitea on my own server","title":"I launched my website","type":"blog"},{"content":"","externalUrl":null,"permalink":"/en/authors/","section":"Authors","summary":"","title":"Authors","type":"authors"},{"content":"","externalUrl":null,"permalink":"/en/categories/","section":"Categories","summary":"","title":"Categories","type":"categories"},{"content":"","externalUrl":null,"permalink":"/en/series/","section":"Series","summary":"","title":"Series","type":"series"}]