In short: FuryBee brings together about twenty public websites and applications, served by 35 containers on a single server that I run on my own. The recipe: Docker Compose, a Traefik reverse proxy, Cloudflare in front of everything, full observability with Prometheus, Grafana and Loki, and monitored encrypted backups. And above all, a list of pitfalls I'd rather share.
The architecture
Everything fits in one Docker Compose project on a 12-core server:
- Traefik receives all the traffic and routes it to the right container. Each service declares its own domain name through Docker labels, which Traefik discovers automatically. TLS certificates are issued through DNS validation with Cloudflare.
- The applications: static websites, Node.js and Nuxt applications, real-time multiplayer games, developer tools.
- The data: Redis for the applications that need it, and a dedicated PostgreSQL database for the one that handles financial data.
- Observability: Prometheus, Grafana, Loki, Alloy, cAdvisor and node-exporter.
An unknown subdomain is redirected to the main portal, and every deployment ends with a Cloudflare cache purge.
Cloudflare in front, and nothing else
Every domain goes through Cloudflare, and the server only accepts Cloudflare's IP addresses. A Traefik middleware filters the HTTPS entry point using Cloudflare's official list of IP ranges, refreshed automatically every week by a script.
In practice:
- a direct call to the server's address gets a 403 error: there's no way to bypass Cloudflare's firewall and cache;
- the visitor's real address is read from the
Cf-Connecting-Ipheader in the access logs, not from the connection address, which is Cloudflare's; - every new domain must be proxied through Cloudflare, otherwise it's blocked.
Observability
- Metrics: Prometheus collects container, server and Traefik metrics every 15 seconds, kept for 30 days.
- Logs: Grafana Alloy ships container logs to Loki, kept for 30 days, along with Traefik's JSON access logs.
- Alerts: managed by Grafana. Container rules are dynamic: adding a service doesn't require updating a list.
Backups
Every night, a script backs up the databases, Redis, the Grafana configuration and the certificates. It sends them as a single encrypted archive (AES-256) to versioned S3 storage, with a 90-day retention period. The account used for uploads cannot delete backups: a compromised server couldn't wipe the history.
The script also publishes its own metrics, and an alert fires if the last backup is more than 48 hours old. It proved its worth: two nights in a row, the upload failed because the upload tool wasn't in cron's minimal PATH. The alert went off, and the script now adds that path itself.
The pitfalls
Certificates that expire silently. After about two months without a restart, Traefik stopped renewing its certificates, with no error at all in its logs. The symptom for visitors: a Cloudflare 526 error (invalid origin certificate). A simple restart wasn't enough: the container had to be recreated, and it renewed every certificate within seconds. The fastest diagnosis: the modification date of the file where Traefik stores its certificates, frozen for weeks.
The log collector reading itself. A collector that also ships its own logs creates an infinite loop: with the previous tool (Promtail), 9 million lines in one evening. Excluding those containers with a simple filtering rule isn't enough; they have to be excluded as early as reading the list of containers.
The major Traefik upgrade (v2 → v3). I validated it with a second, test Traefik v3 on other ports, reading the same labels, using Let's Encrypt's staging environment. Only one incompatibility: the syntax of the rule that catches unknown subdomains, which now uses Go regular expressions.
Broken IPv6 slowing everything down. The server's IPv6 interface was configured but carried no traffic: some outgoing calls waited for IPv6 to fail before falling back to IPv4. Forcing an IPv4 preference brought one of those calls down from 3 minutes to 0.6 seconds.
Small rules that prevent big outages:
- pinned image versions, never
latest, so that an update is always a deliberate choice; - Docker log rotation declared service by service, for lack of admin access to the daemon;
- zero-downtime deployment for the real-time games, whose players stay connected over WebSocket: a temporary container takes over during the update.
What I take away
The same discipline as in a company applies at a small scale, without an ops team to catch what slips through: anything that isn't automated, monitored and documented eventually breaks, usually without warning. That's the experience I bring to teams who want to make their infrastructure more reliable.