An MVP, live and in front of real users, in five working days. See how →

← All work

Platform & Infrastructure Engineering

Zaltev

Zaltev is a multi-tenant automated trading platform. Every subscriber gets their own strategy engine running against a live broker connection, with the risk limits, billing, and operational tooling needed to run that fleet unattended. It is a distributed system first and a web product second — the dashboard is the smallest part of it.

{ Our Role }
Product design, full-stack engineering, and production infrastructure
{ Timeline }
Ongoing — first production deploy 2026
{ Scope }
Distributed Systems · Full-Stack Engineering · DevOps & Infrastructure · Billing & Payments
Zaltev
{ The Problem }

Automated trading products fail in boring ways. One user's strategy crashes and takes the shared process down with it. A broker connection drops overnight and nobody notices until the morning. A backtest promises one edge and the live engine trades a subtly different one. Money moves on a code path nobody tested.

Zaltev needed to run many users' engines against a real broker, continuously, with money on the line — and needed the operators to know something was wrong before the customers did.

{ Our Approach }

We treated isolation and observability as product requirements, not infrastructure chores. Each user's engine runs in its own OS subprocess so a bad strategy, a wedged broker socket, or a memory leak is contained to one account and can be killed and respawned without touching anyone else.

The second principle was a single source of truth for anything that could disagree with itself: one strategy module shared by the live engine and the backtester, one alert computation shared by the operations dashboard and the Telegram alerter, one risk gate that every trade must pass through.

{ The Solution }

The platform ships as eleven containerised services behind Caddy: two React SPAs (a user dashboard and an admin console), a FastAPI backend on multiple workers, an engine worker that owns the fleet, an ops-collector sidecar, two Telegram bots, PostgreSQL, and Redis.

Users onboard, connect their broker, pick a plan, and start an engine. Operators get a single Mission Control view spanning infrastructure health, per-user engine state, signal integrity, revenue, and support tickets — with revenue stripped out for roles that have no business seeing it.

{ Inside the Build }
01

Process-per-user engine isolation

An engine pool spawns, reconciles, and kills one OS subprocess per active user, with respawn caps so a crash loop cannot consume the box. Indicator math runs in a separate process pool so heavy pandas work never blocks the event loop. The Telegram bot deliberately runs no engines, so restarting the control plane never interrupts live trading.

02

One strategy, two consumers

The five-indicator confluence strategy (Bollinger Bands, dual ADX, Awesome Oscillator, MACD) is a pure-pandas module imported by both the live engine and the grid backtester. Backtest results describe the same code that trades, down to the edge-trigger semantics — there is no second implementation to drift.

03

Risk controls that claim before they allow

A risk manager performs an atomic check-and-claim on every entry: max concurrent trades, daily trade caps, a daily-loss kill switch, and a consecutive-loss pause. Plan entitlements are re-enforced at the runtime gate rather than trusted from the UI, and a demo/live safety guard makes trading real money an explicit decision.

04

Built to survive the network

Broker connections reconnect forever with backoff, refresh session credentials automatically, and run a candle-stall watchdog that detects a socket that is technically open but no longer delivering data. Signal-funnel telemetry records why each candle did or did not produce a trade, so a quiet engine can be explained rather than guessed at.

05

Mission Control that cannot lie

The dashboard and the Redis-locked Telegram alert loop call the same alert computation, so the page can never show green while the alerter is paging. It includes a signal-integrity check that flags same-direction streaks — the shape a strategy bug makes before it shows up in the P&L.

06

Deploys that never touch the trading box

Images build on CI runners and are pushed to GHCR with immutable prod-<sha> tags for one-command rollback; the Hetzner host only ever pulls. A host-level flock serialises deploys so a docker build's CPU spike can never land on a machine holding live broker connections. Money-path and migration tests gate every deploy.

07

Encrypted credentials and sliding sessions

Broker session credentials are Fernet-encrypted at rest and revealed only through a permission-gated admin path that writes to the audit log. Sessions re-mint their token on every request behind a 60-minute idle cap and a 7-day absolute cap, so an abandoned tab expires on its own.

08

Crypto billing with a real ledger

NOWPayments subscriptions with IPN handling, plans, trials, coupons, referrals, and a wallet ledger. A Redis-locked sweep renews subscriptions every 30 minutes and downgrades gracefully after a grace period, so billing state and entitlements never disagree.

{ Built With }

Frontend

  • React 18
  • TypeScript
  • Vite
  • TailwindCSS
  • React Query
  • Recharts

Backend

  • FastAPI
  • SQLAlchemy 2.0 async
  • PostgreSQL 16
  • Redis 7
  • Alembic

Trading Engine

  • Python
  • pandas
  • ta
  • Process pools
  • WebSockets

Infrastructure

  • Docker Compose
  • Caddy
  • Terraform
  • Hetzner
  • GitHub Actions
  • GHCR
  • Sentry

Integrations

  • NOWPayments
  • Telegram Bot API
  • Yahoo Finance
  • Anthropic Claude
  • Resend
{ The Results }
11
Containerised services in production
24/7
Unattended engine uptime
7
Permission-gated admin roles
{ Next Project }

MedFlow

View Case Study