F15 — Two unrelated frontend weaknesses. A render error unmounted the whole React tree: a white window, no navigation, no indication of what happened, and the only way out was knowing to reload. ErrorBoundary now shows the message with a retry and a reload, and clears itself when resetKey (the route) changes, so navigating to a working page just works instead of staying stuck. There are two: one inside AppShell around the routed pages, one at the root for the shell itself and the login screen, which sit outside it. The bundle was one 841 kB file (234 kB gzipped) every visitor downloaded in full, with Vite warning about it on every build. Routes are lazy now and the entry chunk is 377 kB (120 kB gzipped) — a 55% cut, warning gone. Measuring first changed what to split. Monaco turned out not to be in the bundle at all: @monaco-editor/react loads it from cdn.jsdelivr.net, so only the small wrapper ships. xterm.js *is* bundled, all 294 kB of it, and it was reachable from ContainerCard — which renders on every stack detail page — so every visitor paid for a terminal most never open. It is lazy now and lands in its own chunk. (Worth knowing separately: the compose editor therefore needs jsdelivr.net reachable. For a self-hosted tool on an air-gapped network that is a real limitation, but vendoring Monaco means +3 MB and is its own change.) Adding a boundary whose behaviour I could only reason about was not good enough, and the missing frontend test runner was already flagged as the gap from 0.49.0. So this also sets up vitest + jsdom + testing-library and covers the boundary: that it renders the error rather than a blank page, offers a way out, clears on navigation, and stays put on an unrelated re-render. CI runs `npm test` next to pytest. One snag worth recording: installing the dev dependencies triggered npm's optional-dependency pruning and dropped @rollup/rollup-linux-x64-gnu, which broke the build. Reinstalling it directly put a linux-x64-glibc binary in package.json, which would have broken `npm ci` on every other platform — so that was backed out and the lockfile now carries the bindings as rollup's optional deps, where they belong. Verified with a clean `npm ci` in a scratch copy: install, typecheck, build and test all pass from the committed lockfile. 758 backend tests, 7 frontend tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Dk43rmEeRfYi5wsLDbmfyG
37 KiB
StackPilot
A self-hosted Docker Compose manager for power users and homelab enthusiasts — as intuitive as Dockge, as capable as Portainer for Compose workflows.
Status: Phase 1 (Core) + Phase 2 (Volumes & GPU) + Phase 3 (Quality of Life) + Phase 4 (Operations) + Phase 6 (Backup destinations) + Phase 7 (Scheduled backups) + Phase 9 (Networks) + Phase 10 (iGPU passthrough) + Phase 11 (Network attach) + Phase 12 (File browser) + Phase 13 (Network & image management) + Phase 15 (Dashboard stack resource usage) + Phase 16 (Volumes page) + Phase 18 (Image prune) + Phase 19 (Compose validate & diff) + Phase 20 (Container management) + Phase 21 (Container terminal) + Phase 22 (Auto-update) + Phase 23 (Secrets & configs) + Phase 24 (Design System v2) complete.
Upgrading to 0.50.0 — nothing to do
Two robustness fixes, no configuration changes.
- A render error no longer blanks the window. Any exception thrown while rendering used to unmount the whole React tree: a white page with no navigation and no clue what happened. An error boundary now shows what broke with a way out, and clears itself when you navigate to another page.
- The bundle is split. It was one 841 kB file every visitor downloaded in full; the entry chunk is now 377 kB (120 kB gzipped) with each page fetched on first open. xterm.js, at 294 kB the single largest piece, loads only when somebody actually opens a container terminal.
Upgrading to 0.49.0 — nothing to do
The UI now refreshes when Docker changes instead of asking every few seconds.
One WebSocket (/ws/events) carries container, image, network and volume
events; the client drops the matching query caches and re-renders. The polling
intervals stay behind it as a safety net for a dropped socket, at 30–60s rather
than 5s.
Live CPU and memory keep their fast poll on purpose — usage drifts continuously and Docker emits no event for it. That request is served from a 4-second server-side cache, so it costs one sample per interval no matter how many tabs are open.
The endpoint existed before this release but nothing used it, and it was not
usable as it stood: it forwarded every event including three exec_* per web
terminal session, never said what had changed, and leaked its reader thread on
every disconnect. All three are fixed.
Upgrading to 0.48.0 — remote hosts are gone
The multi-host feature (the stackpilot-agent sidecar and everything that
proxied to it) has been removed. StackPilot now manages exactly one Docker
host: the one it runs on.
If you never registered a remote host, nothing changes for you. Otherwise:
- Stop and remove your
stackpilot-agentcontainers. They will simply sit there unused; nothing talks to them any more. Thestackpilot-agentimage is no longer built or published. - Registered hosts are deleted from the database on first start, together with the bearer tokens they held. That is deliberate: leaving credentials for a feature that no longer exists is worse than dropping them.
- Backup schedules that targeted a remote stack will fail with "stack not
found" until you delete them under Settings → Scheduled backups. Their
agent_idcolumn is dropped where SQLite supports it. - Stacks that lived on a remote host are untouched on that host — StackPilot
just no longer sees them. Their compose files are still in that host's stacks
directory and
docker composestill works there, which is the point of the file-is-the-truth model.
Gone with it: the host switcher on the Files page, the per-host sections on
Stacks / Networks / Images / Volumes / Dashboard, the host selector in the New
Stack editor and the template dialog, Settings → Remote hosts, and the
/api/agents/* and /ws/agent-* endpoints (57 API routes and 3 WebSocket
routes in total).
Upgrading to 0.47.0 — nothing to do
Three pieces of state moved out of process memory and into the database, and the live-stats endpoint got a cache. No configuration changes, no migration steps; the new tables are created on first start.
- Stacks can only run one compose operation at a time. A second
start/update/downon a busy stack answers409instead of racing the first one over the same containers. Auto-update skips a stack you are already deploying and picks it up next cycle. GET /api/stacks/statsis cached for four seconds. It sampled every running container on every call, and the dashboard and the stacks list both poll it every five seconds — two open tabs on a 40-container host meant a sustained ~16 daemon calls a second.- The update cache and the login rate limiter persist. Both used to reset on
restart; the rate limiter also used to multiply by the worker count, so
neither behaved as documented with
--workersset.
Upgrading to 0.46.0 — everyone is signed out once
Sessions now hold their access token in memory and the refresh token in an httpOnly cookie, so the upgrade signs everybody out exactly once. Sign back in and it behaves as before, including staying signed in across restarts.
What changed and why:
- Tokens can be revoked. Each account carries a
token_versionthat every token is minted with and every request checks. Resetting a password, changing a role or disabling an account now bumps it, which cuts off the tokens that account already holds — previously a password reset was cosmetic and whoever had the old refresh token kept full access for up to 30 days. - The refresh token left
localStorage. It is an httpOnly cookie (SameSite=Lax, scoped to/api/auth), so a successful XSS can act inside the open page but cannot walk off with 30 days of access. The cookie is markedSecureonly when the request arrived over HTTPS, so a plain-HTTP homelab keeps working. Any refresh token left inlocalStorageby an older build is deleted on first load. - "Sign out everywhere" in the user menu revokes every token the account holds, on every device. Plain "Sign out" only ends the session on that device.
- WebSockets re-check the database. Log streams and the container terminal read the role from the live user instead of the token's claim, so a demotion or a disabled account takes effect immediately — the terminal is root-equivalent on the host.
Scripted clients that cannot hold a cookie can still get the refresh token in
the response body with ?in_body=true on login and refresh.
Upgrading to 0.44.0 / 0.45.0 — two defaults changed
0.44.0 closes a privilege-escalation hole and tightens two defaults. Both changes can affect an existing install:
- The
userrole loses read access to secrets. The file browser (page and/api/files/*), the host-path picker, the audit log,GET /api/stacks/{id}/export, a stack's.envand the template detail route (0.45.0 — "save stack as template" snapshots the stack's real.env) are now admin-only. The template listing stays open. Previously any logged-in account could downloadstackpilot.db, every.envand every.secrets/*file — and none of it was audit-logged. If you gave someone auseraccount so they could look at stacks, they still can; they just no longer get the credentials. Nothing changes for admins. /is no longer a default browse root. The new default is/mnt,/media,/srv,/opt,/home. A/entry makes the sandbox allow every path, which is why it is gone — if you relied on it, setALLOWED_BROWSE_ROOTSexplicitly in your.env. StackPilot's ownDATA_DIRis refused either way.
Two things also get fixed without any action on your part: the backend now runs
uvicorn with --proxy-headers, so the login rate limit works per client IP
instead of globally and the audit log records real IPs; and backup-destination
credentials are encrypted at rest, with existing rows migrated on first start.
That encryption is keyed off SECRET_KEY, which is now persisted to
${DATA_DIR}/secret_key when you have not set one — so restarts no longer log
everyone out. If you have never set SECRET_KEY, do not delete that file;
it is what your saved destination credentials are encrypted with.
What works today (Phase 1)
- File-first stacks — every stack is a plain
compose.yaml(+ optional.env) on disk. The DB only stores metadata; nothing is locked in. - Auth — JWT access/refresh tokens, bcrypt hashing, admin/user roles, and a
first-launch setup wizard that creates the initial admin account. The access
token is held in memory; the refresh token is an httpOnly cookie. Every token
carries the account's
token_version, so a password reset, role change or disable revokes the tokens that account already holds — on every device. - Stack lifecycle — create, edit, clone, delete, and
up / down / start / stop / restart / pull / updateviadocker compose. - Live status — running / partial / stopped / error / updating, computed from Docker container labels.
- One operation per stack — a lifecycle call takes a lock (a row, so it
holds across workers and across a restart) and a second one gets
409while it is held; auto-update skips a stack somebody is already deploying. Locks carry an expiry, so a worker killed mid-deploy does not strand a stack. - Resilient UI — a render error shows what broke and offers a way out instead of blanking the window, and clears itself when you navigate away. Routes are code-split, so the entry bundle is 377 kB rather than 841 kB and the container terminal's xterm.js only loads when a terminal is opened.
- Event-driven UI — a single
/ws/eventsconnection carries Docker's own container / image / network / volume events; the client drops the matching caches so pages refresh the moment something changes, instead of every page polling on a timer. The intervals remain as a slow fallback. Live CPU and memory still poll, because usage drifts with no event to announce it. - Real-time logs — streamed over WebSocket, color-coded per service.
- Live deploy console — deploying from the editor streams
compose upoutput (image pulls, container creation) over a WebSocket in real time instead of a blind spinner; the deploy keeps running server-side if the modal is closed. stream through to the browser. Compose runs with--progress json, so the console shows a real progress bar (download bytes per layer, weighted by layer size, then container create/start) plus a per-image bar; the raw output is kept below it and updates one line per layer instead of scrolling past. - Inline update progress (0.42.0) — hitting Update on a stack streams
compose pull && up -dover/ws/update/{stack_id}and folds it into a progress bar inside that stack's row (same--progress jsonweighting as the deploy console: "Pulling images · 3/7 layers · 88 MB / 190 MB"). Several stacks can update at once — busy state and progress are tracked per stack, so the rows advance independently. The RESTPOST /api/stacks/{id}/updatestays for non-interactive callers and as the fallback when no token is available. - Monaco editor — YAML editing with an
.envtab and adocker run→ compose converter. - Dashboard — system resource bar, stack grid with quick actions, and a recent-activity audit feed.
- Auto-discovery — stacks created outside the UI (any folder under the stacks dir containing a compose file) are picked up automatically.
- Dark / light theme.
Phase 2 — Volumes & GPU
- Volume Wizard in the editor (right-hand helper panel): Bind / Named / NFS /
SMB-CIFS / tmpfs, with a sandboxed host path browser for bind mounts and a
live YAML preview. NFS/SMB
driver_optsare generated for you. - GPU assignment per service: auto-detects NVIDIA (
nvidia-smi) and AMD/Intel (/dev/dri+ sysfs), injects the right YAML (NVIDIAdeploy.reservations, or/dev/dripassthrough + render/video groups +LIBVA_DRIVER_NAME=iHDfor Intel). - Device passthrough: lists host USB / serial-TTY / DRI nodes, add per service,
plus a guarded
privilegedtoggle. - All wizard edits are merged into the compose YAML server-side (robust, validated) and returned to the editor for review before saving.
GPU/device detection needs host visibility. The bundled compose bind-mounts
/dev:/dev:ro; NVIDIA additionally requires the NVIDIA container runtime on the host.
Phase 3 — Quality of Life
- Env editor: table mode with sensitive-value masking (
PASS/SECRET/KEY/… auto-detected) + raw mode, plus PUID/PGID/TZ quick-insert. - Image update checker: background task compares the local manifest digest with the registry (Docker Hub / ghcr / lscr / private v2 with token auth); update badges on the Images page + an "updates available" banner on the dashboard. Results are cached in the database, so a restart shows the badges immediately instead of blanking them until the next sweep — and does not re-announce updates it already notified about.
- Port conflict detector: pre-deploy check against host-bound ports
(
/proc/net/tcp[6]) and running container bindings, with a confirm dialog. - Resource limits: CPU/memory sliders in the editor →
deploy.resources.limits. - Template library: each template is a ready-to-run stack folder in git
(
backend/templates/<slug>/—compose.yaml+.env.example+template.json, plus any extra config the app needs, e.g.prometheus.ymlor aCaddyfile). 83 bundled homelab apps across media, *arr automation, networking, reverse proxies, VPN, SSO, monitoring, dashboards, files/backup, notes and wikis, home automation, dev tooling, databases, local AI and more — searchable and filterable by tag on the Templates page. "Pull" copies the whole folder into a new stack (.env.example→.env) which you then edit and deploy. Save any stack back as a custom template (stored under${DATA_DIR}/templates/). Add your own by dropping a folder into the templates dir. - Healthcheck status surfaced per container in the stack overview.
Phase 4 — Operations
- Self-update (0.32.0): the top-bar version badge checks the registry for a
newer StackPilot release on page load (
GET /api/system/update, anonymous v2 token flow, 10 min cache) and shows an amber update pill. One click (POST /api/system/update, admin) spawns a detached helper container that runsdocker compose pull && up -don StackPilot's own compose project (resolved from its container labels) — the helper survives the backend being recreated; the UI polls/api/healthand reloads when the new version answers. Installs not managed by compose get a clear "update manually" error instead. - Backup & restore: per-stack
.tar.gzbackups covering the whole stack — the stack folder, every bind-mounted data directory (./config,/mnt/appdata/…) and every named volume. Bind sources and volumes are read through a throwaway helper container, i.e. by host path, so data that StackPilot itself cannot see is captured too (that is the case wheneverSTACKS_HOST_DIRdiffers from the container'sSTACKS_DIR— compose then creates the data directories at the container path on the host, and a naive backup would only find the compose file). The Backup dialog shows the full inventory with sizes and lets you pick what goes in; NFS/CIFS-backed volumes are unchecked by default because they live on a NAS and restoring one would overwrite the share. Restore puts bind folders back at their host paths and preserves permissions, ownership and symlinks (PUID/PGID-based images such as the *arr suite need this), with optional rename and overwrite/conflict detection. - Notification webhooks: ntfy, Discord, Slack, Gotify, or generic JSON, each
subscribed to chosen events (image update available, stack start/stop/error,
pull failed). Managed in Settings → Notifications; env
NOTIFY_WEBHOOKSstill supported for generic endpoints. - Settings page: tune the update-check interval, manage webhooks, and manage users (create/disable/delete, promote/demote, with last-admin safeguards).
- Audit log page (admin): searchable, paginated view of all recorded actions.
- Mobile-responsive layout: off-canvas sidebar + adaptive spacing.
Phase 6 — Backup destinations
- Off-box backups: define SFTP, S3-compatible (MinIO, Backblaze B2,
AWS S3, …) or NFS share (0.33.0) destinations under Settings → Backup
destinations (with a Test button; secrets are masked in API responses).
NFS needs no privileges in the StackPilot container: the Docker daemon mounts
the export as a named volume (
stackpilot-nfs-dest-<id>, recreated when the config changes) and file I/O runs through throwaway helper containers. - Push & restore: the stack Backup dialog can push straight to a destination instead of downloading; the Restore dialog can browse a destination's backups and restore (volumes included) directly from it. Backups can also be deleted from the UI.
Phase 7 — Scheduled backups
- Recurring backups: schedule a stack to back up to a destination hourly, daily, or weekly (UTC) under Settings → Scheduled backups. A background scheduler runs due jobs every minute and records last/next run + status.
- Retention: keep the newest N backups per stack on the destination; older ones are pruned automatically.
- Run now for an on-demand run, plus a
backup_failednotification event wired into the webhook system.
Phase 9 — Networks
- Network management: the Networks page lists Docker networks (driver, scope, subnet, attached containers / in-use, owning stack), with create (bridge / macvlan / ipvlan / overlay, optional subnet+gateway, internal/attachable), delete (default networks protected; in-use guarded by Docker), and prune unused.
- Stack delete: local stacks can now be deleted from the UI (stack detail and the stack card), with a confirm dialog and an optional "keep files on disk".
Phase 10 — iGPU passthrough
- Render/video group detection: for a passed-through Intel/AMD iGPU, StackPilot
reads the host group ownership of the
/dev/drinodes (render node →renderGID, pairedcardnode →videoGID) and injects them as numericgroup_addentries (e.g.group_add: ["991", "44"]). Group names rarely resolve inside images, so the numeric GID is what actually grants access. The GPU selector shows the detected GIDs; removal cleans them (andLIBVA_DRIVER_NAME).
Phase 11 — Network attach
- Network attach/detach: each network row on the Networks page expands to an
inspect view listing connected containers, with admin controls to disconnect a
container or connect any container on the host (
POST /api/networks/{id}/connect//disconnect).
Phase 24 — Design System v2 (analytics-style UI)
- New shell: the sidebar is gone — a fixed 60px top bar carries a pill navigation (active route = dark pill), the logo mark, a version badge, a theme toggle and an avatar menu. Narrow screens get an off-canvas drawer.
- Design tokens (
frontend/src/styles/tokens.css): one CSS-variable set for surfaces, borders, text tiers, brand colours, radii and type scales, with class-based dark-mode overrides. Existing Tailwind aliases (bg/card/accent) are remapped onto the tokens so all pages reskin consistently. Typeface: Schibsted Grotesk (bundled, offline-friendly). - Stack Health funnel — the dashboard centrepiece. Five stages
(
discovered → running → healthy → updated → auto-managed) fromGET /api/dashboard/funnel(30 s server cache,?refresh=trueto bust; the last stage — API keymonitored— counts stacks with an enabled auto-update policy). SVG waterfall with alternating gradient / diagonal-hatch bars, value chips and hover conversion/drop-off tooltips. - Summary widgets from
GET /api/dashboard/summary: compose-containers card and an "Insights" chip (healthy-rate %), a 30-day uptime line chart (sampled every 5 min intoDATA_DIR/uptime.jsonl, charted as daily averages), and an ops contribution grid from audit-log activity with the peak weekday. A 30/7-day range selector slices both series client-side. - Explore prompt bar under the funnel: typed queries or
/running,/stopped,/attentiontags deep-link to the Stacks page, which now honours?q=and?filter=(plus a new status-filter select).
Phase 23 — Secrets & configs (compose file-based)
- A Secrets tab on the stack detail page manages per-stack Docker secrets and configs: create one by name + content, list them (name, kind, size — content is never returned by the API), and delete. Content is write-only: once saved it is cleared from the form and cannot be read back.
- Files are stored inside the stack's own directory (
<stack_dir>/.secrets/<name>/.configs/<name>, dir0700/ file0600) and referenced from the compose file with a relativefile:path, so the Docker daemon reads them with noHOST_ROOT_PREFIXdependency — exactly as if dropped next tocompose.yaml. - Attach/detach wires a stored secret/config into a chosen service: secrets
appear at
/run/secrets/<name>, configs mount at a target path you specify. The compose file is rewritten in place (top-levelsecrets:/configs:defs are pruned when no service still uses them); redeploy the stack to apply. - Admin-only (secrets are sensitive); every write/delete/attach/detach is
audited (
secret.*). Names are validated against path traversal (single component, no.., no leading dot); content capped at 1 MiB.
Phase 22 — Auto-update (Watchtower-style)
- A per-stack Auto-update policy (on the stack Overview tab): when the background image-update check finds a newer registry digest for one of the stack's images, the stack is either pulled + redeployed or merely flagged ("notify only"), with a Check now button for an on-demand run.
- Runs inside the existing image-update-check cycle (reuses the freshly-computed digest cache, no extra registry calls). Only running stacks are auto-redeployed — a stopped stack is never silently started ("skipped").
- New
stack_auto_updatednotification event. Last run + status (updated / up-to-date / update-available / skipped / error) are shown inline.
Phase 21 — Container terminal (web exec)
- An interactive terminal into any running, compose-managed container,
opened from the terminal button on its container card (stack Overview tab).
Streams an exec session over WebSocket into xterm.js — pick
/bin/sh,/bin/bash, or/bin/ash; full TTY with resize. - Admin-only (exec is root-equivalent): a non-admin token is rejected at the
WebSocket handshake (
4403). Only containers with thecom.docker.compose.projectlabel can be reached.
Phase 20 — Container management
- The stack Overview tab now renders each service as an expandable
container card instead of a static row. Expanding it fetches a curated
single-container inspect view (image, state + exit code, restart count,
started-at, networks, mounts, and environment) via
GET /api/containers/{id}. - Admins get per-container start / stop / restart buttons directly on the
card (
POST /api/containers/{id}/{action}), so a single misbehaving service can be bounced without touching the rest of the stack. - Only containers carrying the
com.docker.compose.projectlabel are exposed, so this never becomes a generic "control any container on the host" backdoor.
Phase 19 — Compose validate & diff
- The editor can validate a compose file (
docker compose config) before deploying and show a diff against the currently deployed definition, so you can see exactly what a re-deploy will change.
Phase 18 — Image prune
- Prune images (dangling, or all unused) from the Images page, on the local host.
Phase 16 — Volumes page
- New Volumes page (sidebar). Lists Docker volumes with driver, owning stack, in-use containers and mountpoint.
- Admin actions: delete a volume (with an in-use warning + force option) and Prune unused; an Only unused filter.
- Volume sizes are loaded on demand via a Compute sizes button (runs
docker system df, which walks volume contents and can take a few seconds); results are cached ~60s. EndpointGET /api/volumes/sizes. - The Volume Wizard in the stack editor (bind/named/NFS/SMB/tmpfs YAML generation) is unchanged — the new page is for managing/cleaning up volumes.
Phase 15 — Dashboard stack resource usage
- The dashboard now lists stacks in a table (status, services) with live
CPU and memory usage per stack, sampled from
docker statsand aggregated by compose project. - When a stack has
deploy.resources.limitsassigned, the bar fills toward that limit and shows usage vs the limit (e.g.0.42 / 1 cores,310 MB / 512 MB); otherwise it shows absolute usage against the host total. Inline start/stop/ restart actions per row for admins. New endpointGET /api/stacks/stats.
Phase 13 — Network & image management
- Network management: list, inspect, create, delete, prune, and connect/disconnect containers — including a Prune unused button, which resolves the common "all predefined address pools have been fully subnetted" deploy error without SSH.
- Images: list image tags (with using-stacks) and run on-demand update checks.
Phase 12 — File browser
- Files page (sidebar): a full host filesystem browser with breadcrumb navigation, clickable browse-root chips, an Up control, and a show/hide hidden-files toggle. Listings show size, permissions and modified time.
- View & edit: clicking a text file opens it in a Monaco editor (with syntax highlighting picked from the extension). Binary and oversized files are detected and offered as a download instead. Admins can edit and Save.
- Admin only: the whole page, including listing, viewing and downloading.
Reads are not less sensitive than writes here — the browser reaches whatever
the backend container can see, which includes every stack's
.envand.secrets/*. Reading and downloading a file are audit-logged (file.read,file.download); directory listing is not, because the page polls it. - Manage (admin): create folders/files, rename, delete (recursive for folders), upload files or whole folders (the directory tree is recreated server-side), and download any file. Copy/cut & paste moves files and folders between directories (clipboard bar + per-row copy/cut, with an overwrite prompt on conflict). Every mutation is audit-logged.
- Sandboxed: all access is confined to
ALLOWED_BROWSE_ROOTS; path traversal and deleting a browse root are refused. StackPilot's ownDATA_DIRis refused regardless of the setting — it holdsstackpilot.dbwith password hashes and backup-destination credentials, none of which the API itself ever hands out. Note that a single/entry inALLOWED_BROWSE_ROOTSswitches the sandbox off entirely; it is no longer part of the default. To reach the real host filesystem, mount it into the backend and setHOST_ROOT_PREFIX(see the commented/:/host_rootvolume indocker-compose.yml). Endpoints live under/api/files/*(list,read,write,mkdir,touch,rename,copy,move,upload— with optionalrel_pathfor folder uploads —,download,DELETE).
Architecture
frontend (React + Vite + Tailwind, served by nginx)
│ proxies /api and /ws
▼
backend (FastAPI + docker-py + SQLite)
│ docker-py + `docker compose` CLI
▼
Docker Engine (via /var/run/docker.sock — never exposed to the browser)
Quick start
cd stackpilot
cp .env.example .env
# edit .env and set a strong SECRET_KEY: openssl rand -base64 48
docker compose up -d --build
Open http://localhost:5009 and complete the first-launch setup wizard to create your admin account.
Configuration
All backend settings are environment variables (see backend/config.py). The
most important ones:
| Variable | Default | Purpose |
|---|---|---|
SECRET_KEY |
(auto, persisted) | JWT + at-rest encryption key (see below) |
STACKS_DIR |
/opt/stacks |
Where stack folders live (in-container) |
DATA_DIR |
/data |
SQLite DB + app data |
CORS_ORIGINS |
localhost | Allowed API origins (comma separated) |
SECRET_KEY signs JWTs and derives the key that encrypts backup-destination
credentials in the database. Leave it unset and one is generated and written to
${DATA_DIR}/secret_key (mode 0600) on first start, so sessions and stored
credentials survive restarts — that file is then part of your backup. Setting it
explicitly always wins and nothing is written.
The host path for stacks is set via STACKS_HOST_DIR in .env, and it should
be the same path as STACKS_DIR (/opt/stacks by default). Compose runs
inside the backend container, so a stack's relative bind mounts (./config) are
resolved against the container path and the daemon creates those directories at
that path on the host. Point STACKS_HOST_DIR somewhere else and every stack's
data lives at /opt/stacks/<stack>/… on the host while StackPilot looks at a
different folder — the file browser and editor then show only the compose file.
Backups cover the data either way (they read bind sources by host path through a
helper container) and the Backup dialog warns when the two paths diverge.
Local development
Backend:
cd backend
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
SECRET_KEY=dev STACKS_DIR=../data/stacks DATA_DIR=../data uvicorn main:app --reload --port 5008
Frontend (proxies to the backend on :5008):
cd frontend
npm install
npm run dev # http://localhost:5173
Tests & linting
Same three commands the CI runs — build-and-push only starts once they pass.
cd backend
pip install -r requirements-dev.txt
pytest # 758 tests, no Docker daemon needed
ruff check .
cd ../frontend && npx tsc --noEmit -p tsconfig.json && npm test
The frontend has a small vitest suite alongside it (npm test, jsdom). It
covers behaviour the typechecker cannot see — currently the error boundary:
that it renders the error instead of a blank page, offers a way out, and clears
on navigation so one broken page does not strand you.
The backend suite drives the app through TestClient without the lifespan, so it
never opens a Docker socket and never starts the background loops; conftest.py
points DATA_DIR/STACKS_DIR at a temp directory before anything is imported.
The load-bearing one is tests/test_route_authorization.py. Authorization lives
in the routers — each route independently picks require_admin or
get_current_user, and nothing checked that the choice was right, which is how
0.43.0 shipped a read-only role that could download the auth database. That file
states the policy once — every route requires admin unless it is listed — and
fails on any route that disagrees. Adding a route the user role may reach means
adding it to USER_READABLE with a note on why it cannot return a credential.
tests/test_token_revocation.py covers the revoke switch — that each of the
three authority changes kills the account's tokens, that a cosmetic re-save does
not, and that the refresh cookie is httpOnly and not marked Secure over plain
HTTP. tests/test_schema_migration.py builds a database with the old user
table and asserts the added column is backfilled rather than left NULL, which is
what would otherwise have signed out every user on every install.
tests/test_docker_events.py drives the event stream against a fake daemon: that
exec_* noise is dropped before it reaches the client, that the payload names
the resource so the client knows what to invalidate, and that the stream is
closed on disconnect — cancelling the executor future does not interrupt a
thread already inside a blocking read, so without that close every page load
leaked one.
tests/test_stack_locking.py and tests/test_runtime_state.py cover the state
that moved into the database: that a busy stack answers 409 without ever
reaching Docker, that an expired lock is taken over rather than stranding the
stack, that the stats cache serves repeat callers from one sweep, and that the
update cache and rate limiter survive a restart. One of them asserts that
update_service never imports the database — it is pure registry logic, and
persistence stays opt-in so the module remains testable without one.
tests/test_bundled_templates.py covers the 83 shipped templates: each must
parse, name an image per service, keep .env.example in sync with the variables
compose actually reads, ship every file it bind-mounts, and never come with a
working default password.
API surface (Phase 1)
POST /api/auth/setup | login | refresh GET /api/auth/me | needs-setup
POST /api/auth/logout | logout-everywhere
GET /api/stacks POST /api/stacks
GET /api/stacks/{id} PUT /api/stacks/{id} DELETE /api/stacks/{id}
POST /api/stacks/{id}/{start|stop|restart|pull|update|down|clone}
GET /api/stacks/{id}/logs GET /api/stacks/{id}/export
POST /api/stacks/convert (docker run → compose)
GET /api/system/info | gpus | devices GET /api/audit
GET /api/system/update POST /api/system/update (self-update)
WS /ws/logs/{stack_id}[/{service}] WS /ws/events
WS /ws/deploy/{stack_id} WS /ws/update/{stack_id}
Phase 2 endpoints
GET /api/volumes | /orphaned DELETE /api/volumes/{name}
POST /api/volumes/prune POST /api/volumes/generate-yaml
GET /api/host/paths?path=&show_hidden= (sandboxed browser)
POST /api/editor/services | add-volume | set-gpu | add-device | remove-device | set-privileged
Phase 3 endpoints
GET /api/images | /updates POST /api/images/check
POST /api/ports/conflicts POST /api/editor/set-resources
GET /api/templates | /{id} POST /api/templates/{id}/instantiate
POST /api/templates | /from-stack DELETE /api/templates/custom/{slug}
Phase 4 endpoints
GET /api/stacks/{id}/backup?include_volumes=&stop_first= POST /api/stacks/restore
GET /api/settings PUT /api/settings
GET /api/settings/webhooks POST /api/settings/webhooks
PUT /api/settings/webhooks/{id} DELETE /api/settings/webhooks/{id}
POST /api/settings/webhooks/{id}/test
GET /api/auth/users POST /api/auth/users
PATCH /api/auth/users/{id} DELETE /api/auth/users/{id}
Phase 6 endpoints
GET /api/backups/destinations POST /api/backups/destinations
PUT /api/backups/destinations/{id} DELETE /api/backups/destinations/{id}
POST /api/backups/destinations/{id}/test GET /api/backups/destinations/{id}/backups
DELETE /api/backups/destinations/{id}/backups/{name}
POST /api/stacks/{id}/backup/push POST /api/stacks/restore-from
Phase 7 endpoints
GET /api/backups/schedules POST /api/backups/schedules
PUT /api/backups/schedules/{id} DELETE /api/backups/schedules/{id}
POST /api/backups/schedules/{id}/run
Phase 9 endpoints
GET /api/networks | /{id} POST /api/networks
DELETE /api/networks/{id} POST /api/networks/prune
DELETE /api/stacks/{id}?delete_files= (stack delete, now surfaced in the UI)
Phase 11 endpoints
GET /api/networks/{id}/containers POST /api/networks/{id}/connect | /disconnect
POST /api/templates/{id}/instantiate (create a stack from a template)
Phase 24 endpoints
GET /api/dashboard/funnel[?refresh=true] (stack-health funnel, 30s TTL cache)
GET /api/dashboard/summary (containers, uptime series, ops activity)
Security notes
- The Docker socket is only ever touched by the backend process; it is never proxied to the browser.
- Login is rate-limited (10/min/IP).
- Compose files are backed up to
*.bakbefore every overwrite. - Generated YAML never includes the obsolete
version:field and uses Compose v2 (docker compose) syntax.