Files
stackpilot/backend/services/stack_lock_service.py
T
menzeljandClaude Opus 5 09bed274eb
CI / check (push) Successful in 7m4s
CI / build-and-push (push) Successful in 1m47s
Lock stacks during compose runs, cache stats, persist runtime state (0.47.0)
F7 — Nothing stopped two compose operations landing on the same stack. There
was a busy flag, but is_busy() was only ever read to colour the status column;
no lifecycle handler consulted it before acting. Two tabs, or auto-update
picking up a stack somebody had just clicked, both ran pull + up -d against the
same project and raced over recreating containers.

Lifecycle calls, the two deploy WebSockets and the auto-update pass now take a
real lock; a second caller gets 409 (or an error frame and close 4409) and
auto-update skips and retries next cycle. The lock is a row rather than a set
in one worker's memory, so it holds across workers and across a restart, and it
carries an expiry — a worker killed mid-deploy would otherwise strand the stack
with no fix short of editing the database.

F10 — /api/stacks/stats sampled every running container on every call, one
blocking daemon request each, and both the dashboard and the stacks list poll
it every five seconds. Two tabs on a 40-container host meant a sustained ~16
samples a second. Cached for 4s behind a lock so concurrent callers share one
sweep, the same shape dashboard_service already used for its fleet aggregate.

F11 — Three module dicts assumed exactly one uvicorn worker without saying so
and were lost on restart. The busy set is the lock above. The image update
cache is now mirrored to SQLite, so a restart shows the badges immediately
instead of blanking them for up to an hour, and the already-notified marks come
back with them rather than re-announcing the same updates. The login rate
limiter is a table, so it cannot be cleared by getting the process to restart
and no longer multiplies by the worker count.

The constraint that shaped this: compose_service and update_service are shared
with the agent, which has no database. Neither may import one. So the lock is a
separate service the central app enforces at its own entry points, and update
persistence is an opt-in callback the central app registers in its lifespan —
the agent registers nothing and behaves exactly as before. A test asserts
update_service never imports the database, since that is the kind of thing a
later change breaks silently.

Both new nets were checked by reverting the fix: dropping the lock from
_lifecycle fails six tests, removing the stats cache fails the one that names
the behaviour.

Also wires up cache pruning in the same sweep — without it both the dict and
the table grew one entry per image tag ever run, for the life of the install.

31 new tests (729 total).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dk43rmEeRfYi5wsLDbmfyG
2026-08-31 13:43:39 +02:00

149 lines
4.9 KiB
Python

"""One compose operation per stack at a time.
``docker compose`` does no locking. Two ``update`` calls against the same
project — two open browser tabs, or the auto-update pass landing on a stack
somebody just clicked — both run ``pull`` and then ``up -d``, and race each
other recreating the same containers.
There *was* a busy flag (``compose_service._BUSY``), but it only ever fed the
status column: no lifecycle handler consulted it before acting. This module is
the actual guard, and it lives in the database so it holds across workers and
across a restart.
``compose_service`` keeps its in-process set because it is shared with the
agent, which has no database. The agent is a single process managing one host,
and the central app holds this lock before calling it, so the two do not
conflict.
"""
from __future__ import annotations
import logging
from contextlib import contextmanager
from datetime import datetime, timedelta, timezone
from typing import Optional
from sqlalchemy.exc import IntegrityError
from sqlmodel import Session, delete, select
from models.runtime_state import StackLock
logger = logging.getLogger("stackpilot.stack_lock")
#: Long enough to outlast the slowest legitimate operation (compose commands
#: time out at 600s, a full pull of a large stack can chain several), short
#: enough that a lock orphaned by a killed worker clears itself within an hour.
DEFAULT_TTL = timedelta(minutes=30)
class StackBusy(Exception):
"""The stack is already running an operation."""
def __init__(self, stack_id: str, action: str):
self.stack_id = stack_id
self.action = action
super().__init__(f"Stack '{stack_id}' is busy: {action} in progress")
def _now() -> datetime:
return datetime.now(timezone.utc)
def _aware(value: Optional[datetime]) -> Optional[datetime]:
"""SQLite hands datetimes back naive; compare them as UTC."""
if value is not None and value.tzinfo is None:
return value.replace(tzinfo=timezone.utc)
return value
def acquire(
session: Session,
stack_id: str,
action: str,
owner: str = "",
ttl: timedelta = DEFAULT_TTL,
) -> None:
"""Take the lock for ``stack_id`` or raise :class:`StackBusy`.
An expired lock is taken over — that is the recovery path for a worker that
died mid-deploy, which would otherwise leave the stack unusable.
"""
now = _now()
existing = session.get(StackLock, stack_id)
if existing is not None:
if (_aware(existing.expires_at) or now) > now:
raise StackBusy(stack_id, existing.action)
logger.warning(
"Taking over an expired %s lock on '%s' (held by %r since %s)",
existing.action, stack_id, existing.owner, existing.acquired_at,
)
session.delete(existing)
session.commit()
session.add(
StackLock(
stack_id=stack_id,
action=action,
owner=owner,
acquired_at=now,
expires_at=now + ttl,
)
)
try:
session.commit()
except IntegrityError as exc:
# Another worker inserted between our check and our commit. The primary
# key is what actually makes this safe; the read above is only there to
# give a useful error and to clear stale rows.
session.rollback()
raise StackBusy(stack_id, action) from exc
def release(session: Session, stack_id: str) -> None:
"""Drop the lock. Safe to call when it is not held."""
existing = session.get(StackLock, stack_id)
if existing is not None:
session.delete(existing)
session.commit()
@contextmanager
def hold(session: Session, stack_id: str, action: str, owner: str = ""):
"""Hold the lock for the duration of the block.
Raises :class:`StackBusy` if somebody else has it. Always releases, so a
failed deploy does not leave the stack locked.
"""
acquire(session, stack_id, action, owner)
try:
yield
finally:
try:
release(session, stack_id)
except Exception: # noqa: BLE001 - never mask the original error
logger.exception("Failed to release the lock on '%s'", stack_id)
def active(session: Session) -> dict[str, str]:
"""``{stack_id: action}`` for every lock still in force.
One query for the whole stacks list, rather than a lookup per row.
"""
now = _now()
return {
lock.stack_id: lock.action
for lock in session.exec(select(StackLock)).all()
if (_aware(lock.expires_at) or now) > now
}
def is_busy(session: Session, stack_id: str) -> bool:
lock = session.get(StackLock, stack_id)
return lock is not None and (_aware(lock.expires_at) or _now()) > _now()
def prune_expired(session: Session) -> int:
"""Drop locks that have timed out. Called at startup and by the scheduler."""
result = session.exec(delete(StackLock).where(StackLock.expires_at < _now()))
session.commit()
return result.rowcount or 0