Lock stacks during compose runs, cache stats, persist runtime state (0.47.0)
F7 — Nothing stopped two compose operations landing on the same stack. There was a busy flag, but is_busy() was only ever read to colour the status column; no lifecycle handler consulted it before acting. Two tabs, or auto-update picking up a stack somebody had just clicked, both ran pull + up -d against the same project and raced over recreating containers. Lifecycle calls, the two deploy WebSockets and the auto-update pass now take a real lock; a second caller gets 409 (or an error frame and close 4409) and auto-update skips and retries next cycle. The lock is a row rather than a set in one worker's memory, so it holds across workers and across a restart, and it carries an expiry — a worker killed mid-deploy would otherwise strand the stack with no fix short of editing the database. F10 — /api/stacks/stats sampled every running container on every call, one blocking daemon request each, and both the dashboard and the stacks list poll it every five seconds. Two tabs on a 40-container host meant a sustained ~16 samples a second. Cached for 4s behind a lock so concurrent callers share one sweep, the same shape dashboard_service already used for its fleet aggregate. F11 — Three module dicts assumed exactly one uvicorn worker without saying so and were lost on restart. The busy set is the lock above. The image update cache is now mirrored to SQLite, so a restart shows the badges immediately instead of blanking them for up to an hour, and the already-notified marks come back with them rather than re-announcing the same updates. The login rate limiter is a table, so it cannot be cleared by getting the process to restart and no longer multiplies by the worker count. The constraint that shaped this: compose_service and update_service are shared with the agent, which has no database. Neither may import one. So the lock is a separate service the central app enforces at its own entry points, and update persistence is an opt-in callback the central app registers in its lifespan — the agent registers nothing and behaves exactly as before. A test asserts update_service never imports the database, since that is the kind of thing a later change breaks silently. Both new nets were checked by reverting the fix: dropping the lock from _lifecycle fails six tests, removing the stats cache fails the one that names the behaviour. Also wires up cache pruning in the same sweep — without it both the dict and the table grew one entry per image tag ever run, for the life of the install. 31 new tests (729 total). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Dk43rmEeRfYi5wsLDbmfyG
This commit is contained in:
@@ -0,0 +1,148 @@
|
||||
"""One compose operation per stack at a time.
|
||||
|
||||
``docker compose`` does no locking. Two ``update`` calls against the same
|
||||
project — two open browser tabs, or the auto-update pass landing on a stack
|
||||
somebody just clicked — both run ``pull`` and then ``up -d``, and race each
|
||||
other recreating the same containers.
|
||||
|
||||
There *was* a busy flag (``compose_service._BUSY``), but it only ever fed the
|
||||
status column: no lifecycle handler consulted it before acting. This module is
|
||||
the actual guard, and it lives in the database so it holds across workers and
|
||||
across a restart.
|
||||
|
||||
``compose_service`` keeps its in-process set because it is shared with the
|
||||
agent, which has no database. The agent is a single process managing one host,
|
||||
and the central app holds this lock before calling it, so the two do not
|
||||
conflict.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from contextlib import contextmanager
|
||||
from datetime import datetime, timedelta, timezone
|
||||
from typing import Optional
|
||||
|
||||
from sqlalchemy.exc import IntegrityError
|
||||
from sqlmodel import Session, delete, select
|
||||
|
||||
from models.runtime_state import StackLock
|
||||
|
||||
logger = logging.getLogger("stackpilot.stack_lock")
|
||||
|
||||
#: Long enough to outlast the slowest legitimate operation (compose commands
|
||||
#: time out at 600s, a full pull of a large stack can chain several), short
|
||||
#: enough that a lock orphaned by a killed worker clears itself within an hour.
|
||||
DEFAULT_TTL = timedelta(minutes=30)
|
||||
|
||||
|
||||
class StackBusy(Exception):
|
||||
"""The stack is already running an operation."""
|
||||
|
||||
def __init__(self, stack_id: str, action: str):
|
||||
self.stack_id = stack_id
|
||||
self.action = action
|
||||
super().__init__(f"Stack '{stack_id}' is busy: {action} in progress")
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
return datetime.now(timezone.utc)
|
||||
|
||||
|
||||
def _aware(value: Optional[datetime]) -> Optional[datetime]:
|
||||
"""SQLite hands datetimes back naive; compare them as UTC."""
|
||||
if value is not None and value.tzinfo is None:
|
||||
return value.replace(tzinfo=timezone.utc)
|
||||
return value
|
||||
|
||||
|
||||
def acquire(
|
||||
session: Session,
|
||||
stack_id: str,
|
||||
action: str,
|
||||
owner: str = "",
|
||||
ttl: timedelta = DEFAULT_TTL,
|
||||
) -> None:
|
||||
"""Take the lock for ``stack_id`` or raise :class:`StackBusy`.
|
||||
|
||||
An expired lock is taken over — that is the recovery path for a worker that
|
||||
died mid-deploy, which would otherwise leave the stack unusable.
|
||||
"""
|
||||
now = _now()
|
||||
existing = session.get(StackLock, stack_id)
|
||||
if existing is not None:
|
||||
if (_aware(existing.expires_at) or now) > now:
|
||||
raise StackBusy(stack_id, existing.action)
|
||||
logger.warning(
|
||||
"Taking over an expired %s lock on '%s' (held by %r since %s)",
|
||||
existing.action, stack_id, existing.owner, existing.acquired_at,
|
||||
)
|
||||
session.delete(existing)
|
||||
session.commit()
|
||||
|
||||
session.add(
|
||||
StackLock(
|
||||
stack_id=stack_id,
|
||||
action=action,
|
||||
owner=owner,
|
||||
acquired_at=now,
|
||||
expires_at=now + ttl,
|
||||
)
|
||||
)
|
||||
try:
|
||||
session.commit()
|
||||
except IntegrityError as exc:
|
||||
# Another worker inserted between our check and our commit. The primary
|
||||
# key is what actually makes this safe; the read above is only there to
|
||||
# give a useful error and to clear stale rows.
|
||||
session.rollback()
|
||||
raise StackBusy(stack_id, action) from exc
|
||||
|
||||
|
||||
def release(session: Session, stack_id: str) -> None:
|
||||
"""Drop the lock. Safe to call when it is not held."""
|
||||
existing = session.get(StackLock, stack_id)
|
||||
if existing is not None:
|
||||
session.delete(existing)
|
||||
session.commit()
|
||||
|
||||
|
||||
@contextmanager
|
||||
def hold(session: Session, stack_id: str, action: str, owner: str = ""):
|
||||
"""Hold the lock for the duration of the block.
|
||||
|
||||
Raises :class:`StackBusy` if somebody else has it. Always releases, so a
|
||||
failed deploy does not leave the stack locked.
|
||||
"""
|
||||
acquire(session, stack_id, action, owner)
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
try:
|
||||
release(session, stack_id)
|
||||
except Exception: # noqa: BLE001 - never mask the original error
|
||||
logger.exception("Failed to release the lock on '%s'", stack_id)
|
||||
|
||||
|
||||
def active(session: Session) -> dict[str, str]:
|
||||
"""``{stack_id: action}`` for every lock still in force.
|
||||
|
||||
One query for the whole stacks list, rather than a lookup per row.
|
||||
"""
|
||||
now = _now()
|
||||
return {
|
||||
lock.stack_id: lock.action
|
||||
for lock in session.exec(select(StackLock)).all()
|
||||
if (_aware(lock.expires_at) or now) > now
|
||||
}
|
||||
|
||||
|
||||
def is_busy(session: Session, stack_id: str) -> bool:
|
||||
lock = session.get(StackLock, stack_id)
|
||||
return lock is not None and (_aware(lock.expires_at) or _now()) > _now()
|
||||
|
||||
|
||||
def prune_expired(session: Session) -> int:
|
||||
"""Drop locks that have timed out. Called at startup and by the scheduler."""
|
||||
result = session.exec(delete(StackLock).where(StackLock.expires_at < _now()))
|
||||
session.commit()
|
||||
return result.rowcount or 0
|
||||
Reference in New Issue
Block a user