- Automatic test / admin lock —
lager pythonand the box-mutating admin commands (lager install,lager uninstall,lager update,lager install-wheel) reserve the box for the lifetime of the command. - User lock —
lager boxes lockexplicitly reserves a box until you unlock it.
Automatic test lock
Everylager python <runnable> invocation automatically acquires the box
lock at start and releases it at end. This includes failures, Ctrl+C,
crashes, and signal-killed runs. The lock is released through a finally
block, a signal handler, an atexit net, and (worst case) a server-side
TTL reap.
Which commands auto-lock
Read-only commands (
lager hello, lager boxes list, lager boxes lock /
unlock itself, net-listing paths like lager supply --box X with no
subcommand, status / dry-run paths, etc.) do not acquire the auto-lock.
The v0.12–0.13.3 design used a shared decorator on every command, and
v0.13.4 reverted it. This implementation instead uses the same
TTL + heartbeat + atexit infrastructure as lager python. That avoids the
three corner cases that motivated the revert. See Backward
compatibility below for the full history.
The lock identity is CI-aware so concurrent test runs in CI mutually
exclude correctly. Holder formats:
The
:pid (and @runner / @host) suffix guarantees that two parallel
matrix items in the same workflow run get distinct holder strings.
Collision behavior
Whenlager python tries to acquire a lock that another holder owns:
- On dev: prints an error and exits 1 immediately (no waiting).
- In CI: waits up to
LAGER_LOCK_WAITseconds (default1800, i.e. 30 min), polling every 2s, and only fails if the wait elapses. This lets matrix jobs queue against the same self-hosted box.
lager boxes lock before you ran
lager python, the CLI sees the lock as already-ours. It does not
release the lock on exit, so your explicit reservation survives the test.
TTL & heartbeat
Each test lock is written withttl_seconds: 1800 and refreshed every
60 seconds by a background heartbeat thread inside the CLI. The TTL is
not a cap on test runtime — as long as the heartbeat keeps refreshing
last_heartbeat, the lock stays valid indefinitely.
What the TTL actually bounds is the worst-case stale-lock dwell time
after a CLI crash. If your laptop loses network, or the CI runner is
hard-killed, the box reaps the lock once last_heartbeat + ttl_seconds
falls in the past. Another caller therefore waits at most one TTL.
--detach hands the lock to the box
lager python script.py --detach has no CLI left to hold its lock — the
client is answered and goes away, which is the point. So the box takes the
lock’s lifetime over: it heartbeats while the detached job runs and releases
when the job ends, however it ends. Nothing has to be unlocked by hand.
lager boxes lock reservation that a
detached run merely resumed is never handed over and never released. That
reservation is the whole point of taking one.
Against a box too old to know about the handoff, the CLI keeps the previous
behavior. That is an eternal hold, with the old “release with
lager boxes unlock” message. The CLI arms the lapse TTL only after the box
confirms that it heartbeats. So a newer CLI can never leave a lock that expires
underneath a job that still runs.
Escape hatches
User lock
A user lock is an explicit, persistent reservation you place on a box. Unlike the automatic test lock, user locks never expire — you must manually unlock when you’re done. Use cases:- Reserving a box for an extended debugging session.
- Preventing others from using a box during maintenance.
- Claiming a box when you’re not actively running a command.
lager boxes lock
--box(required) — name of the box to lock.--user— username to lock as (useful when running inside Docker where the user is otherwiseroot).
lager boxes unlock
--box(required) — name of the box to unlock.--force— force unlock even if the box was locked by another user (use this to clear a stalelager boxes lockleft by a teammate).
Management operations skip the lock
The following sub-commands oflager python are management operations
on already-running processes and intentionally skip both lock checks and
auto-acquire:
lager python --kill <ID>lager python --kill-alllager python --reattach <ID>lager python --continue <ID>lager python --console <ID>
--kill it without first having to fight an unrelated user lock.
lager boxes shows lock holders
When boxes are locked, lager boxes shows an extra column:
github lager run 9182 job test on runner-3) rather than printed as raw colon-delimited strings.
CI workflow example
The always-on auto-lock + CI auto-wait combination means a CI matrix job needs no special invocation:...GITHUB_JOB=hardware-tests/<runner>:<pid> differs per item. Each item then
POSTs /lock. Whichever item loses the race waits up to 30 minutes for the
winner to finish, and then retries. No lager boxes lock call needed.
Backward compatibility
lager boxes lockandlager boxes unlockbehave exactly as before. The CLI now sendsholder_type: "user"+ttl_seconds: nullon the wire. A legacy client — an older CLI against the new box server — still gets the same eternal-lock behavior. The server treats a payload with neither field as legacy, and applies the same defaults._check_box_lock(the read-only lock check that already gates every command in resolve_and_validate_box) is unchanged.
How this differs from v0.13.0 – v0.13.3 (removed in v0.13.4)
v0.13.0 added an ephemeral “command-in-progress” lock that fired on every CLI command via a shared decorator, gated by a--force-command flag. v0.13.4 removed it because three corner cases
were unfixable in that design:
--force-command is gone. Collision policy is structured (fail-fast
in dev, queue in CI) and the existing lager boxes lock --force is the
escape hatch when you genuinely need to override.
