Skip to content
cdkd

State locking

cdkd locks a stack while a command changes it, so that two commands cannot write the same state record at once. The lock is a small object, lock.json, stored next to the stack's record. The command that takes the lock deletes it when it finishes.

# which stacks are locked
cdkd state list --long

# the record and its lock
cdkd state show MyStack --json

# delete a lock whose owner is gone
cdkd force-unlock MyStack --stack-region us-east-1

This page is part of State Management.

Which commands take the lock

A lock covers one stack in one region. Every command that writes the stack's record takes it:

  • cdkd deploy, cdkd destroy and cdkd rollback
  • cdkd import, cdkd export, cdkd orphan and cdkd scrub
  • the cdkd state and cdkd drift forms that write

cdkd diff takes no lock.

The life of a lock

Lifecycle of a stack lockA command that starts work on a stack writes lock.json only if the key is absent. While the lock is held, the command renews it about every 2 minutes, and each renewal moves expiresAt 30 minutes ahead. From there the lock is either released, when the command finished, failed, or took a first Ctrl-C or SIGTERM, or left behind, after a SIGKILL, a crash, a second Ctrl-C or a cancelled CI job. A lock left behind is cleared when the next command takes it over once expiresAt passes, or when cdkd force-unlock deletes it.A command starts work on a stackWrites lock.json only if the key is absentLock heldRenewed about every 2 minutes; eachrenewal moves expiresAt 30 minutes aheadReleasedThe commandfinished, failed,or took a firstCtrl-C or SIGTERMLeft behindSIGKILL, a crash,a second Ctrl-C, acancelled CI jobClearedThe next commandtakes it overonce expiresAtpasses, or cdkdforce-unlockdeletes it nowLifecycle of a stack lockA command that starts work on a stack writes lock.json only if the key is absent. While the lock is held, the command renews it about every 2 minutes, and each renewal moves expiresAt 30 minutes ahead. From there the lock is either released, when the command finished, failed, or took a first Ctrl-C or SIGTERM, or left behind, after a SIGKILL, a crash, a second Ctrl-C or a cancelled CI job. A lock left behind is cleared when the next command takes it over once expiresAt passes, or when cdkd force-unlock deletes it.A command starts work on a stackWrites lock.json only if the key is absentLock heldRenewed about every 2 minutes; eachrenewal moves expiresAt 30 minutes aheadReleasedThe commandfinished, failed,or took a firstCtrl-C or SIGTERMLeft behindSIGKILL, a crash,a second Ctrl-C, acancelled CI jobClearedThe next commandtakes it over onceexpiresAt passes, orcdkd force-unlockdeletes it now

lock.json names the process that holds the lock:

{
  "owner": "goto@macbook:12345",
  "timestamp": 1710835200000,
  "expiresAt": 1710837000000,
  "operation": "deploy"
}
Field Meaning
owner user@hostname:pid of the holding process
timestamp When the lock was taken, in Unix milliseconds
expiresAt The time by which the holder must renew the lock
operation The command holding the lock, such as deploy

How long a lock is held

A command holds the lock for as long as it works on that stack. A run that covers several stacks releases each stack's lock as that stack finishes.

The lock expires 30 minutes after its last renewal, and the holder renews it about every 2 minutes for as long as it runs. The 30 minutes therefore measure how long the holder has been silent. They do not limit how long an operation can run: a deploy that takes two hours stays locked for two hours.

The lock is released when cdkd's own work ends, which can be before the resources have finished coming up. A second deploy can therefore start while resources from the first are still provisioning. cdkd deploy describes what that second deploy meets.

When another command holds the lock

A command that finds a lock whose expiresAt has not passed does not proceed. cdkd deploy and cdkd rollback try three more times, 2 seconds apart, and then fail. Most other commands fail at once.

The error names the lock's owner, its operation and its expiry time. It also prints the cdkd force-unlock command for that exact lock.

If the owner is a deploy that is still running, wait for it to finish.

What an interrupted command leaves behind

Whether a lock is left behind depends on how the command stopped. This table applies to cdkd deploy, cdkd destroy, cdkd state destroy and cdkd rollback.

How the command stops Lock State record
It finishes or fails Released Saved
First Ctrlc or SIGTERM Released Saved
Second Ctrlc or SIGTERM May be left behind As of the last save
SIGKILL, a crash, a machine that sleeps Left behind As of the last save

On the first Ctrlc or SIGTERM, cdkd lets the operations already in progress finish, saves the record and releases the lock. A deploy also writes the rollback journal, the file cdkd rollback reads. You can run the command again immediately.

On the second Ctrlc or SIGTERM, the process exits with code 130 without waiting. The record holds whatever the last save after a resource operation wrote.

An interrupted destroy leaves a record that lists only the resources that still exist. Running the destroy again does not repeat the deletes that completed.

A cancelled CI job

A cancelled CI job can leave the lock behind even though the runner sends SIGTERM first. Runners follow SIGTERM with SIGKILL after a short grace period, about 10 seconds on GitHub Actions, and a long AWS call may not finish in that time. Troubleshooting has workflow patterns that avoid this.

Clear a lock whose owner is gone

A lock that was left behind is no longer renewed. The next command that needs the lock takes it over once expiresAt passes, and warns with the name of the previous owner.

To clear it sooner, find the locked stack and delete the lock:

cdkd state list --long                               # which stacks are locked
cdkd force-unlock MyStack --stack-region us-east-1   # delete the lock now

Caution

cdkd force-unlock deletes the lock whether or not its owner is alive. If you run it against a live deploy, two processes end up writing the same stack. Confirm the owner is gone first.

cdkd force-unlock covers the cases where a lock does not clear by itself.

Last updated: