The deploy pipeline
This page follows one cdkd deploy through its five stages and names the
module that owns each one. Read Architecture first for the
overview and the diagram of the whole run.
cdkd deploy MyStack
| Stage | What happens | Layer |
|---|---|---|
| 1 | Resolve the configuration | CLI |
| 2 | Synthesize the CDK app | Synthesis |
| 3 | Schedule assets and stacks | Deployment |
| 4 | Publish the assets | Assets |
| 5 | Deploy each stack | State, Analysis, Deployment, Provisioning |
Stage 1: resolve the configuration
The CLI needs two values before anything else can run: the command that runs the CDK app, and the S3 bucket that holds cdkd's state. Each value is looked up in several places, and the first place that has it wins.
| Value | Looked up in, in order |
|---|---|
| App command | --app, the CDKD_APP environment variable, cdk.json |
| State bucket | --state-bucket, the environment, cdk.json, then the default cdkd-state-{accountId} |
Synthesis does not need the state bucket, so cdkd resolves the bucket while synthesis is running. The stages on this page are the logical order.
Stage 2: synthesize the CDK app
Synthesis turns the CDK app into CloudFormation templates. cdkd does not use
the CDK CLI or the CDK toolkit libraries for it. The app itself, through
aws-cdk-lib, generates the templates. cdkd runs the app, reads what the app
wrote, and answers the app's questions about the AWS account.
Running the app
AppExecutor (src/synthesis/app-executor.ts) starts the app as a child
process. It passes the app what it needs through four environment variables:
| Variable | Value |
|---|---|
CDK_OUTDIR |
The output directory, cdk.out by default |
CDK_CONTEXT_JSON |
The merged context, serialized |
CDK_DEFAULT_REGION |
The AWS region |
CDK_DEFAULT_ACCOUNT |
The AWS account id |
The context in CDK_CONTEXT_JSON is merged from five sources. When two
sources set the same key, the source lower in this table wins.
| Order | Source |
|---|---|
| 1 | CDK defaults (aws:cdk:enable-path-metadata, aws:cdk:enable-asset-metadata, aws:cdk:version-reporting, aws:cdk:bundling-stacks) |
| 2 | ~/.cdk.json, the context field |
| 3 | cdk.json, the context field |
| 4 | cdk.context.json, the cached lookup results, reloaded on each loop iteration |
| 5 | -c key=value on the command line |
When --app names an existing directory, cdkd treats the directory as a cloud
assembly that is already synthesized and does not run the app.
Reading the cloud assembly
The app writes its output, the cloud assembly, to cdk.out/. AssemblyReader
(src/synthesis/assembly-reader.ts) parses cdk.out/manifest.json and finds
four things for each stack:
- the template,
{StackName}.template.json, - the asset manifest,
{StackName}.assets.json, - the stacks this stack depends on,
- the CDK annotations (
Annotations.addError,addWarning,addInfo).
synth and deploy print the warnings. They refuse to continue when a
selected stack carries an error annotation.
Context lookups run the app again
A construct such as Vpc.fromLookup() needs a value from the AWS account, and
the app cannot fetch it. So the app records the missing key in the manifest
and exits. Synthesizer (src/synthesis/synthesizer.ts) looks the value up
through the AWS SDK, saves the answer to cdk.context.json, and runs the app
again. The loop ends when the manifest reports nothing missing.
The next diagram shows the same loop by the module that performs each step:
Every CDK context provider type is supported. The implementations are in
src/synthesis/context-providers/.
Stage 3: schedule assets and stacks
A stack cannot deploy before its assets exist, and it cannot deploy before the
stacks it depends on. WorkGraph (src/deployment/work-graph.ts) handles both
rules with one graph. Every asset and every stack is a node, and a node becomes
ready when all of its dependencies have completed.
There are three node types, and each has its own concurrency limit:
| Node type | Work | Limit | Flag |
|---|---|---|---|
asset-build |
Build a Docker image | 4 | --image-build-concurrency |
asset-publish |
Upload a file to S3 or push an image to ECR | 8 | --asset-publish-concurrency |
stack |
Deploy one stack | 4 | --stack-concurrency |
A file asset has one node, asset-publish. A Docker asset has two, because the
image is built before it is pushed. A stack waits for all of its assets and for
the stacks CDK says it depends on. When a node fails, the nodes downstream of
it are skipped.
Stage 4: publish the assets
Each asset node is run by a publisher in src/assets/:
FileAssetPublisherchecks for the object withHeadObjectand skips the upload when the object exists.DockerAssetPublisherbuilds the image and pushes it.
A third class, AssetPublisher, is the orchestrator behind the standalone
cdkd publish-assets command. deploy does not use it, because deploy
drives the individual asset nodes through the work graph.
Where assets are stored
Assets go to one of two sets of locations. cdkd decides per region, from
whether cdkd bootstrap has written its bootstrap marker for that region.
| cdkd-assets mode | Legacy mode | |
|---|---|---|
| Used when | The region has the marker | The region has no marker |
| S3 bucket | cdkd-assets-${AccountId}-${Region} |
cdk-hnb659fds-assets-${AccountId}-${Region} |
| ECR repository | cdkd-container-assets-${AccountId}-${Region} |
cdk-hnb659fds-container-assets-${AccountId}-${Region} |
A synthesized template refers to the CDK bootstrap locations, which are the legacy ones. In cdkd-assets mode cdkd rewrites those references to the cdkd locations. The rewrite rules are in Architecture internals.
Stage 5: deploy each stack
DeployEngine deploys one stack. Its code is in
src/deployment/deploy-engine.ts and the mixins in deploy-engine/. For each
stack it does nine things in order:
- Acquire the stack lock.
- Load the stack's state from S3.
- Parse the template and drop the resources whose
Conditionis false. - Build the dependency graph.
- Diff the template against the state. Each resource becomes
CREATE,UPDATE,DELETEorNO_CHANGE. - Print the plan. With
--dry-run, stop here. - Run the creates and updates in dependency order, then the deletes in reverse dependency order.
- Resolve the template's
Outputs. - Save the state and release the lock.
State is also saved after each resource completes. A crash in the middle of a deploy therefore leaves a state file that matches what is in AWS.
Which resources depend on which
The dependency graph (src/analyzer/dag-builder.ts) has an edge wherever one
resource names another in the template. Three things create an edge:
- a
DependsOnattribute, - a
Ref, - an
Fn::GetAtt.
The graph also has two kinds of edge that the template does not state:
- An IAM policy attached to a custom resource's handler role gets an edge to the custom resource. Without it, the handler could be invoked before the policy exists.
- A Lambda function in a VPC gets edges from its subnets and security groups. On delete the function then goes first, and its network interfaces have time to detach.
The order resources run in
A resource starts as soon as all of its own dependencies have completed. It
does not wait for unrelated resources at the same depth of the graph. This
event-driven dispatch is in src/deployment/dag-executor.ts.
--concurrency (default 10) caps the number of operations in flight. When more
resources are ready than the cap allows, the one with the most transitive
dependents starts first. A slow chain, such as an Elastic IP followed by a NAT
gateway, therefore starts early.
The log line DAG: <n> levels reports the depth of the graph. Levels are not
used to schedule anything.
When a resource fails
When one operation fails, the resources downstream of the failed one are
skipped. Operations that are already in flight finish, and nothing new starts.
The engine then rolls back what this run changed, unless the deploy ran with
--no-rollback. Rollback describes what a rollback undoes.
How the diff decides what changed
DiffCalculator (src/analyzer/diff-calculator.ts) compares each resource's
properties in state with the same resource in the template. Three rules matter
when you change anything near it:
- State holds resolved values and the template holds intrinsic functions such
as
Ref. So the template side is resolved against the current state before the two are compared. - Only the keys the template declares are compared. An extra key in state is a value AWS added, so it does not count as a change.
- Some properties cannot be updated in place. A change to one of them makes
the change a replacement: the new resource is created and the old one is
deleted. Replacing a stateful resource requires
--force-stateful-recreation.
A replacement gives the resource a new physical id. The resources that refer to
it must receive the new id, so the diff promotes them from NO_CHANGE to
UPDATE. The promotion rules, and the separate diff of the Outputs section,
are in
Architecture internals.
What an update sends to AWS
The two kinds of provider update a resource differently. An SDK provider's
update() calls the service's own update APIs. On the Cloud Control route,
cdkd generates a JSON Patch (RFC 6902) from the old and the new properties and
calls UpdateResource with it:
Providers and intrinsic functions describes the two kinds of provider.
Related
- Architecture: the layers and the overview of a deploy
- Destroy, state and locking
- Providers and intrinsic functions
- Architecture internals: the detailed rules behind the diff and asset storage
- Rollback