Deploying

Deployments

A deployment is one attempt to get a commit or an image running. It builds, starts a new container beside the old one, checks it is healthy, then moves traffic over. If anything fails, the old container keeps serving.

The pipeline

Each service has one worker, so its deployments run one at a time. A new deployment cancels any older one that hasn't finished: a queued one is canceled outright, a running one is interrupted with the reason superseded by a newer deployment. Deployments of different services run in parallel, though only one build runs at a time across the whole instance.

waitingWait for CIpoll 10s ≤ 60mbuildingBuildclone · builddeployingStartno alias yetdeployingHealth checkprobe ≤ 120sdeployingSwitchalias · routesactiveActiveold → removedSkippedCI failedred CIFailederror recorded · candidate removed · previous container keeps servingA CI wait that outlasts 60 minutes fails the deployment.A newer deployment, the Cancel button, or stopping the service cancels it from any step.Services that don't wait for CI skip the first step.
Solid arrows are the happy path. Dashed arrows are failure exits.
  1. Wait for CI. Only for apps with waitForCi and a commit. See Waiting for CI.
  2. Build. Repo apps are cloned at the commit and built into shed/<serviceID>:<deploymentID>. Image apps and databases pull their image and record its local image ID. See Builds.
  3. Start. Variables are resolved and the container is created without the service's network alias.
  4. Health check. Up to 120 seconds for a TCP connect or a 2xx/3xx on the health check path.
  5. Switch. The alias moves, a second probe runs, routes are applied, the deployment turns active, and the old container is removed.

Deployment statuses

A deployment's status only moves forward. Four statuses mean it is still in the pipeline; the rest are final, and a final deployment never changes status again, except that active becomes removed when the next one goes live.

queuedwaitingbuildingdeployingactiveremovedno CI waitsupersededcanceledcancelskippedCI redfailedCanceled can also happen from waiting, building, and deploying. Failed from waiting is the 60-minute CI timeout.
Dashed transitions end the deployment without it going live.
StatusMeaning
queuedCreated and waiting for its service's worker.
waitingHolding for the commit's CI results.
buildingCloning and building, or pulling the image.
deployingStarting, health checking, or switching traffic.
activeLive. A service has at most one, enforced by the database.
removedWas active; a newer deployment replaced it.
failedEnded with an error, recorded on the deployment.
canceledStopped by a newer deployment, the Cancel button, stopping the service, or deleting it.
skippedCI failed, so the deployment never built.
crashedReserved. shed doesn't record it today; a container that dies shows up as a crashed service.

Zero-downtime switchover

Replacing a container without dropping requests is mostly about ordering. shed never routes to a container it hasn't verified twice, and never retires the old one before the new one is in the routing table.

Services that can share nothing, those with volumes or a public port, can't overlap: two Postgres containers can't open one data directory, and two containers can't publish one host port. For them shed stops the old container first and accepts a short outage. Everything else overlaps.

Step through a switchover
network shed-<projectID>Caddypublic routesOld containerrunning · has aliasNew containernot created yetOther servicesweb:3000

The build or pull finished. The old container is serving all traffic, publicly through Caddy and privately through its network alias.

If any step fails, the candidate is removed and the deployment is marked failed. When the old container had been stopped first, shed starts it again, but only after confirming that every other container of the service is gone, since two of them could share storage. If that can't be confirmed the old one stays stopped, and the error says so.

Route updates are serialized through activation. If Caddy rejects the new config, the candidate fails and the previous deployment keeps its routes. Once routing has succeeded, activation finishes even if you cancel at that moment, so a routed candidate is never removed mid-switch.

Reading a build log

Every deployment writes a build log to <data>/logs/<deploymentID>.log, whatever its outcome. Each stage prints a line starting with ==> , followed by detail lines. The headings are stable, which makes them easy to grep.

==> Waiting for CI on a1b2c3d
==> CI passed
Only for apps with waitForCi and a commit. Status is waiting. Errors reaching GitHub print as Checking CI status: … and polling continues.
==> Building acme/web@a1b2c3d
Status becomes building. Image apps and databases print ==> Pulling <image> instead, and a redeploy prints ==> Reusing image <ref> and skips ahead.
==> Cloning
[git init, fetch --depth 1, checkout]
==> Building with Dockerfile
[docker buildx build output]
Printed by the builder. The second heading is ==> Building with Railpack when there is no Dockerfile. Tool output follows each heading, with variable values masked.
==> Detected port 3000 from the image
Only when the service's port is 0. The lowest exposed TCP port is saved on the service.
==> Stopping previous deployment x5k2m7q1d9ab
Volumes and published ports cannot be shared, so the previous container stops first
Only for services with volumes or a public port. Status is already deploying.
==> Starting container
Name: shed-p4n8…-r7c1…
Image: shed/p4n8…:r7c1…
Network: shed-k2…, private address web:3000 once healthy
Volume: shed-vol-t9… → /data
Environment: 9 variables
Started container 3f9a1c2b7d10
What the container runs with. Variables are counted, never listed. A start command and published ports add lines of their own.
==> Waiting for port 3000 to become healthy
Probing TCP connect to 172.18.0.4:3000 (timeout 2m0s)
Not ready yet: dial tcp 172.18.0.4:3000: connect: connection refused
Listening on :3000
Healthy after 2.3s
With a health check path the heading is Waiting for /healthz to become healthy and the probe is a GET. The Not ready yet line repeats every 5 seconds. The container's own output is copied in (here, Listening on :3000), and the heading is the only line shed adds a space to if it starts with ==>.
==> Watching the container start
No port, so no health check; watching for 3s
Still running after 3s
The same stage for a service without a port.
==> Switching traffic
Still healthy at 172.18.0.4:3000
Private host web resolves to the new container
Routing web.example.com → port 3000
Removing previous deployment x5k2m7q1d9ab (container 0d8e1f2a3b4c)
The alias, the second probe, then routes. Everything is written before the deployment turns active, because log followers stop then.
==> Deployment failed: health check timed out after 2m0s (last error: …)
The last line of a deployment that ended without going live. The word after Deployment is the status: failed, skipped, or canceled.

Output from the container itself is copied in from the moment it starts until the health check ends, without Docker's timestamps, so its boot messages and last words are right there. A container line that begins with ==> gets a leading space, so it can't be mistaken for a heading. Logs are capped at deployments.log_max_mb (10 by default); past the cap, output is dropped and the log ends with ==> Log truncated. Values of the service's variables are masked.

Paste a log into the reader below to see which stage it reached. A deployment that failed is shown stopped at the last stage that printed a heading.

Place a build log on the pipeline
  1. Wait for CINot used
  2. Build or pullDone
  3. Start containerDone
  4. Health checkStopped here
  5. Switch trafficNot reached

Ended failed: health check timed out after 2m0s

  1. ==> Building acme/web@a1b2c3dBuild or pull
  2. ==> CloningBuild or pull
  3. ==> Building with DockerfileBuild or pull
  4. ==> Starting containerStart container
  5. ==> Waiting for port 3000 to become healthyHealth check
  6. ==> Deployment failed: health check timed out after 2m0s

Redeploys and rollbacks

Redeploying an old deployment is how rollbacks work. It creates a new deployment with the trigger redeploy, the same commit details, and the same recorded image, then runs it from the start stage: no CI wait and no build. The old deployment isn't revived, so your history keeps reading in order.

A redeploy reuses the image but not the configuration. It starts with the service's current variables, port, and limits. Rolling back code doesn't roll back settings.

The image must be named immutably. Built images (shed/<serviceID>:<deploymentID>) and pulled images pinned by ID or digest qualify. Older records that stored only a mutable tag such as postgres:18-alpine are refused, since the tag may have moved; deploy those afresh.

The image must also still be on the server. shed keeps only the five newest built images per service, plus the active one, while history keeps 50 deployments, so most old deployments can't be redeployed. shed checks before queueing: a redeploy of a missing image is refused with a conflict, nothing is queued, and the running container is left alone. The pipeline checks again before it stops anything, in case the image was removed meanwhile. To get an old commit back, deploy it again so it is rebuilt.

History and retention

After each deployment ends, shed trims the service's history: finished deployments (failed, removed, canceled, skipped) beyond the newest deployments.keep are deleted, along with their build logs. The default is 50; 0 keeps everything. Active, crashed, and in-progress deployments are never deleted, however old.

The dashboard and GET /api/services/{id}/deployments list the newest 50.

Stop, start, restart

Stop, start, and restart act on the active deployment's container. They aren't deployments and write no build log.

ActionWhat happens
stopCancels any deployment in progress (canceled, "service stopped"), sets the stopped flag, stops the container with a 30-second grace period but doesn't remove it, and drops the service's routes. The deployment stays active.
startClears the flag, starts the active deployment's container (recreating it as deployed if it's gone, see below), and restores routes. Refused if nothing was ever deployed, or if recreating its container would require a missing image or volume.
restartRestarts the active container, then reapplies routes, because the container may come back at a different address. Refused if the service is stopped or has nothing deployed.

A stopped service stays stopped across a shed restart or reboot. You can still deploy it: when the new deployment goes live, it clears the flag. All three are refused with a conflict while the service is fenced by a restore or held for a backup.

Reconcile on boot

Containers use the unless-stopped restart policy, so Docker brings them back after a reboot by itself. shed still checks its own books on every start, because a crash can leave the database and Docker disagreeing. Before it accepts requests, shed does four things:

  1. Marks every deployment still queued, waiting, building, or deploying as failed with the error interrupted by restart, and removes its containers. Pipelines don't resume; deploy again.
  2. Ensures the active deployment of every service that is not stopped and not fenced has a running container, recreating a missing one as deployed (see below).
  3. Keeps one deployment per service, the newest active one, and removes the service's other containers. That is how a predecessor left behind by a crash mid-switchover disappears.
  4. Applies the proxy routes.

A graceful shutdown works the same way from the other side: running pipelines are interrupted and recorded with the error interrupted by shutdown.

Recreating a missing container

When the active deployment goes live, shed records the configuration its container was started with: the image, port, command, resolved variables, CPU and memory limits, public port, and volumes. A container that has gone missing, at boot, on start, or after a restore, is recreated from that record, not from settings saved since. A start command, variable, or limit you saved without deploying takes effect only with the next deployment, so a typo in it can't break a service that was running.

  • Deleted volumes stay deleted. A volume deleted since the deployment is not mounted, and its data is removed rather than recreated empty. A volume added since is mounted from the next deployment.
  • Deployed volumes must still exist. If a volume is missing from Docker, shed refuses to recreate the container and reports the service as crashed. Start returns 409. Restore the volume from a backup, or deploy again to explicitly start with an empty volume. Volumes deleted in shed are not required.
  • The image must be on the server. shed never builds or pulls while recreating. If the image is gone, nothing is created, not even the volumes, and the error is logged. The service shows as crashed; start fails with a conflict. Deploy it again.
  • Older deployments. A deployment that went live before shed kept this record is recreated from the service's current settings, as before.

Moving to a new server

Recovery on boot repairs a server that still has its Docker images and volumes. It doesn't move shed to a new one: shed.db holds the records, not the images or the data. With only shed.db restored on a fresh server, missing images and volumes block recovery of active containers. Affected services show as crashed; recovery does not create empty replacement volumes. Pending pushes and new deployments still follow the normal deployment pipeline. That is deliberate: a database started on an empty volume would look healthy and hide the loss. Recover the Docker state separately, or rebuild services and restore data from their backups. See new-server recovery for prerequisites and limits.