Restores
A restore replaces a service's data with a backup. shed takes a safety backup first, checks the archive before touching anything, and keeps a copy of what it replaces until the new data is verified.
How a restore runs
You start a restore with POST /api/backups/{id}/restore. It creates a restores row (running while it waits in the queue) and answers 202. The work happens in the background, in the same single-file queue as backups. Only successful service backups can be restored. For shed.db, follow the manual steps. The request fails with 409 if shed is busy or the service is fenced, and with 400 for shed.db, unsuccessful, or vanished backups.
- Pre-restore backup. shed takes a backup of the current state with the trigger
pre-restore, using the policy's compression and upload settings. It runs inside the restore's job. If it fails, the restore fails, and nothing has changed. - Check the archive. The archive is opened from the local file, or downloaded from S3 into
<restoreID>.download.partial, and decoded all the way through to prove it decrypts and decompresses. Volume archives are also scanned for unsafe entries: absolute names, names leaving the root, entries beneath a symlink of the archive, and hard links leaving their volume. Any one rejects the whole archive. - Hold the service. shed holds it for the rest of the restore. Deployments in progress are canceled. Deploy, redeploy, start, stop, restart, delete, and volume deletion are rejected, and the API answers 409. Pushes that arrive during a hold get 202 and wait; the newest one deploys once the hold ends (see pushes that can't deploy yet).
- Replace the data. Volumes and redis take the volume path. postgres, mysql, and mongo take the dump path.
- Release the hold. Unless the service is stopped (by you, or by a fence a failure left in place), its active container is started again, recreated if it was removed, and routes are applied. If that fails after the data was restored, the restore is recorded as failed with that error.
Volume restores
The volume path is used for volume backups and for redis dumps. It replaces the service's volumes whose mount path appears in the archive (for redis, the /data volume). Archive paths that are not a volume of the service are logged and left alone, and so are volumes the archive doesn't contain.
One rule drives the design: a volume's data is never removed before a verified copy of it exists.
- Fence. In one transaction, shed inserts a
restore_fencesrow (phaseretaining) and setsservices.stopped, remembering its old value. Nothing starts a stopped service, including boot reconcile, so the fence holds across restarts. - Stop and remove the active container. The deployment stays
active. - Copy each volume to a fresh
shed-vol-<volumeID>-pre-restorevolume, check the copy, then set the phase toreplacing. - Empty each volume (remove and recreate it), extract the archive into it at
/, and check it. - Lift the fence (restore
services.stoppedand delete the row, in one transaction), then remove the pre-restore volumes.
Copies and extractions stream a tar through the Docker archive API between helper containers labeled shed.backup=<restoreID>. shed-restore-<id> mounts the volumes writable for the extraction, and shed-restore-<id>-src (read-only) and shed-restore-<id>-dst mount the source and target of a copy at the service's mount paths.
What "check" means
A check reads the target back and compares it with the tar that was written. Every entry must be present with the same type, the same symlink target, and, for files, the same size and SHA-256. Hard links compare as the file they link to. Extra entries are allowed, since Docker may fill an empty volume with what the image has at the mount path.
When something fails
| Failure | Result |
|---|---|
| Stopping or copying aside fails | Volumes are unchanged. The fence is lifted and the service starts again. |
| Emptying, extracting, or the check fails | The pre-restore copies are put back (empty, copy, check) and the fence is lifted. The error says so. |
| Putting back fails too | The fence and the pre-restore volumes are kept. The service stays stopped and fenced, and the error names the volumes holding the previous data. |
| shed shuts down while replacing | The fence is kept. The next start puts the data back. |
Redis
A redis RDB is written to /data/dump.rdb and, as a hard link, to /data/appendonlydir/appendonly.aof.1.base.rdb with a manifest naming it as the base of a fresh multi-part AOF. A server with appendonly yes loads only the AOF and would otherwise start empty.
Dump restores
For postgres, mysql, and mongo dump backups, shed loads the dump into an isolated copy of the database, so no app can read or write half-restored data. The service's container must be running; otherwise the restore fails with "the service is not running; start it to restore a database dump".
- shed creates
shed-restore-<restoreID>-db, labeledshed.backup=<restoreID>, from the service's container: the same image, entrypoint, command, environment (so the same credentials), volumes, and resource limits. It has network modenone, so it has only loopback, no network alias, and no published ports, and its restart policy isno, so Docker never starts it on its own, not even after a reboot. - The service is fenced in phase
loading, and its containers are stopped and removed. Its private domain and published ports now lead nowhere. - The copy is started on the service's volumes, and shed waits up to 5 minutes for the server to run as PID 1 and accept connections.
- The decoded dump is streamed into the database's own client in the copy with
docker exec. - The copy is stopped, with up to 2 minutes to shut down cleanly, and removed. Only then is the fence lifted, so the service's own container never runs on the volumes at the same time as the copy.
A dump restore replaces databases, it doesn't merge into them. The dumps drop and recreate only what they contain, so shed first drops every other database. Databases created after the backup do not survive.
| Engine | Loaded with | Dropped first | Kept |
|---|---|---|---|
| postgres | psql -v ON_ERROR_STOP=1 | Every database via DROP DATABASE … WITH (FORCE) (postgres 13 or newer). First, new connections are disallowed and other sessions are terminated. | postgres and the templates. Roles and objects created after the backup. |
| mysql | mysql -uroot | Every database, in the loading session, with foreign key checks off. The client stops at the first error. lock_wait_timeout=300 fails a restore blocked by a metadata lock after 5 minutes. | mysql, sys, information_schema, performance_schema. |
| mongo | mongorestore --archive --drop | Every database, with mongosh, which reads credentials from the environment. | admin, config, local. |
Postgres has two extra rules. The dump's DROP ROLE and CREATE ROLE for the connected user always fail, so shed filters them out. Connections are allowed again afterwards, also when the load fails. Without the terminate step, DROP DATABASE would fail while an app is connected and the rest of the dump would mix into the old data.
Failure and cancellation
| What fails | Result |
|---|---|
| Stopping the service, or the copy never gets ready | The data is unchanged. The copy is removed, the fence is lifted, and the service starts again. |
| The load fails, or a shutdown cancels it | The database may hold part of the dump. The copy is stopped and removed, which ends the load, the fence row is deleted, and the service stays stopped. The error tells you to restore a backup again, or to start the service to keep the data as it is. |
| Removing the copy fails | The fence stays, so nothing can start the service while the copy may still run. The next boot removes the copy first. |
Restore fences
A restore_fences row means a restore is changing, or failed while changing, the service's data. While it exists, shed refuses everything that could start or replace the service, and the row is authoritative: it survives restarts.
| Blocked while fenced | Result |
|---|---|
| Deploy and redeploy (manual deploys, rollbacks) | 409 Conflict |
| Start and restart | 409 Conflict |
| Container recreation (boot reconcile, releasing a hold) | Refused |
| A new restore into the service | 409 Conflict |
| A GitHub push for the service | Stored and answered 202. It deploys once you clear the fence, unless a newer deployment was made first. |
Stopping and deleting the service still work. The service API exposes the fence as restoreFence, with restoreId, phase (retaining, replacing, or loading), and createdAt.
Clearing a fence
The fence goes away when its restore finishes, when boot recovery resolves it, or when you clear it with POST /api/services/{id}/restore-fence/clear (the dashboard's confirmed Keep current data action). Clearing deletes the row and nothing else: the service stays stopped with the data it has now, and pre-restore volumes are left for you to inspect or remove by hand. The next restore of the service overwrites them. Clearing is refused with 409 while the service is held, so a restore in progress can't be cleared, and clearing a service without a fence does nothing.
Crash recovery
On boot, before the deployer reconciles, shed settles whatever a crash or shutdown interrupted. Backup rows first: queued and running backups and running restores become failed ("interrupted by restart"), and uploading backups are settled by what survives. Leftover .partial files and shed.backup helper containers are removed. Then each fence is resolved by its phase.
| Fence phase | On the next boot | Service afterwards |
|---|---|---|
retaining | Volumes are unchanged, so the fence is lifted and the pre-restore volumes are removed. | Starts as before the restore. |
replacing | The pre-restore copies are put back and checked first, then the fence is lifted. If that fails, the error is logged and the fence stays. | Starts on success. Stays stopped and fenced on failure, and the next boot tries again. |
loading | The isolated copy is removed, because the load may still be running in it, and the active container is stopped. Then the fence row is deleted. If the copy cannot be removed, the error is logged and the fence stays. | Stays stopped. Its data may hold part of the dump. |
One more case: a fenced service that is no longer stopped (only a shed that did not enforce fences could have started it). Its data may have changed since, so its active container is stopped and stopped is set again, but its fence, data, and pre-restore volumes are kept and the error is logged. Recovery never drops a fence because the service looks started.
The same sweep settles uploading backups: with the local file present, the backup is succeeded and local, with the interruption as its remote_error. Without it, the backup is remote-only if the object at its recorded remote_key can be read, and failed otherwise. A succeeded backup therefore always has a local file or a remote object. On shutdown, the running job is canceled, and it and the queued ones are marked failed ("interrupted by shutdown").
Recovering shed.db
shed.db can't be restored from the dashboard, because shed is running on it. Recovery of the metadata database is manual: stop shed, replace the database, then start shed. These steps assume the server still has its Docker images and volumes. They do not recover a lost server's workloads or data; see the new-server section below. You need your age secret key if backups are encrypted, and the original shed.key: shed.db holds secrets encrypted with it, so put it back next to the config file before starting shed, or shed refuses to start. The GitHub App credentials come back with shed.db, but if the new server has a different server.url, update the App's webhook and callback URLs to match. See the encryption section about storing that key.
# Run in bash; stop if decryption or decompression fails. set -euo pipefail systemctl stop shed # fetch the newest shed.db archive from S3 (any S3 client works)... aws s3 ls s3://<bucket>/<prefix>/system/ --endpoint-url <endpoint> aws s3 cp s3://<bucket>/<prefix>/system/<id>.db.zst.age . --endpoint-url <endpoint> # ...or from the old server's disk: <data.dir>/backups/system/ # decrypt (drop age if encryption was off) and decompress age -d -i key.txt <id>.db.zst.age | zstd -d -o shed.db # Keep the existing database and its WAL together before replacing them. recovery_copy=$(mktemp -d /var/lib/shed/before-recovery.XXXXXX) for file in /var/lib/shed/shed.db /var/lib/shed/shed.db-wal /var/lib/shed/shed.db-shm; do if [ -e "$file" ]; then cp -a "$file" "$recovery_copy/"; fi done mv shed.db /var/lib/shed/shed.db rm -f /var/lib/shed/shed.db-wal /var/lib/shed/shed.db-shm systemctl start shed
shed.db runs in WAL mode, so shed.db-wal and shed.db-shm sit beside it. They belong to the database you are replacing, so delete them as shown. /var/lib/shed is the default data.dir; use yours if you changed it.
Recovering on a new server
A shed.db backup restores records: projects, services, variables, domains, deployment history, backup locations, pending pushes, and GitHub App credentials. It does not contain Docker images, containers, or volume data. Keep the matching TOML configuration, archive files or S3 access, shed.key, and encryption key separately.
For a host-level recovery, restore the Docker images and named volumes from your host backup before starting shed. Volume names must match the IDs in the restored database (shed-vol-<volumeID>). A missing image or declared volume blocks container recreation instead of creating an empty replacement. Boot logs identify the failure, and the affected service appears crashed. Restored pending pushes can still trigger new deployments, so keep the server isolated from production traffic until recovery is checked.
If the Docker state was lost and only shed's archives remain, rebuilding services is a separate manual recovery operation. Keep dependent apps stopped and auto-deploy disabled while recovering databases. A new deployment builds the current branch head or pulls the configured image tag; it does not recover the previous image or its data, and missing volumes start empty. Database dump restores require a running, compatible database container. Restore data from a reachable backup after that container exists, verify it, then bring dependent apps and traffic back. A full fresh-host recovery through this sequence has not been integration-tested.
Local archive files must be copied into the matching <data.dir>/backups/<serviceID>/ paths recorded by shed; restoring their database rows does not copy the files. S3 archives need their recorded destination credentials, and encrypted archives need the original age key. If the address changed, update server.url, DNS, and the GitHub App's webhook and callback URLs.