Back up and restore Fungi
Plan recovery for a self-hosted Fungi installation and test backups before using them.
For a self-hosted Fungi installation, back up every state owner and keep the decryption keys somewhere you can reach after losing the host. A copy on the same disk does not protect you from losing that disk.
This guide describes the recovery boundaries. Installation-specific commands, credentials and cutover belong to your operator’s runbook. Restore into a fresh isolated target first. A restore test must not overwrite your running service.
Know what you need to recover
Section titled “Know what you need to recover”| State | What to preserve | Recovery boundary |
|---|---|---|
| Hub account database | PostgreSQL backups, archived WAL and logical dumps | Account/authentication state is separate from Team cells |
| Connection-service database | The database used by your configured connection service | Include it when that service is part of your installation |
| Hub/State fleet | The store’s celld cells, D1 data and stored objects | Preserve database identities, bucket names and configuration with the data |
| Exec App state | Every configured shard’s store bucket | Include all shards, with their own bucket-scoped credentials |
| App Releases and artifacts | Release catalog, captured artifacts, media bytes and source repositories used by the deployment | Hosted Cloudflare resources need their own recovery plan |
| Configuration and credentials | Render inputs, service definitions, signing/auth secrets and resource identities | Keep encrypted copies and a separate key escrow |
| Computer files | A separately arranged export of files you need | The shipped host backup jobs do not back up Computer disks |
| Local CLI work | Project and home .fungi/ state, plus ~/.config/fungi/ or the configured credential location on the client’s machine |
Keep client credentials in encrypted backups. A server backup does not contain them or local Sessions |
An App’s source, published Release, installed permissions and mutable backend state are different records. Saving its source repository alone does not restore its users’ data. Database identities also matter: restoring stored objects under a different celld D1 identity does not recover the same database.
Understand the backup jobs
Section titled “Understand the backup jobs”The PostgreSQL jobs keep physical backups with archived WAL and encrypted logical dumps. WAL is PostgreSQL’s change log. Physical recovery needs the base backup and the required WAL, not just the newest archive file. Logical dumps provide another database recovery path.
The fleet jobs mirror the Hub/State store and each Exec shard bucket. They compare the mirror to its source and retain daily copies. Local mirrors and daily copies are plaintext. Restrict their access. The offsite copy uses encrypted storage.
A mirror made while cells are writing is not a coordinated snapshot of all running cells. It also does not create a transaction across PostgreSQL, fleet stores, App artifacts and external providers. Plan a recovery point across those owners rather than promising that the latest files restore one exact instant.
The configuration job encrypts its archive and verifies that it can decrypt the stored copy. It deliberately excludes the backup decryption keys. Keep those keys in separate escrow, outside the host being protected. Include the keys for physical database backups, encrypted archives and encrypted offsite storage. Losing a required key makes that backup unusable.
Check that backups are usable
Section titled “Check that backups are usable”Check each job’s last successful completion, backup age and offsite copy. An enabled timer or a successful upload is not a restore result. Confirm that all configured databases and Exec buckets are included. A missing shard credential must not silently turn a full backup into a partial one.
Keep an inventory alongside the encrypted backups: source revision, celld pin, database version, schema/migration versions, resource identities, selected backup files and their checksums. Record key fingerprints without recording the keys in ordinary logs or source control.
Rehearse recovery
Section titled “Rehearse recovery”- Obtain the offsite copies and the separately escrowed keys. Verify their checksums and that each encrypted backup can be read.
- Prepare a fresh scratch database cluster, directory and ports. Refuse an existing target or any target that overlaps a live database, backup or key directory.
- Restore the PostgreSQL physical backup and replay its WAL. Separately test the logical dumps in another fresh scratch cluster. Check schemas and per-table row counts.
- Restore the Hub/State and all Exec store copies into scratch. Compare object inventories and contents. Preserve their original identity mapping.
- Extract configuration into scratch and compare its checksums. Do not unpack an archive with absolute paths over the running host.
- Restore or reconnect the separate Release/artifact/source resources from your deployment’s recovery inventory. Include any saved Computer files separately.
- Start an isolated service using the matching runtime and configuration. Verify sign-in, Team access, an Agent Session, App state and artifact reads. Keep email, payments and other external effects isolated during the test.
The maintained offsite restore test checks database row counts, fleet object readback and a configuration file in guarded scratch. Those checks do not prove that the whole product is healthy. The isolated service checks above are a separate part of your recovery rehearsal.
Recover the service
Section titled “Recover the service”Choose a verified restore target and record the accepted recovery point before cutover. Recreate the required provider resources, service configuration and identity mappings. Apply the current migration contract rather than starting an empty schema over restored data.
Verify authentication with the required retained auth secret. Reconcile pending work and billing through their owners before resuming external effects. A restored pending request is not proof that an earlier payment, message or generation never completed.
Switch traffic only after the isolated target is qualified. Restore scheduled backups and monitoring as part of the recovery. Keep a protected prior copy until the new service and its next backup have been checked.
Computer files need their own recovery plan. See Use Computers for their lifetime and Privacy for the service’s data boundaries.