Manage instance backups from Server settings: list, back up now, download, delete, schedule and retention #242

Open
opened 2026-09-15 21:40:04 +00:00 by tiagoagueda · 1 comment
Owner

Administrators should be able to manage instance backups from Server settings: see which archives exist, take one now, download it, delete old ones, and set a schedule with retention. Today all of that is a shell command and a cron line on the host.

What exists today

The hard part is already built and tested (#32):

  • core/backup.py:
    • write_backup writes a manifest (Postulo version, engine, counts), the database (SQLite's backup API, or pg_dump) and the media directory, then verifies the archive;
    • verify_backup re-checks an archive against its checksum;
    • restore_backup refuses an archive from the other engine and refuses a non-empty instance without force.
  • manage.py backup [target] and manage.py restore <archive> [--force].
  • POSTULO_BACKUP_DIR (config/settings/base.py:452): data/backups by default, /app/data/backups in the container.
  • Server settings → Overview already shows the newest archive (core/server_views.py:75,131), but links nowhere.
  • The wiki (Backups and your data) tells operators to add a host cron line and to encrypt archives themselves in transit.

What is missing is any of it in the interface.

Proposal: a Backups section in Server settings

A new SettingsSection (core/server_sections.py), staff only, between Logs and Defaults.

1. The list

  • Every archive in POSTULO_BACKUP_DIR, newest first: when, size, Postulo version, engine, record counts from the manifest, and whether it verifies.
  • Verification is on demand, or cached by file mtime; never re-hashed on every page load.
  • Free disk space on that volume, and a warning when the next backup is unlikely to fit (estimated from the newest archive).
  • An archive that fails verification says so in words, not only in colour.

2. Back up now

  • Runs as a background task, never inside the request. On SQLite a request holds the write lock for its whole run (#220), and an archive with media can take minutes.
  • The page shows working… and polls with htmx; the result arrives in a live region.
  • One backup at a time: a second press while one runs says so.

3. Download

  • The archive holds everyone's data, unencrypted. Downloading is the most sensitive action on the instance, so it needs:
    • a recent re-authentication (allauth's reauth flow, as Delete my account uses);
    • a log line naming who downloaded which archive;
    • a streamed FileResponse with Content-Disposition: attachment.
  • Only files matching the backup name pattern inside POSTULO_BACKUP_DIR, addressed by name: never a path from the request, never a symlink out of the directory.
  • Optional passphrase encryption of the download (the passphrase is never stored). The wiki currently says Postulo leaves encryption to restic, borg or age. That was reasonable for a file on the host; a button that sends everyone's data to a browser is a different exposure. Decide here.

4. Delete

A confirmation page naming the archive and its date, re-authentication, and a log line. The newest verified archive cannot be deleted while it is the only one.

5. Schedule and retention

  • Off, daily or weekly, at a chosen hour in the instance's time zone; keep the newest N (default 7) and delete older ones only after a new archive has verified.
  • Run by the scheduler, so it depends on the scheduler being reliable (#221).
  • Last scheduled run, its outcome, and the next run shown on the page. A failed or missed run shows on Overview too.
  • Metrics: postulo_backup_last_success_timestamp_seconds and postulo_backup_failures_total, documented on Health, metrics and logs.
  • Settings live on SiteSettings, and the host cron line in the wiki stays valid for people who prefer it.

6. Restore: decide the scope

Restoring from the web is the dangerous half. restore_backup overwrites the database and media under a running instance, with gunicorn workers and the scheduler still connected, and #234 asks restore to refuse exactly that. Two ways forward:

  • (a) Recommended for this issue:

    • upload an archive (size-limited, into POSTULO_BACKUP_DIR) or pick one from the list;
    • verify it and show its manifest;
    • show the exact stop-restore-start commands for this installation (Compose or bare), with the archive path filled in.

    The restore itself stays on the command line.

  • (b) Later, as its own issue: a real maintenance mode (every request answers 503 with a page, the scheduler pauses), then a restore task that runs under it and requires re-authentication plus typing the instance's name.

Depends on

  • #220: background tasks, so nothing runs in the request.
  • #221: a scheduler that doesn't double-run or die, for scheduled backups.
  • #234: the archive should include the plugins record and a field-key fingerprint, so an archive shown here is actually restorable; restore safety.
  • #219: pg_dump in the image, or backups on PostgreSQL fail from the web exactly as from the shell.

Checks

  • Security tests:
    • non-staff get 404 on every Backups URL;
    • download and delete require recent re-authentication;
    • names with .., /, absolute paths or symlinks are refused;
    • an upload over the limit or not a Postulo archive is refused before it is kept.
  • Page coverage: add the page to the browser walk. The download and task-status URLs go in EXCUSED with reasons.
  • Accessibility: the list is a real table, progress is announced in a live region, and destructive actions use the shared danger styles (#227).
  • Translations: every new string, in the working catalogues.
  • Wiki: Backups and your data gains a section for the page, and Configuration documents the schedule and retention settings.
Administrators should be able to manage instance backups from *Server settings*: see which archives exist, take one now, download it, delete old ones, and set a schedule with retention. Today all of that is a shell command and a cron line on the host. ## What exists today The hard part is already built and tested (#32): - `core/backup.py`: - `write_backup` writes a manifest (Postulo version, engine, counts), the database (SQLite's backup API, or `pg_dump`) and the media directory, then verifies the archive; - `verify_backup` re-checks an archive against its checksum; - `restore_backup` refuses an archive from the other engine and refuses a non-empty instance without `force`. - `manage.py backup [target]` and `manage.py restore <archive> [--force]`. - `POSTULO_BACKUP_DIR` (`config/settings/base.py:452`): `data/backups` by default, `/app/data/backups` in the container. - *Server settings → Overview* already shows the newest archive (`core/server_views.py:75,131`), but links nowhere. - The wiki (*Backups and your data*) tells operators to add a host cron line and to encrypt archives themselves in transit. What is missing is any of it in the interface. ## Proposal: a *Backups* section in Server settings A new `SettingsSection` (`core/server_sections.py`), staff only, between *Logs* and *Defaults*. ### 1. The list - Every archive in `POSTULO_BACKUP_DIR`, newest first: when, size, Postulo version, engine, record counts from the manifest, and whether it verifies. - Verification is on demand, or cached by file mtime; never re-hashed on every page load. - Free disk space on that volume, and a warning when the next backup is unlikely to fit (estimated from the newest archive). - An archive that fails verification says so in words, not only in colour. ### 2. *Back up now* - Runs as a background task, never inside the request. On SQLite a request holds the write lock for its whole run (#220), and an archive with media can take minutes. - The page shows *working…* and polls with htmx; the result arrives in a live region. - One backup at a time: a second press while one runs says so. ### 3. Download - The archive holds **everyone's** data, unencrypted. Downloading is the most sensitive action on the instance, so it needs: - a recent re-authentication (allauth's reauth flow, as *Delete my account* uses); - a log line naming who downloaded which archive; - a streamed `FileResponse` with `Content-Disposition: attachment`. - Only files matching the backup name pattern inside `POSTULO_BACKUP_DIR`, addressed by name: never a path from the request, never a symlink out of the directory. - Optional passphrase encryption of the download (the passphrase is never stored). The wiki currently says Postulo leaves encryption to restic, borg or age. That was reasonable for a file on the host; a button that sends everyone's data to a browser is a different exposure. Decide here. ### 4. Delete A confirmation page naming the archive and its date, re-authentication, and a log line. The newest verified archive cannot be deleted while it is the only one. ### 5. Schedule and retention - Off, daily or weekly, at a chosen hour in the instance's time zone; keep the newest *N* (default 7) and delete older ones only after a new archive has verified. - Run by the scheduler, so it depends on the scheduler being reliable (#221). - Last scheduled run, its outcome, and the next run shown on the page. A failed or missed run shows on *Overview* too. - Metrics: `postulo_backup_last_success_timestamp_seconds` and `postulo_backup_failures_total`, documented on *Health, metrics and logs*. - Settings live on `SiteSettings`, and the host cron line in the wiki stays valid for people who prefer it. ### 6. Restore: decide the scope Restoring from the web is the dangerous half. `restore_backup` overwrites the database and media under a running instance, with gunicorn workers and the scheduler still connected, and #234 asks `restore` to refuse exactly that. Two ways forward: - **(a) Recommended for this issue:** - upload an archive (size-limited, into `POSTULO_BACKUP_DIR`) or pick one from the list; - **verify** it and show its manifest; - show the **exact** stop-restore-start commands for this installation (Compose or bare), with the archive path filled in. The restore itself stays on the command line. - **(b) Later, as its own issue:** a real maintenance mode (every request answers 503 with a page, the scheduler pauses), then a restore task that runs under it and requires re-authentication plus typing the instance's name. ## Depends on - #220: background tasks, so nothing runs in the request. - #221: a scheduler that doesn't double-run or die, for scheduled backups. - #234: the archive should include the plugins record and a field-key fingerprint, so an archive shown here is actually restorable; restore safety. - #219: `pg_dump` in the image, or backups on PostgreSQL fail from the web exactly as from the shell. ## Checks - **Security tests:** - non-staff get 404 on every Backups URL; - download and delete require recent re-authentication; - names with `..`, `/`, absolute paths or symlinks are refused; - an upload over the limit or not a Postulo archive is refused before it is kept. - **Page coverage:** add the page to the browser walk. The download and task-status URLs go in `EXCUSED` with reasons. - **Accessibility:** the list is a real table, progress is announced in a live region, and destructive actions use the shared danger styles (#227). - **Translations:** every new string, in the working catalogues. - **Wiki:** *Backups and your data* gains a section for the page, and *Configuration* documents the schedule and retention settings.
tiagoagueda added this to the 0.6.0 milestone 2026-09-15 21:40:04 +00:00
Author
Owner

On the maintenance mode this issue defers

Restoring from the web is left open above, with a real maintenance mode named as something
that "comes later". Worth recording what that would actually take, because the obvious
package solves about a third of it and the other two thirds are ours whatever we pick.

django-maintenance-mode is the standard answer: middleware that returns a 503 with a
maintenance template while a flag is set, management commands to set and clear it, and
exemptions for staff users, IP ranges and named URLs. Nothing in this repository does that
today.

It is a reasonable fit for the HTTP half. Three things it does not solve, all of which
this issue would meet:

  1. The flag cannot live in the database. Its backends include a local file, the cache
    and the database; a restore replaces the database underneath the running process, so a
    flag stored there disappears at exactly the moment it is load-bearing — and comes back
    holding whatever the restored archive thought. It has to be the file backend on a
    volume every web container can see, or the cache.

  2. It is HTTP-only, and we have two other processes. core/scheduler.py holds a lease
    and a heartbeat, and db_worker executes queued tasks. Neither goes through the request
    middleware, so both keep running and both keep writing while the archive is being laid
    down. Quiescing them — and knowing they have actually stopped, not merely been asked —
    is the part with no package behind it, and it is the part that decides whether a web
    restore is safe at all.

  3. A process does not survive its database being swapped. On SQLite the file is
    replaced; on PostgreSQL the connections are to a database being dropped and recreated.
    Open connections in every gunicorn worker have to be closed and reopened, which in
    practice means the restore ends in a restart rather than in a redirect.

So the recommendation in the issue stands unchanged: verify the archive and show the
command.
That is honest about where the operation actually happens. django-maintenance- mode is what would let that become a real button later, and it is worth naming here so
whoever picks that up starts from the file-backed flag and the two unguarded processes
rather than discovering them.

One smaller thing it would earn its place for sooner, independent of restore: "Back up
now" on a large media directory
. The issue already says that runs as a background task
and can take minutes. It does not need the site down, but it does need somebody not to be
editing a document that is halfway into the archive. A URL-scoped maintenance flag is one
way to say so; a banner is another and cheaper one. Worth a line in the design either way.

## On the maintenance mode this issue defers Restoring from the web is left open above, with a real maintenance mode named as something that "comes later". Worth recording what that would actually take, because the obvious package solves about a third of it and the other two thirds are ours whatever we pick. **`django-maintenance-mode`** is the standard answer: middleware that returns a 503 with a maintenance template while a flag is set, management commands to set and clear it, and exemptions for staff users, IP ranges and named URLs. Nothing in this repository does that today. It is a reasonable fit for the HTTP half. Three things it does **not** solve, all of which this issue would meet: 1. **The flag cannot live in the database.** Its backends include a local file, the cache and the database; a restore replaces the database underneath the running process, so a flag stored there disappears at exactly the moment it is load-bearing — and comes back holding whatever the *restored* archive thought. It has to be the file backend on a volume every web container can see, or the cache. 2. **It is HTTP-only, and we have two other processes.** `core/scheduler.py` holds a lease and a heartbeat, and `db_worker` executes queued tasks. Neither goes through the request middleware, so both keep running and both keep writing while the archive is being laid down. Quiescing them — and knowing they have actually stopped, not merely been asked — is the part with no package behind it, and it is the part that decides whether a web restore is safe at all. 3. **A process does not survive its database being swapped.** On SQLite the file is replaced; on PostgreSQL the connections are to a database being dropped and recreated. Open connections in every gunicorn worker have to be closed and reopened, which in practice means the restore ends in a restart rather than in a redirect. So the recommendation in the issue stands unchanged: **verify the archive and show the command.** That is honest about where the operation actually happens. `django-maintenance- mode` is what would let that become a real button later, and it is worth naming here so whoever picks that up starts from the file-backed flag and the two unguarded processes rather than discovering them. One smaller thing it *would* earn its place for sooner, independent of restore: **"Back up now" on a large media directory**. The issue already says that runs as a background task and can take minutes. It does not need the site down, but it does need somebody not to be editing a document that is halfway into the archive. A URL-scoped maintenance flag is one way to say so; a banner is another and cheaper one. Worth a line in the design either way.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#242
No description provided.