Manage instance backups from Server settings: list, back up now, download, delete, schedule and retention #242
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#242
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Administrators should be able to manage instance backups from Server settings: see which archives exist, take one now, download it, delete old ones, and set a schedule with retention. Today all of that is a shell command and a cron line on the host.
What exists today
The hard part is already built and tested (#32):
core/backup.py:write_backupwrites a manifest (Postulo version, engine, counts), the database (SQLite's backup API, orpg_dump) and the media directory, then verifies the archive;verify_backupre-checks an archive against its checksum;restore_backuprefuses an archive from the other engine and refuses a non-empty instance withoutforce.manage.py backup [target]andmanage.py restore <archive> [--force].POSTULO_BACKUP_DIR(config/settings/base.py:452):data/backupsby default,/app/data/backupsin the container.core/server_views.py:75,131), but links nowhere.What is missing is any of it in the interface.
Proposal: a Backups section in Server settings
A new
SettingsSection(core/server_sections.py), staff only, between Logs and Defaults.1. The list
POSTULO_BACKUP_DIR, newest first: when, size, Postulo version, engine, record counts from the manifest, and whether it verifies.2. Back up now
3. Download
FileResponsewithContent-Disposition: attachment.POSTULO_BACKUP_DIR, addressed by name: never a path from the request, never a symlink out of the directory.4. Delete
A confirmation page naming the archive and its date, re-authentication, and a log line. The newest verified archive cannot be deleted while it is the only one.
5. Schedule and retention
postulo_backup_last_success_timestamp_secondsandpostulo_backup_failures_total, documented on Health, metrics and logs.SiteSettings, and the host cron line in the wiki stays valid for people who prefer it.6. Restore: decide the scope
Restoring from the web is the dangerous half.
restore_backupoverwrites the database and media under a running instance, with gunicorn workers and the scheduler still connected, and #234 asksrestoreto refuse exactly that. Two ways forward:(a) Recommended for this issue:
POSTULO_BACKUP_DIR) or pick one from the list;The restore itself stays on the command line.
(b) Later, as its own issue: a real maintenance mode (every request answers 503 with a page, the scheduler pauses), then a restore task that runs under it and requires re-authentication plus typing the instance's name.
Depends on
pg_dumpin the image, or backups on PostgreSQL fail from the web exactly as from the shell.Checks
..,/, absolute paths or symlinks are refused;EXCUSEDwith reasons.On the maintenance mode this issue defers
Restoring from the web is left open above, with a real maintenance mode named as something
that "comes later". Worth recording what that would actually take, because the obvious
package solves about a third of it and the other two thirds are ours whatever we pick.
django-maintenance-modeis the standard answer: middleware that returns a 503 with amaintenance template while a flag is set, management commands to set and clear it, and
exemptions for staff users, IP ranges and named URLs. Nothing in this repository does that
today.
It is a reasonable fit for the HTTP half. Three things it does not solve, all of which
this issue would meet:
The flag cannot live in the database. Its backends include a local file, the cache
and the database; a restore replaces the database underneath the running process, so a
flag stored there disappears at exactly the moment it is load-bearing — and comes back
holding whatever the restored archive thought. It has to be the file backend on a
volume every web container can see, or the cache.
It is HTTP-only, and we have two other processes.
core/scheduler.pyholds a leaseand a heartbeat, and
db_workerexecutes queued tasks. Neither goes through the requestmiddleware, so both keep running and both keep writing while the archive is being laid
down. Quiescing them — and knowing they have actually stopped, not merely been asked —
is the part with no package behind it, and it is the part that decides whether a web
restore is safe at all.
A process does not survive its database being swapped. On SQLite the file is
replaced; on PostgreSQL the connections are to a database being dropped and recreated.
Open connections in every gunicorn worker have to be closed and reopened, which in
practice means the restore ends in a restart rather than in a redirect.
So the recommendation in the issue stands unchanged: verify the archive and show the
command. That is honest about where the operation actually happens.
django-maintenance- modeis what would let that become a real button later, and it is worth naming here sowhoever picks that up starts from the file-backed flag and the two unguarded processes
rather than discovering them.
One smaller thing it would earn its place for sooner, independent of restore: "Back up
now" on a large media directory. The issue already says that runs as a background task
and can take minutes. It does not need the site down, but it does need somebody not to be
editing a document that is halfway into the archive. A URL-scoped maintenance flag is one
way to say so; a banner is another and cheaper one. Worth a line in the design either way.