Move the slow work onto django-tasks-db, and show a "working…" state #247

Closed
opened 2026-09-16 13:33:32 +00:00 by tiagoagueda · 2 comments
Owner

django-tasks-db is configured (config/settings/base.py, TASKS["default"]["BACKEND"]),
django_tasks_db is installed, and nothing anywhere enqueues anything.
core/server_views.py:_queued_tasks already counts DBTaskResult rows with status READY
for the Overview page, so server settings can report a queue that is structurally always
empty.

#220 took the slow views out of the request's transaction, so one of them no longer holds
the SQLite write lock while it works. That was the half that was a bug. It did not make them
fast, and it did not stop them occupying a gunicorn worker for the whole time: a capture
still waits up to fifteen seconds with a tab spinning, a CV is still rendered while somebody
watches, and an export of a large account is still one long request. This is the other half.

What moves

  • Capture fetching — jobs/capture_views.py, CaptureCreateView.post. fetch_page
    allows ten seconds for the page and five for robots.txt.
  • Logo fetching — jobs/views.py, CompanyLogoActionView.post. Find logo reads the
    company's site and then tries up to six images, one round trip each.
  • PDF rendering — documents/views.py, CVExportView.post and SendDocumentsView.post;
    applications/views.py, ReportPDFView. On the Chromium backend this is a browser launch
    plus a render; on a small machine it is why the container now allows a worker two minutes
    instead of thirty seconds (#220).
  • Export archives — core/views_export.py, export_download. Written to a temporary
    file rather than the BytesIO in core/export.py:write_archive, and handed over as a
    download once it exists.
  • Notification delivery from the API — api/api.py:_capture calls notify()
    synchronously, so a batch of forty from the browser extension waits on every notifier's
    network timeout, forty times. notifications/service.py:notify is already written to
    survive a failing notifier, so it is the delivery that is in the wrong place.

Two things the queue needs before any of this is useful

  1. Something has to run manage.py db_worker. Neither Compose file mentions it and the
    Dockerfile runs only gunicorn, so queueing work today means queueing it for ever. The
    obvious shape is the one the scheduler already has: a second service under a profile,
    sharing the volume and the database URL, with POSTULO_SKIP_MIGRATE=1 and
    POSTULO_SKIP_PLUGIN_SYNC=1 (#221), and a healthcheck that notices when it has stopped.
  2. It must be optional. Postulo is self-hosted and some people run one container. Either
    the work falls back to running inline when no worker is configured, or the person is told
    plainly, on the page, that what they pressed is waiting for a worker that is not running.
    A button that silently does nothing is worse than a slow button.

The UI half

Each of these is a POST that today answers with the finished thing. Queued, each has to
answer with a pending state that htmx polls until it resolves:

  • somewhere for the state to live so a poll can find it — a task id on the record, or a
    small model holding the task, its owner, its kind and its outcome;
  • a partial rendering "working…", "done" with the result, or "failed" with the reason, in a
    live region so the change is announced rather than merely painted;
  • ownership checked on the poll endpoint, not only on the POST, because the poll is a new
    address that names a task;
  • what happens when the browser is closed, and where the person finds the result afterwards
    — a rendered document is already in Sent documents, an export archive is nowhere yet;
  • what happens to a task whose record was deleted while it waited.

Worth deciding deliberately

  • The export archive is a file that has to be reaped. One per export, holding every
    document in an account, left on disk if nobody downloads it. It needs an owner, an expiry,
    something that deletes it, and it must be reachable only through an ownership-checked view
    like every other personal document.
  • The rate limit is spent before the fetch, not in the worker. throttle.capture is
    consumed in the view, and #220 deliberately left it outside any transaction. Moving the
    fetch must not move the counting, or an account can queue a thousand fetches for one.
  • Idempotency. api/idempotency.py claims a key before the work starts and gives it up
    if the request failed. With the work in a worker, "the request failed" and "the work
    failed" stop being the same event.
  • A retry is not always wanted. Rendering a CV twice files two snapshots.
    snapshot_report already dedupes on source_text; nothing else does.
  • db_worker is another writer on the SQLite file, and should follow the rule #220
    established: short transactions around writes, nothing slow inside one.
  • A draft PDF cache already exists in documents/pdf.py, in-process and bounded, used
    only by ReportPDFView.get. If rendering moves to a worker, a per-worker cache is a
    different cache from the one the web process would read, and the note there explaining why
    it is not in Django's cache — the default cache is a table in the same database — still
    applies.

Not in this

The non_atomic_requests work is done (#220), except for the capture API: django-ninja
resolves one Django view per path, inside PathView.get_view(), so opting that one endpoint
out means either reaching into the router or opting the whole API out at once and making
every endpoint responsible for its own transactions. Worth doing, a change of its own, and
possibly moot if the fetch and the notify both move onto the queue here.

`django-tasks-db` is configured (`config/settings/base.py`, `TASKS["default"]["BACKEND"]`), `django_tasks_db` is installed, and nothing anywhere enqueues anything. `core/server_views.py:_queued_tasks` already counts `DBTaskResult` rows with status `READY` for the Overview page, so server settings can report a queue that is structurally always empty. #220 took the slow views out of the request's transaction, so one of them no longer holds the SQLite write lock while it works. That was the half that was a bug. It did not make them fast, and it did not stop them occupying a gunicorn worker for the whole time: a capture still waits up to fifteen seconds with a tab spinning, a CV is still rendered while somebody watches, and an export of a large account is still one long request. This is the other half. ## What moves - **Capture fetching** — `jobs/capture_views.py`, `CaptureCreateView.post`. `fetch_page` allows ten seconds for the page and five for `robots.txt`. - **Logo fetching** — `jobs/views.py`, `CompanyLogoActionView.post`. *Find logo* reads the company's site and then tries up to six images, one round trip each. - **PDF rendering** — `documents/views.py`, `CVExportView.post` and `SendDocumentsView.post`; `applications/views.py`, `ReportPDFView`. On the Chromium backend this is a browser launch plus a render; on a small machine it is why the container now allows a worker two minutes instead of thirty seconds (#220). - **Export archives** — `core/views_export.py`, `export_download`. Written to a temporary file rather than the `BytesIO` in `core/export.py:write_archive`, and handed over as a download once it exists. - **Notification delivery from the API** — `api/api.py:_capture` calls `notify()` synchronously, so a batch of forty from the browser extension waits on every notifier's network timeout, forty times. `notifications/service.py:notify` is already written to survive a failing notifier, so it is the delivery that is in the wrong place. ## Two things the queue needs before any of this is useful 1. **Something has to run `manage.py db_worker`.** Neither Compose file mentions it and the Dockerfile runs only gunicorn, so queueing work today means queueing it for ever. The obvious shape is the one the scheduler already has: a second service under a profile, sharing the volume and the database URL, with `POSTULO_SKIP_MIGRATE=1` and `POSTULO_SKIP_PLUGIN_SYNC=1` (#221), and a healthcheck that notices when it has stopped. 2. **It must be optional.** Postulo is self-hosted and some people run one container. Either the work falls back to running inline when no worker is configured, or the person is told plainly, on the page, that what they pressed is waiting for a worker that is not running. A button that silently does nothing is worse than a slow button. ## The UI half Each of these is a POST that today answers with the finished thing. Queued, each has to answer with a pending state that htmx polls until it resolves: - somewhere for the state to live so a poll can find it — a task id on the record, or a small model holding the task, its owner, its kind and its outcome; - a partial rendering "working…", "done" with the result, or "failed" with the reason, in a live region so the change is announced rather than merely painted; - ownership checked on the poll endpoint, not only on the POST, because the poll is a new address that names a task; - what happens when the browser is closed, and where the person finds the result afterwards — a rendered document is already in *Sent documents*, an export archive is nowhere yet; - what happens to a task whose record was deleted while it waited. ## Worth deciding deliberately - **The export archive is a file that has to be reaped.** One per export, holding every document in an account, left on disk if nobody downloads it. It needs an owner, an expiry, something that deletes it, and it must be reachable only through an ownership-checked view like every other personal document. - **The rate limit is spent before the fetch, not in the worker.** `throttle.capture` is consumed in the view, and #220 deliberately left it outside any transaction. Moving the fetch must not move the counting, or an account can queue a thousand fetches for one. - **Idempotency.** `api/idempotency.py` claims a key before the work starts and gives it up if the request failed. With the work in a worker, "the request failed" and "the work failed" stop being the same event. - **A retry is not always wanted.** Rendering a CV twice files two snapshots. `snapshot_report` already dedupes on `source_text`; nothing else does. - **`db_worker` is another writer on the SQLite file**, and should follow the rule #220 established: short transactions around writes, nothing slow inside one. - **A draft PDF cache already exists** in `documents/pdf.py`, in-process and bounded, used only by `ReportPDFView.get`. If rendering moves to a worker, a per-worker cache is a different cache from the one the web process would read, and the note there explaining why it is not in Django's cache — the default cache is a table in the same database — still applies. ## Not in this The `non_atomic_requests` work is done (#220), **except for the capture API**: django-ninja resolves one Django view per path, inside `PathView.get_view()`, so opting that one endpoint out means either reaching into the router or opting the whole API out at once and making every endpoint responsible for its own transactions. Worth doing, a change of its own, and possibly moot if the fetch and the `notify` both move onto the queue here.
tiagoagueda added this to the 0.4.0 milestone 2026-09-16 13:33:32 +00:00
Author
Owner

Three refinements, all found inside what is already installed

Checked against the versions in the lock: Django 6.1.1 and django-tasks-db 0.13.0.

1. "It must be optional" is a settings value, not a fallback branch

Point 2 above asks for the work to run inline when no worker is configured, or for the
person to be told plainly. Django ships the first of those: django.tasks.backends holds
immediate and dummy alongside base.

So a single-container install sets ImmediateBackend and every enqueue() runs in the
request, on the same code path, with nothing in the views asking whether a worker exists.
That removes the branch this issue was bracing for.

Two things to say plainly about it:

  • Inline means the request waits again, which is exactly today's behaviour. Not a
    regression, and the honest answer for somebody running one container — but it means the
    "working…" state described below has to render correctly for a task that is already
    finished by the time the response is written.
  • DummyBackend is the wrong choice here. It accepts work and discards it, which is
    precisely the "button that silently does nothing" this issue warns against. It belongs in
    tests and nowhere else.

The setting then becomes part of the self-hosting documentation: one container, immediate;
worker running, database backend.

2. Nothing prunes DBTaskResult, and the command already exists

django-tasks-db ships manage.py prune_db_task_results, taking --min-age-days and a
separate minimum age for failed results (defaulting to the same). Nothing runs it: the
scheduler service runs send_due_reminders --loop --every 300 and nothing else.

The moment this issue lands, every capture, logo fetch, PDF render, export and API
notification writes a row that is never removed — an unbounded table, one row per
background action, on SQLite by default. It should be a second pass in the scheduler
alongside the reminders, and the two ages should differ: a succeeded result is worth days,
a failed one is worth however long somebody might reasonably come looking for it.

Worth deciding here rather than discovering on an instance a year old.

3. There are no task-level retries, and this issue is what makes that matter

django_tasks_db.utils.retry is an internal decorator guarding the worker's own database
writes against lock contention. It is not "re-run this task". Nothing re-runs a failed task.

That is fine today because nothing is queued. It stops being fine with the list above,
because two of the five are network operations against an address somebody else chose:
fetch_page allows ten seconds for the page and five for robots.txt, and Find logo
makes up to six round trips to a site we do not control.

Today a timeout is a visible failure and the person presses the button again. Queued, it
becomes a row with FAILED status, behind a spinner that stops, with no button. A retry
policy is a decision this issue has to make, not one it inherits
— whether that is
re-enqueue with a delay, or no retry and a clear failure the person can act on. Either is
defensible; silence is not.

One line for the Overview page

_queued_tasks counts rows with status READY. Once work is genuinely queued, the number
that matters is FAILED — because a queue that is always empty and a queue full of
failures produce the same Overview today. That is one more count and one more sentence, and
it belongs with this change rather than after it.

And the consequence for a self-hosted install

With db_worker added, a full deployment is three processes: gunicorn, the scheduler, and
the worker. Both extras are Compose profiles, so a default install gets neither.
ImmediateBackend is what keeps that default working rather than quietly broken, which is
the same argument point 2 was already making — it just has a name now.

## Three refinements, all found inside what is already installed Checked against the versions in the lock: Django **6.1.1** and `django-tasks-db` **0.13.0**. ### 1. "It must be optional" is a settings value, not a fallback branch Point 2 above asks for the work to run inline when no worker is configured, or for the person to be told plainly. Django ships the first of those: `django.tasks.backends` holds **`immediate`** and **`dummy`** alongside `base`. So a single-container install sets `ImmediateBackend` and every `enqueue()` runs in the request, on the same code path, with nothing in the views asking whether a worker exists. That removes the branch this issue was bracing for. Two things to say plainly about it: - **Inline means the request waits again**, which is exactly today's behaviour. Not a regression, and the honest answer for somebody running one container — but it means the "working…" state described below has to render correctly for a task that is already finished by the time the response is written. - **`DummyBackend` is the wrong choice here.** It accepts work and discards it, which is precisely the "button that silently does nothing" this issue warns against. It belongs in tests and nowhere else. The setting then becomes part of the self-hosting documentation: one container, immediate; worker running, database backend. ### 2. Nothing prunes `DBTaskResult`, and the command already exists `django-tasks-db` ships `manage.py prune_db_task_results`, taking `--min-age-days` and a *separate* minimum age for failed results (defaulting to the same). Nothing runs it: the scheduler service runs `send_due_reminders --loop --every 300` and nothing else. The moment this issue lands, every capture, logo fetch, PDF render, export and API notification writes a row that is never removed — an unbounded table, one row per background action, on SQLite by default. It should be a second pass in the scheduler alongside the reminders, and the two ages should differ: a succeeded result is worth days, a failed one is worth however long somebody might reasonably come looking for it. Worth deciding here rather than discovering on an instance a year old. ### 3. There are no task-level retries, and this issue is what makes that matter `django_tasks_db.utils.retry` is an internal decorator guarding the worker's own database writes against lock contention. It is not "re-run this task". Nothing re-runs a failed task. That is fine today because nothing is queued. It stops being fine with the list above, because two of the five are network operations against an address somebody else chose: `fetch_page` allows ten seconds for the page and five for `robots.txt`, and *Find logo* makes up to six round trips to a site we do not control. Today a timeout is a visible failure and the person presses the button again. Queued, it becomes a row with `FAILED` status, behind a spinner that stops, with no button. **A retry policy is a decision this issue has to make, not one it inherits** — whether that is re-enqueue with a delay, or no retry and a clear failure the person can act on. Either is defensible; silence is not. ### One line for the Overview page `_queued_tasks` counts rows with status `READY`. Once work is genuinely queued, the number that matters is `FAILED` — because a queue that is always empty and a queue full of failures produce the same Overview today. That is one more count and one more sentence, and it belongs with this change rather than after it. ### And the consequence for a self-hosted install With `db_worker` added, a full deployment is three processes: gunicorn, the scheduler, and the worker. Both extras are Compose profiles, so a default install gets neither. `ImmediateBackend` is what keeps that default working rather than quietly broken, which is the same argument point 2 was already making — it just has a name now.
Author
Owner

Landed on main as 46ef9919d.

Landed on `main` as 46ef9919d.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#247
No description provided.