A catch-up cursor cannot walk past rows that share one updated_at #245

Closed
opened 2026-09-16 12:14:50 +00:00 by tiagoagueda · 0 comments
Owner

#230 gave every list an updated_since cursor: ask for what changed at or after a moment,
oldest change first, take the updated_at of the last row you read and ask again from
there. It works until several rows carry the same updated_at, and then it stops dead.

The filter is updated_at__gte=since. A client whose last row is one of a group sharing a
timestamp asks again from that timestamp and is handed the same group again. Where the group
is larger than limit, every page after that is rows it has already seen, for ever: the
client either loops or, if it stops when a page brings nothing new, silently never reads the
rest of the account. Nothing raises, nothing logs, and the copy it is keeping is quietly
short.

This is not a rare shape. updated_at is auto_now, so a bulk edit, an import, a migration
backfill or a spreadsheet import all write a run of rows within the same microsecond — and
the faster the machine, the longer the run. It surfaced as
tests/test_api_lists.py::test_catching_up_reads_forward_so_a_cursor_can_advance failing on
an idle machine while passing on a loaded one, which is the same fault wearing a disguise:
twelve rows created in a tight loop landed four-deep on one timestamp, and a four-row page
could not get past them.

Proposal

Make the cursor a keyset rather than a timestamp: alongside updated_since, an optional
after_id, and the filter

Q(updated_at__gt=since) | Q(updated_at=since, pk__gt=after_id)

which is exactly the order the list is already sorted in (updated_at, pk), so it names a
position rather than a moment. Without after_id the behaviour is what it is today, so
nothing that works now breaks.

  • Every row already carries updated_at; it carries id too, so a client has both halves
    and needs no new field.
  • The handbook's catch-up recipe changes to "remember the last updated_at and id".
  • The test walks with both and asserts it reaches every row, on any machine.

Found while landing #224 on top of #230, 2026-09-16.

#230 gave every list an `updated_since` cursor: ask for what changed at or after a moment, oldest change first, take the `updated_at` of the last row you read and ask again from there. It works until several rows carry the same `updated_at`, and then it stops dead. The filter is `updated_at__gte=since`. A client whose last row is one of a group sharing a timestamp asks again from that timestamp and is handed the same group again. Where the group is larger than `limit`, every page after that is rows it has already seen, for ever: the client either loops or, if it stops when a page brings nothing new, silently never reads the rest of the account. Nothing raises, nothing logs, and the copy it is keeping is quietly short. This is not a rare shape. `updated_at` is `auto_now`, so a bulk edit, an import, a migration backfill or a spreadsheet import all write a run of rows within the same microsecond — and the faster the machine, the longer the run. It surfaced as `tests/test_api_lists.py::test_catching_up_reads_forward_so_a_cursor_can_advance` failing on an idle machine while passing on a loaded one, which is the same fault wearing a disguise: twelve rows created in a tight loop landed four-deep on one timestamp, and a four-row page could not get past them. ## Proposal Make the cursor a keyset rather than a timestamp: alongside `updated_since`, an optional `after_id`, and the filter Q(updated_at__gt=since) | Q(updated_at=since, pk__gt=after_id) which is exactly the order the list is already sorted in (`updated_at`, `pk`), so it names a position rather than a moment. Without `after_id` the behaviour is what it is today, so nothing that works now breaks. - Every row already carries `updated_at`; it carries `id` too, so a client has both halves and needs no new field. - The handbook's catch-up recipe changes to "remember the last `updated_at` **and** `id`". - The test walks with both and asserts it reaches every row, on any machine. Found while landing #224 on top of #230, 2026-09-16.
tiagoagueda added this to the 0.4.0 milestone 2026-09-16 12:14:50 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#245
No description provided.