A results page holds forty listings and Postulo can capture one of them #177
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#177
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observation
#160 opens by saying what a listings page is for:
You cannot get forty in. Capture reads one posting per page, so the gesture the listings
page is being rebuilt around has no way to fill it.
What exists
postulo-chromium/src/lib/parse.jsreturns exactly one posting:and
bestPostingpicks the one whoseurl,@idoridentifiernames the page beinglooked at, discarding the rest. That was added deliberately in #176 — before it, a capture
from a results page brought back whichever advert the board happened to list first — but it
solves "which one" by throwing the other thirty-nine away.
popup.jsis one capture per click: read the tab, fill one form,keeporsend. Thestore keeps the page it was read from per capture (
page:<id>inlib/store.js).The server already does the hard half
POST /capturesdoes not trust the client's reading._read()inapi/api.pyparses thesupplied HTML again and the payload's
datais applied as corrections over that:Which means the shape of this is already in place. The extension can send, for each posting
somebody ticked, the same results-page HTML with
urlset to that posting's ownaddress, and
bestPostingpicks the right one server-side by the correlation #176 added.One posting per request, no new endpoint, no schema change, and a partial failure loses one
capture rather than forty. The
search-resultsfixture intests/fixtures/postings/isexactly this case and already passes.
So this is an extension feature with three server-side consequences, not an API redesign.
What has to be solved
The page is kept and sent once per capture.
store.add({url, html, ...})writes apage:<id>for every capture, and the outbox sends it with each. Forty captures off oneresults page means forty copies of that page in browser storage and forty copies on the
wire.
fetching.MAX_BYTESis 2,000,000 per page. The page wants to be stored once andreferenced by the captures read from it, in the store and in the request — and
slimPagealready makes it much smaller than the original, which helps and does not solve it.
Notifications fire per capture.
create_capturecallsnotify()for every one:That reasoning holds for a capture arriving on its own and inverts completely for forty
arriving because somebody just pressed a button. Forty notifications for one deliberate
gesture is worse than none.
The rate limits disagree, and this must not be "fixed" by accident.
throttle.capture(POSTULO_CAPTURE_RATE, 30/h) is called in exactly one place —jobs/capture_views.py:88, the web form. The API path is bounded byPOSTULO_API_RATE(600/h per token) and nothing else. So forty POSTs fit today. Anyone tempted to make the API
consistent with the form should read the separate issue about that first: the form's
comment says it is tight because capture is "the only surface that makes this server talk to
somebody else's", and that is about the fetch, not the parse. Applying 30/h to captures
that arrive with their HTML would cap this feature at thirty.
A posting on a results page is a stub. Boards put a title, a company and a place in
their list markup and leave the description to the advert's own page. Those captures will
arrive thin. Decide whether that is fine — a listing is a thing to triage, and #160 says
most are discarded — or whether accepting one should fetch its page. Fine is the cheaper
answer and probably the right one; fetching forty pages is the thing the capture throttle
exists to prevent.
Roughly, the work
parse.js: a reader that returns every posting on the page, withreadPostingkept asthe one-posting entry the popup uses today.
bestPostingstays; it is how the server picks.popup.js: when a page holds more than one, offer them as a list to tick rather than aform to correct. One posting stays exactly as it is now.
lib/store.js: one stored page per page, not per capture, with captures referring to it.api/api.py: quiet the per-capture notification when several arrive together, or summarise.Not in this issue: reviewing them once they arrive, which is its own gesture and its own
issue, and telling you that you have captured one before, which is another.
Landed on both sides:
0f2efba88—POST /api/v1/capturestakes an optionalbatch: {size, position}.The first of a batch is announced with the count — Captured 40 postings from
careers.example.org, linking to the review list — and the rest arrive quietly; a batch
that makes no sense (one of one, a position past the end) is a
422before anything isstored. Announced on the first rather than the last on purpose: the last may never come, a
refused one is retried later on its own, and a promise of forty is nearer the truth than
silence.
d35eb4d—readPostings()reads every posting a page lists, each asthe address it states, and says which of them is this page: an advert that also lists
similar jobs is still captured as one; only a page about none of them becomes a list to
tick, nothing ticked, All and None to hand. Each ticked posting is its own capture off
the one page: the store keeps the page once and drops it when the last capture read from
it has gone; the outbox sends the shared page with each posting's own address and its
place in the batch. Eight strings in all forty locales. README says what a results page
becomes.
batch, and that a capture read off a results page isa stub to triage rather than a listing to read; nothing fetches the forty adverts, which
is what the capture throttle exists to prevent.
The rate-limit disagreement this issue warned about is #194, filed on purpose rather
than "fixed by accident" here: the 30/h limit is about the fetch, and a capture arriving with
its HTML fetches nothing.
Two things the tests caught before shipping: the store lists newest first and the outbox
sends oldest first, so a batch indexed in page order went out last-to-first — position 1
arriving last and the announcement with it — and is now indexed with its last posting as the
newest. Not reviewed here, as the issue said: reviewing them once they arrive, and saying you
captured one before.