A results page holds forty listings and Postulo can capture one of them #177

Closed
opened 2026-09-12 08:50:21 +00:00 by tiagoagueda · 1 comment
Owner

Observation

#160 opens by saying what a listings page is for:

A listing is the shortest-lived thing Postulo holds. Most are discarded; the useful
gesture is look at forty, keep three.

You cannot get forty in. Capture reads one posting per page, so the gesture the listings
page is being rebuilt around has no way to fill it.

What exists

postulo-chromium/src/lib/parse.js returns exactly one posting:

export function readPosting(doc, url, parse) { ... return {data, source} ... }

and bestPosting picks the one whose url, @id or identifier names the page being
looked at, discarding the rest. That was added deliberately in #176 — before it, a capture
from a results page brought back whichever advert the board happened to list first — but it
solves "which one" by throwing the other thirty-nine away.

popup.js is one capture per click: read the tab, fill one form, keep or send. The
store keeps the page it was read from per capture (page:<id> in lib/store.js).

The server already does the hard half

POST /captures does not trust the client's reading. _read() in api/api.py parses the
supplied HTML again
and the payload's data is applied as corrections over that:

The page is still read; each field given replaces what was read

Which means the shape of this is already in place. The extension can send, for each posting
somebody ticked, the same results-page HTML with url set to that posting's own
address
, and bestPosting picks the right one server-side by the correlation #176 added.
One posting per request, no new endpoint, no schema change, and a partial failure loses one
capture rather than forty. The search-results fixture in tests/fixtures/postings/ is
exactly this case and already passes.

So this is an extension feature with three server-side consequences, not an API redesign.

What has to be solved

The page is kept and sent once per capture. store.add({url, html, ...}) writes a
page:<id> for every capture, and the outbox sends it with each. Forty captures off one
results page means forty copies of that page in browser storage and forty copies on the
wire. fetching.MAX_BYTES is 2,000,000 per page. The page wants to be stored once and
referenced by the captures read from it, in the store and in the request — and slimPage
already makes it much smaller than the original, which helps and does not solve it.

Notifications fire per capture. create_capture calls notify() for every one:

The one event a person cannot see coming: something arrived from outside.

That reasoning holds for a capture arriving on its own and inverts completely for forty
arriving because somebody just pressed a button. Forty notifications for one deliberate
gesture is worse than none.

The rate limits disagree, and this must not be "fixed" by accident.
throttle.capture (POSTULO_CAPTURE_RATE, 30/h) is called in exactly one place —
jobs/capture_views.py:88, the web form. The API path is bounded by POSTULO_API_RATE
(600/h per token) and nothing else. So forty POSTs fit today. Anyone tempted to make the API
consistent with the form should read the separate issue about that first: the form's
comment says it is tight because capture is "the only surface that makes this server talk to
somebody else's", and that is about the fetch, not the parse. Applying 30/h to captures
that arrive with their HTML would cap this feature at thirty.

A posting on a results page is a stub. Boards put a title, a company and a place in
their list markup and leave the description to the advert's own page. Those captures will
arrive thin. Decide whether that is fine — a listing is a thing to triage, and #160 says
most are discarded — or whether accepting one should fetch its page. Fine is the cheaper
answer and probably the right one; fetching forty pages is the thing the capture throttle
exists to prevent.

Roughly, the work

  • parse.js: a reader that returns every posting on the page, with readPosting kept as
    the one-posting entry the popup uses today. bestPosting stays; it is how the server picks.
  • popup.js: when a page holds more than one, offer them as a list to tick rather than a
    form to correct. One posting stays exactly as it is now.
  • lib/store.js: one stored page per page, not per capture, with captures referring to it.
  • api/api.py: quiet the per-capture notification when several arrive together, or summarise.
  • A fixture in the corpus for a real results page, both sides, as #176 established.

Not in this issue: reviewing them once they arrive, which is its own gesture and its own
issue, and telling you that you have captured one before, which is another.

## Observation #160 opens by saying what a listings page is for: > A listing is the shortest-lived thing Postulo holds. Most are discarded; the useful > gesture is *look at forty, keep three*. You cannot get forty in. Capture reads one posting per page, so the gesture the listings page is being rebuilt around has no way to fill it. ## What exists `postulo-chromium/src/lib/parse.js` returns exactly one posting: export function readPosting(doc, url, parse) { ... return {data, source} ... } and `bestPosting` picks the one whose `url`, `@id` or `identifier` names the page being looked at, discarding the rest. That was added deliberately in #176 — before it, a capture from a results page brought back whichever advert the board happened to list first — but it solves "which one" by throwing the other thirty-nine away. `popup.js` is one capture per click: read the tab, fill one form, `keep` or `send`. The store keeps the page it was read from per capture (`page:<id>` in `lib/store.js`). ## The server already does the hard half `POST /captures` does not trust the client's reading. `_read()` in `api/api.py` **parses the supplied HTML again** and the payload's `data` is applied as corrections over that: > The page is still read; each field given replaces what was read Which means the shape of this is already in place. The extension can send, for each posting somebody ticked, the **same results-page HTML** with `url` set to **that posting's own address**, and `bestPosting` picks the right one server-side by the correlation #176 added. One posting per request, no new endpoint, no schema change, and a partial failure loses one capture rather than forty. The `search-results` fixture in `tests/fixtures/postings/` is exactly this case and already passes. So this is an extension feature with three server-side consequences, not an API redesign. ## What has to be solved **The page is kept and sent once per capture.** `store.add({url, html, ...})` writes a `page:<id>` for every capture, and the outbox sends it with each. Forty captures off one results page means forty copies of that page in browser storage and forty copies on the wire. `fetching.MAX_BYTES` is 2,000,000 per page. The page wants to be stored once and referenced by the captures read from it, in the store and in the request — and `slimPage` already makes it much smaller than the original, which helps and does not solve it. **Notifications fire per capture.** `create_capture` calls `notify()` for every one: > The one event a person cannot see coming: something arrived from outside. That reasoning holds for a capture arriving on its own and inverts completely for forty arriving because somebody just pressed a button. Forty notifications for one deliberate gesture is worse than none. **The rate limits disagree, and this must not be "fixed" by accident.** `throttle.capture` (`POSTULO_CAPTURE_RATE`, 30/h) is called in exactly one place — `jobs/capture_views.py:88`, the web form. The API path is bounded by `POSTULO_API_RATE` (600/h per token) and nothing else. So forty POSTs fit today. Anyone tempted to make the API consistent with the form should read the *separate* issue about that first: the form's comment says it is tight because capture is "the only surface that makes this server talk to somebody else's", and that is about the **fetch**, not the parse. Applying 30/h to captures that arrive with their HTML would cap this feature at thirty. **A posting on a results page is a stub.** Boards put a title, a company and a place in their list markup and leave the description to the advert's own page. Those captures will arrive thin. Decide whether that is fine — a listing is a thing to triage, and #160 says most are discarded — or whether accepting one should fetch its page. Fine is the cheaper answer and probably the right one; fetching forty pages is the thing the capture throttle exists to prevent. ## Roughly, the work - `parse.js`: a reader that returns **every** posting on the page, with `readPosting` kept as the one-posting entry the popup uses today. `bestPosting` stays; it is how the server picks. - `popup.js`: when a page holds more than one, offer them as a list to tick rather than a form to correct. One posting stays exactly as it is now. - `lib/store.js`: one stored page per *page*, not per capture, with captures referring to it. - `api/api.py`: quiet the per-capture notification when several arrive together, or summarise. - A fixture in the corpus for a real results page, both sides, as #176 established. Not in this issue: reviewing them once they arrive, which is its own gesture and its own issue, and telling you that you have captured one before, which is another.
Author
Owner

Landed on both sides:

  • core 0f2efba88 — POST /api/v1/captures takes an optional batch: {size, position}.
    The first of a batch is announced with the count — Captured 40 postings from
    careers.example.org
    , linking to the review list — and the rest arrive quietly; a batch
    that makes no sense (one of one, a position past the end) is a 422 before anything is
    stored. Announced on the first rather than the last on purpose: the last may never come, a
    refused one is retried later on its own, and a promise of forty is nearer the truth than
    silence.
  • postulo-chromium d35eb4d — readPostings() reads every posting a page lists, each as
    the address it states, and says which of them is this page: an advert that also lists
    similar jobs is still captured as one; only a page about none of them becomes a list to
    tick, nothing ticked, All and None to hand. Each ticked posting is its own capture off
    the one page: the store keeps the page once and drops it when the last capture read from
    it has gone; the outbox sends the shared page with each posting's own address and its
    place in the batch. Eight strings in all forty locales. README says what a results page
    becomes.
  • wiki — the API page documents batch, and that a capture read off a results page is
    a stub to triage rather than a listing to read; nothing fetches the forty adverts, which
    is what the capture throttle exists to prevent.

The rate-limit disagreement this issue warned about is #194, filed on purpose rather
than "fixed by accident" here: the 30/h limit is about the fetch, and a capture arriving with
its HTML fetches nothing.

Two things the tests caught before shipping: the store lists newest first and the outbox
sends oldest first, so a batch indexed in page order went out last-to-first — position 1
arriving last and the announcement with it — and is now indexed with its last posting as the
newest. Not reviewed here, as the issue said: reviewing them once they arrive, and saying you
captured one before.

Landed on both sides: - **core `0f2efba88`** — `POST /api/v1/captures` takes an optional `batch: {size, position}`. The first of a batch is announced with the count — *Captured 40 postings from careers.example.org*, linking to the review list — and the rest arrive quietly; a batch that makes no sense (one of one, a position past the end) is a `422` before anything is stored. Announced on the first rather than the last on purpose: the last may never come, a refused one is retried later on its own, and a promise of forty is nearer the truth than silence. - **postulo-chromium `d35eb4d`** — `readPostings()` reads every posting a page lists, each as the address it states, and says which of them *is* this page: an advert that also lists similar jobs is still captured as one; only a page about none of them becomes a list to tick, nothing ticked, *All* and *None* to hand. Each ticked posting is its own capture off the one page: the store keeps the page once and drops it when the last capture read from it has gone; the outbox sends the shared page with each posting's own address and its place in the batch. Eight strings in all forty locales. README says what a results page becomes. - **wiki** — the API page documents `batch`, and that a capture read off a results page is a stub to triage rather than a listing to read; nothing fetches the forty adverts, which is what the capture throttle exists to prevent. The rate-limit disagreement this issue warned about is **#194**, filed on purpose rather than "fixed by accident" here: the 30/h limit is about the fetch, and a capture arriving with its HTML fetches nothing. Two things the tests caught before shipping: the store lists newest first and the outbox sends oldest first, so a batch indexed in page order went out last-to-first — position 1 arriving last and the announcement with it — and is now indexed with its last posting as the newest. Not reviewed here, as the issue said: reviewing them once they arrive, and saying you captured one before.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#177
No description provided.