A capture should be able to keep the page: a rendering of the whole thing, and the source #256

Open
opened 2026-09-17 14:57:07 +00:00 by tiagoagueda · 0 comments
Owner

A capture reads a page, keeps what it understood, and throws the page away. fetching.py
returns a FetchedPage, parse_page turns it into JobPostingData, and
Capture.data stores that JSON — the markup is never written anywhere. The same is true of
the source a browser extension hands over through the API: it is parsed on arrival and
dropped.

A capture should be able to keep two things beside the parsed fields, both optional:

  1. A rendering of the whole page, not the viewport — the advert as it looked, including
    everything below the fold.
  2. The source as it arrived, exactly the bytes that were parsed.

Why it earns the disk

Because the parse is a guess and the page is the evidence. Capture's own docstring
says so: "Parsing somebody else's markup is guesswork often enough that turning the result
straight into a record would put invented job titles into the one place they must not be."
Everything captured waits for a person to look at it — and what they are checking it against
is a browser tab they may have already closed. The rendering puts the original beside the
reading.

Because postings vanish. JobPosting has both closes_at and closed_at: the model
already knows an advert dies. When it does, the URL is a 404 and the parsed JSON is the only
surviving account of what the job was — no salary line as written, no list of requirements,
nothing to check a later dispute against. Postulo already believes in freezing what was true
at a moment; that is what snapshot-on-send does for a document that has left the building. A
capture is the same promise pointing the other way.

Because it is the difference between a record and evidence. The report CSV exists to be
handed to an employment office. "I applied to this on this date" is a stronger statement when
the advert is attached than when it is a row someone typed.

Because a parser improves and old captures do not. This is the one that compounds. Today
the source plugin that matched runs once, at fetch time, and its reading is all that survives.
Keep the source and a capture can be read again — by a better parser, by a plugin that did
not exist when it was taken, by a person fixing a field by hand months later. Every parser fix
currently helps future captures only. This is what "post-processing" means in practice and it
is worth more than the archival argument.

Privacy: two switches, not one

The two artefacts leak differently and must be refusable separately.

  • A rendering is a picture. If it came from the extension it is a picture of a page as
    you were seeing it, which can include your own name in the corner, a signed-in session, a
    personalised salary band, a recruiter's direct line.
  • The source is worse in a quieter way: tracking parameters, embedded identifiers, and
    whatever the page addressed to you personally.

So: keep both, keep one, keep neither. Where the switch lives is already decided by
precedent — CaptureView is a PolicyView over the site policy, with pinned_fields for
what the environment has fixed, exactly as capture_ignore_robots works. An operator's no
has to be final; a user's preference may only narrow it further, never widen it.

Default off, and say why in the settings page the way the robots switch does. Keeping a
copy of somebody else's page is a decision an operator should make on purpose.

The hard parts

The server cannot draw most of these pages, and often cannot see them at all. WeasyPrint
is the default renderer and runs no JavaScript, so it would produce a faithful rendering of
an empty shell. A real full-page rendering needs the Chromium backend — which exists,
ChromiumBackend over Playwright in documents/pdf.py, with a session() that keeps one
browser open across documents — but it is optional, and an instance that has never installed
Playwright must degrade to "source only" rather than to a broken picture.

Worse: #194's finding was that large employers sit behind bot protection that refuses
anything that is not a browser, which is exactly why the API accepts html from the
extension. For those pages a server-side render is a photograph of a login wall. The
extension is the right renderer for the hard cases
— it is the only thing holding the
session — so postulo-chromium and postulo-firefox should be able to send a rendering and
the source alongside the html they already send, and CaptureIn should accept them. That
makes this a change in three repositories, which is worth knowing before it is scheduled.

Never serve it back as a page. #218 is the precedent and the rule it set applies here
with force: text from a stranger's page, stored and later drawn, is that stranger's code
running as the person who opened it. An archived source must leave Postulo as a download
with a content type that is not HTML, or be shown inside a sandboxed frame with a policy of
its own. It must never be rendered on Postulo's own origin. Whatever is built here needs a
test that says so.

Size, and what an instance is agreeing to. The fetch cap is 2 MB of source
(fetching.MAX_BYTES); the extension path has no such cap because it never had to. A
full-page rendering of a long advert is larger than either. This needs its own limit, a
retention rule — captures that were never confirmed are the obvious candidates for expiry —
and it has to show up wherever the instance reports what it is storing. The files are
per-owner and go through the private media path, like every other uploaded file.

Format. A PDF from the existing Chromium backend is one searchable file, printable, and
needs no new machinery. An image is more faithful and much larger. Single-file HTML is
re-viewable and is the thing the paragraph above says not to render. PDF for the rendering
and the source stored as compressed text is the proposal; it is worth arguing with.

robots.txt. Postulo honours it when fetching. Keeping a copy of a page it was allowed to
fetch, for the person who asked for it, does not seem to raise a second question — but the
settings page should say what it does rather than leave a reader to wonder.

A capture reads a page, keeps what it understood, and throws the page away. `fetching.py` returns a `FetchedPage`, `parse_page` turns it into `JobPostingData`, and `Capture.data` stores that JSON — the markup is never written anywhere. The same is true of the source a browser extension hands over through the API: it is parsed on arrival and dropped. A capture should be able to keep two things beside the parsed fields, both optional: 1. **A rendering of the whole page**, not the viewport — the advert as it looked, including everything below the fold. 2. **The source as it arrived**, exactly the bytes that were parsed. ## Why it earns the disk **Because the parse is a guess and the page is the evidence.** `Capture`'s own docstring says so: "Parsing somebody else's markup is guesswork often enough that turning the result straight into a record would put invented job titles into the one place they must not be." Everything captured waits for a person to look at it — and what they are checking it against is a browser tab they may have already closed. The rendering puts the original beside the reading. **Because postings vanish.** `JobPosting` has both `closes_at` and `closed_at`: the model already knows an advert dies. When it does, the URL is a 404 and the parsed JSON is the only surviving account of what the job was — no salary line as written, no list of requirements, nothing to check a later dispute against. Postulo already believes in freezing what was true at a moment; that is what snapshot-on-send does for a document that has left the building. A capture is the same promise pointing the other way. **Because it is the difference between a record and evidence.** The report CSV exists to be handed to an employment office. "I applied to this on this date" is a stronger statement when the advert is attached than when it is a row someone typed. **Because a parser improves and old captures do not.** This is the one that compounds. Today the source plugin that matched runs once, at fetch time, and its reading is all that survives. Keep the source and a capture can be **read again** — by a better parser, by a plugin that did not exist when it was taken, by a person fixing a field by hand months later. Every parser fix currently helps future captures only. This is what "post-processing" means in practice and it is worth more than the archival argument. ## Privacy: two switches, not one The two artefacts leak differently and must be refusable separately. - A **rendering** is a picture. If it came from the extension it is a picture of a page as *you* were seeing it, which can include your own name in the corner, a signed-in session, a personalised salary band, a recruiter's direct line. - The **source** is worse in a quieter way: tracking parameters, embedded identifiers, and whatever the page addressed to you personally. So: keep both, keep one, keep neither. Where the switch lives is already decided by precedent — `CaptureView` is a `PolicyView` over the site policy, with `pinned_fields` for what the environment has fixed, exactly as `capture_ignore_robots` works. An operator's *no* has to be final; a user's preference may only narrow it further, never widen it. **Default off**, and say why in the settings page the way the robots switch does. Keeping a copy of somebody else's page is a decision an operator should make on purpose. ## The hard parts **The server cannot draw most of these pages, and often cannot see them at all.** WeasyPrint is the default renderer and runs no JavaScript, so it would produce a faithful rendering of an empty shell. A real full-page rendering needs the Chromium backend — which exists, `ChromiumBackend` over Playwright in `documents/pdf.py`, with a `session()` that keeps one browser open across documents — but it is optional, and an instance that has never installed Playwright must degrade to "source only" rather than to a broken picture. Worse: #194's finding was that large employers sit behind bot protection that refuses anything that is not a browser, which is exactly why the API accepts `html` from the extension. For those pages a server-side render is a photograph of a login wall. **The extension is the right renderer for the hard cases** — it is the only thing holding the session — so `postulo-chromium` and `postulo-firefox` should be able to send a rendering and the source alongside the `html` they already send, and `CaptureIn` should accept them. That makes this a change in three repositories, which is worth knowing before it is scheduled. **Never serve it back as a page.** #218 is the precedent and the rule it set applies here with force: text from a stranger's page, stored and later drawn, is that stranger's code running as the person who opened it. An archived source must leave Postulo as a download with a content type that is not HTML, or be shown inside a sandboxed frame with a policy of its own. It must never be rendered on Postulo's own origin. Whatever is built here needs a test that says so. **Size, and what an instance is agreeing to.** The fetch cap is 2 MB of source (`fetching.MAX_BYTES`); the extension path has no such cap because it never had to. A full-page rendering of a long advert is larger than either. This needs its own limit, a retention rule — captures that were never confirmed are the obvious candidates for expiry — and it has to show up wherever the instance reports what it is storing. The files are per-owner and go through the private media path, like every other uploaded file. **Format.** A PDF from the existing Chromium backend is one searchable file, printable, and needs no new machinery. An image is more faithful and much larger. Single-file HTML is re-viewable and is the thing the paragraph above says not to render. PDF for the rendering and the source stored as compressed text is the proposal; it is worth arguing with. **robots.txt.** Postulo honours it when fetching. Keeping a copy of a page it was allowed to fetch, for the person who asked for it, does not seem to raise a second question — but the settings page should say what it does rather than leave a reader to wonder.
tiagoagueda added this to the 0.5.0 milestone 2026-09-17 14:57:07 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#256
No description provided.