A capture should be able to keep the page: a rendering of the whole thing, and the source #256
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#256
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
A capture reads a page, keeps what it understood, and throws the page away.
fetching.pyreturns a
FetchedPage,parse_pageturns it intoJobPostingData, andCapture.datastores that JSON — the markup is never written anywhere. The same is true ofthe source a browser extension hands over through the API: it is parsed on arrival and
dropped.
A capture should be able to keep two things beside the parsed fields, both optional:
everything below the fold.
Why it earns the disk
Because the parse is a guess and the page is the evidence.
Capture's own docstringsays so: "Parsing somebody else's markup is guesswork often enough that turning the result
straight into a record would put invented job titles into the one place they must not be."
Everything captured waits for a person to look at it — and what they are checking it against
is a browser tab they may have already closed. The rendering puts the original beside the
reading.
Because postings vanish.
JobPostinghas bothcloses_atandclosed_at: the modelalready knows an advert dies. When it does, the URL is a 404 and the parsed JSON is the only
surviving account of what the job was — no salary line as written, no list of requirements,
nothing to check a later dispute against. Postulo already believes in freezing what was true
at a moment; that is what snapshot-on-send does for a document that has left the building. A
capture is the same promise pointing the other way.
Because it is the difference between a record and evidence. The report CSV exists to be
handed to an employment office. "I applied to this on this date" is a stronger statement when
the advert is attached than when it is a row someone typed.
Because a parser improves and old captures do not. This is the one that compounds. Today
the source plugin that matched runs once, at fetch time, and its reading is all that survives.
Keep the source and a capture can be read again — by a better parser, by a plugin that did
not exist when it was taken, by a person fixing a field by hand months later. Every parser fix
currently helps future captures only. This is what "post-processing" means in practice and it
is worth more than the archival argument.
Privacy: two switches, not one
The two artefacts leak differently and must be refusable separately.
you were seeing it, which can include your own name in the corner, a signed-in session, a
personalised salary band, a recruiter's direct line.
whatever the page addressed to you personally.
So: keep both, keep one, keep neither. Where the switch lives is already decided by
precedent —
CaptureViewis aPolicyViewover the site policy, withpinned_fieldsforwhat the environment has fixed, exactly as
capture_ignore_robotsworks. An operator's nohas to be final; a user's preference may only narrow it further, never widen it.
Default off, and say why in the settings page the way the robots switch does. Keeping a
copy of somebody else's page is a decision an operator should make on purpose.
The hard parts
The server cannot draw most of these pages, and often cannot see them at all. WeasyPrint
is the default renderer and runs no JavaScript, so it would produce a faithful rendering of
an empty shell. A real full-page rendering needs the Chromium backend — which exists,
ChromiumBackendover Playwright indocuments/pdf.py, with asession()that keeps onebrowser open across documents — but it is optional, and an instance that has never installed
Playwright must degrade to "source only" rather than to a broken picture.
Worse: #194's finding was that large employers sit behind bot protection that refuses
anything that is not a browser, which is exactly why the API accepts
htmlfrom theextension. For those pages a server-side render is a photograph of a login wall. The
extension is the right renderer for the hard cases — it is the only thing holding the
session — so
postulo-chromiumandpostulo-firefoxshould be able to send a rendering andthe source alongside the
htmlthey already send, andCaptureInshould accept them. Thatmakes this a change in three repositories, which is worth knowing before it is scheduled.
Never serve it back as a page. #218 is the precedent and the rule it set applies here
with force: text from a stranger's page, stored and later drawn, is that stranger's code
running as the person who opened it. An archived source must leave Postulo as a download
with a content type that is not HTML, or be shown inside a sandboxed frame with a policy of
its own. It must never be rendered on Postulo's own origin. Whatever is built here needs a
test that says so.
Size, and what an instance is agreeing to. The fetch cap is 2 MB of source
(
fetching.MAX_BYTES); the extension path has no such cap because it never had to. Afull-page rendering of a long advert is larger than either. This needs its own limit, a
retention rule — captures that were never confirmed are the obvious candidates for expiry —
and it has to show up wherever the instance reports what it is storing. The files are
per-owner and go through the private media path, like every other uploaded file.
Format. A PDF from the existing Chromium backend is one searchable file, printable, and
needs no new machinery. An image is more faithful and much larger. Single-file HTML is
re-viewable and is the thing the paragraph above says not to render. PDF for the rendering
and the source stored as compressed text is the proposal; it is worth arguing with.
robots.txt. Postulo honours it when fetching. Keeping a copy of a page it was allowed to
fetch, for the person who asked for it, does not seem to raise a second question — but the
settings page should say what it does rather than leave a reader to wonder.