Sites that publish no standard: learn from the corrections, not a recipe per board #267
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#267
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
plugins/builtinreads a page in three tiers: a board recipe, then schema.org in its threespellings, then page metadata. The first two are good and neither helps the case this issue
is about. The third is honest about itself:
That is the floor, and most of the web a job seeker actually visits sits on it: a public
sector board in Portugal, a municipal careers page, a fifteen-person company's WordPress
site, a recruiter's bespoke portal. None of them publish JSON-LD and none of them ever will.
Why more recipes is not the answer
There are two recipes today, LinkedIn and Greenhouse, and the module says why they exist:
boards that "publish nothing a standard can read". They are worth their maintenance because
each covers an enormous number of postings.
The long tail is the opposite trade. A recipe for one municipality's careers page costs the
same to write and to repair as the LinkedIn one and covers a handful of adverts, for a
handful of people, until the site is redesigned. There is no number of recipes that finishes
this job, and #241's source plugins — all of them platforms with public JSON — do not touch
it either.
The asset nobody is using: the corrections
Every capture is already corrected by a person.
Capture.dataholds what was parsed, thecapture stays
PENDINGuntil somebody looks at it, and the review screen is where they fixthe title, the company, the place and the salary. On the API path the same thing happens
explicitly —
api.pymergescorrectionsover the parsed data before validating.That difference — what we guessed, and what the person changed it to — is a labelled example
of how to read that site, produced by the person who was going to have to look at the page
anyway. It is thrown away every time.
Proposal: remembered field hints, per person, per host
When somebody corrects a field on review and the corrected value can be located in the
captured page, remember where it was found, for that owner and that host. The next capture
from the same host tries the remembered places before falling through to page metadata.
(
postulo-chromium/src/lib/parse.jsalready reads the page there). A selector learned inthe browser is reusable by the server-side fetch too — the HTML is the same HTML — so one
correction improves both paths.
for_user()likeeverything else. No registry, no corpus, nothing reported anywhere. A shared recipe
database would be a different product with a different privacy posture, and this project
should not have one.
applying through one national board and three local employers corrects each of them once.
The guards, which matter more than the mechanism
nothing leaves the field empty. A remembered place is a better guess, not a promotion to
fact — everything still goes to review.
poor capture, not a permanently wrong answer that is harder to notice than an empty field.
wrong one is visible and correctable rather than quietly authoritative. This is the
difference between a feature and a trap.
was right — they may not have checked it.
The cheaper half, worth doing first and on its own
The page-metadata tier can be improved with no learning and no state at all, and it helps
every site immediately:
<title>;<title>on the separators boards actually use — a great many are literallyJob Title - Company - Location;existing locale knowledge is an advantage nobody else reading these pages has.
None of that needs the proposal above, and it should probably land first so that the hints
are learned on top of a better floor.
Relation to #241's local-model idea
#241 item 5 proposes a local model for extracting postings. That is the probabilistic answer
to this same problem, and the two are complementary rather than competing: remembered hints
are deterministic, explainable, free, need no model and no hardware, and they get the
specific sites one person uses. A model generalises to sites never seen before and costs
much more to run.
If both are ever built, hints should run first and the model should fill what they leave —
the same order the three tiers already use, for the same reason.