Sites that publish no standard: learn from the corrections, not a recipe per board #267

Open
opened 2026-09-17 19:30:35 +00:00 by tiagoagueda · 0 comments
Owner

plugins/builtin reads a page in three tiers: a board recipe, then schema.org in its three
spellings, then page metadata. The first two are good and neither helps the case this issue
is about. The third is honest about itself:

page-metadata — When there is no structured data, take the title the page declares and
the readable text, and let the person capturing fix the rest. Deliberately unambitious: it
saves typing, and it never pretends to know more than it does.

That is the floor, and most of the web a job seeker actually visits sits on it: a public
sector board in Portugal, a municipal careers page, a fifteen-person company's WordPress
site, a recruiter's bespoke portal. None of them publish JSON-LD and none of them ever will.

Why more recipes is not the answer

There are two recipes today, LinkedIn and Greenhouse, and the module says why they exist:
boards that "publish nothing a standard can read". They are worth their maintenance because
each covers an enormous number of postings.

The long tail is the opposite trade. A recipe for one municipality's careers page costs the
same to write and to repair as the LinkedIn one and covers a handful of adverts, for a
handful of people, until the site is redesigned. There is no number of recipes that finishes
this job, and #241's source plugins — all of them platforms with public JSON — do not touch
it either.

The asset nobody is using: the corrections

Every capture is already corrected by a person. Capture.data holds what was parsed, the
capture stays PENDING until somebody looks at it, and the review screen is where they fix
the title, the company, the place and the salary. On the API path the same thing happens
explicitly — api.py merges corrections over the parsed data before validating.

That difference — what we guessed, and what the person changed it to — is a labelled example
of how to read that site, produced by the person who was going to have to look at the page
anyway. It is thrown away every time.

Proposal: remembered field hints, per person, per host

When somebody corrects a field on review and the corrected value can be located in the
captured page, remember where it was found, for that owner and that host. The next capture
from the same host tries the remembered places before falling through to page metadata.

  • The extension is the natural place to learn, because it has a real DOM
    (postulo-chromium/src/lib/parse.js already reads the page there). A selector learned in
    the browser is reusable by the server-side fetch too — the HTML is the same HTML — so one
    correction improves both paths.
  • Per owner, never shared. A hint is the person's own data, scoped by for_user() like
    everything else. No registry, no corpus, nothing reported anywhere. A shared recipe
    database would be a different product with a different privacy posture, and this project
    should not have one.
  • It improves the sites you personally use, which is exactly the right target: somebody
    applying through one national board and three local employers corrects each of them once.

The guards, which matter more than the mechanism

  • Never invent. The module's existing rule holds without exception: a hint that finds
    nothing leaves the field empty. A remembered place is a better guess, not a promotion to
    fact — everything still goes to review.
  • Decay. Two failures in a row and the hint is dropped. A site redesign should cost one
    poor capture, not a permanently wrong answer that is harder to notice than an empty field.
  • Say so. The review screen should mark a field that came from a remembered hint, so a
    wrong one is visible and correctable rather than quietly authoritative. This is the
    difference between a feature and a trap.
  • Only learn from what was touched. A field the person left alone is not evidence that it
    was right — they may not have checked it.

The cheaper half, worth doing first and on its own

The page-metadata tier can be improved with no learning and no state at all, and it helps
every site immediately:

  • read OpenGraph and Twitter card properties, not only <title>;
  • split a <title> on the separators boards actually use — a great many are literally
    Job Title - Company - Location;
  • per-locale patterns for a salary and a closing date, which is the one place Postulo's
    existing locale knowledge is an advantage nobody else reading these pages has.

None of that needs the proposal above, and it should probably land first so that the hints
are learned on top of a better floor.

Relation to #241's local-model idea

#241 item 5 proposes a local model for extracting postings. That is the probabilistic answer
to this same problem, and the two are complementary rather than competing: remembered hints
are deterministic, explainable, free, need no model and no hardware, and they get the
specific sites one person uses. A model generalises to sites never seen before and costs
much more to run.

If both are ever built, hints should run first and the model should fill what they leave —
the same order the three tiers already use, for the same reason.

`plugins/builtin` reads a page in three tiers: a board recipe, then schema.org in its three spellings, then page metadata. The first two are good and neither helps the case this issue is about. The third is honest about itself: > **page-metadata** — When there is no structured data, take the title the page declares and > the readable text, and let the person capturing fix the rest. Deliberately unambitious: it > saves typing, and it never pretends to know more than it does. That is the floor, and most of the web a job seeker actually visits sits on it: a public sector board in Portugal, a municipal careers page, a fifteen-person company's WordPress site, a recruiter's bespoke portal. None of them publish JSON-LD and none of them ever will. ## Why more recipes is not the answer There are two recipes today, LinkedIn and Greenhouse, and the module says why they exist: boards that "publish nothing a standard can read". They are worth their maintenance because each covers an enormous number of postings. The long tail is the opposite trade. A recipe for one municipality's careers page costs the same to write and to repair as the LinkedIn one and covers a handful of adverts, for a handful of people, until the site is redesigned. There is no number of recipes that finishes this job, and #241's source plugins — all of them platforms with public JSON — do not touch it either. ## The asset nobody is using: the corrections **Every capture is already corrected by a person.** `Capture.data` holds what was parsed, the capture stays `PENDING` until somebody looks at it, and the review screen is where they fix the title, the company, the place and the salary. On the API path the same thing happens explicitly — `api.py` merges `corrections` over the parsed data before validating. That difference — what we guessed, and what the person changed it to — is a labelled example of how to read that site, produced by the person who was going to have to look at the page anyway. It is thrown away every time. ## Proposal: remembered field hints, per person, per host When somebody corrects a field on review and the corrected value can be located in the captured page, remember where it was found, for that owner and that host. The next capture from the same host tries the remembered places before falling through to page metadata. - **The extension is the natural place to learn**, because it has a real DOM (`postulo-chromium/src/lib/parse.js` already reads the page there). A selector learned in the browser is reusable by the server-side fetch too — the HTML is the same HTML — so one correction improves both paths. - **Per owner, never shared.** A hint is the person's own data, scoped by `for_user()` like everything else. No registry, no corpus, nothing reported anywhere. A shared recipe database would be a different product with a different privacy posture, and this project should not have one. - **It improves the sites you personally use**, which is exactly the right target: somebody applying through one national board and three local employers corrects each of them once. ### The guards, which matter more than the mechanism - **Never invent.** The module's existing rule holds without exception: a hint that finds nothing leaves the field empty. A remembered place is a better guess, not a promotion to fact — everything still goes to review. - **Decay.** Two failures in a row and the hint is dropped. A site redesign should cost one poor capture, not a permanently wrong answer that is harder to notice than an empty field. - **Say so.** The review screen should mark a field that came from a remembered hint, so a wrong one is visible and correctable rather than quietly authoritative. This is the difference between a feature and a trap. - **Only learn from what was touched.** A field the person left alone is not evidence that it was right — they may not have checked it. ## The cheaper half, worth doing first and on its own The page-metadata tier can be improved with no learning and no state at all, and it helps every site immediately: - read OpenGraph and Twitter card properties, not only `<title>`; - split a `<title>` on the separators boards actually use — a great many are literally `Job Title - Company - Location`; - per-locale patterns for a salary and a closing date, which is the one place Postulo's existing locale knowledge is an advantage nobody else reading these pages has. None of that needs the proposal above, and it should probably land first so that the hints are learned on top of a better floor. ## Relation to #241's local-model idea #241 item 5 proposes a local model for extracting postings. That is the probabilistic answer to this same problem, and the two are complementary rather than competing: remembered hints are deterministic, explainable, free, need no model and no hardware, and they get the specific sites one person uses. A model generalises to sites never seen before and costs much more to run. If both are ever built, hints should run first and the model should fill what they leave — the same order the three tiers already use, for the same reason.
tiagoagueda added this to the 0.5.0 milestone 2026-09-17 19:30:35 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#267
No description provided.