The sources read the first posting on the page, one spelling of the standard, and the whole page as a description #176

Closed
opened 2026-09-12 07:52:01 +00:00 by tiagoagueda · 2 comments
Owner

Captures come back wrong often enough to be worth measuring, and until now there was
nothing to measure with: the sources were covered by hand-written objects, never by a page.
Reading four real adverts found four separate faults, none of which any existing test could
have caught.

What is wrong

A search page returns the wrong advert. SchemaOrgSource._find_posting takes the first
JobPosting it finds. A results page carries one per hit, so a capture from one brings back
whichever the board listed first. Nothing correlates the posting with the address being
looked at, even though every posting states its own url, @id or identifier.

A posting inside an ItemList is invisible. _flatten follows @graph and nothing
else. Several large aggregators hang their postings off itemListElement → item, two
levels down, and those pages read as having no structured data at all.

Microdata and RDFa are not read. Only <script type="ld+json">. Older recruitment
systems and a good many public-sector boards publish itemprop markup and no script, and
every one of them falls through to the fallback and comes back with the page's <title> as
the job title. It is the same schema.org vocabulary, in the spelling the standard also
defines.

Fields are read from one place when boards write them in two. Checked against live
pages:

  • jobLocation only, so a board naming the office on hiringOrganization.address loses the
    location entirely — We Work Remotely does this on every advert.
  • baseSalary only, never estimatedSalary.
  • jobLocationType only, never applicantLocationRequirements, which schema.org defines as
    where an applicant may be for a job done remotely and is therefore a plain statement of
    remote.
  • One jobLocation, where a posting may name several offices.
  • A currency and a period are kept with no amount behind them. Boards write a USD 0–0 YEAR
    placeholder into every advert, so captures arrive saying "USD, per year" and nothing else.

The fallback reads the whole page. html_to_text(html) on the body: navigation, cookie
banner, "similar jobs", footer, the application form's buttons. On a Greenhouse advert the
description began "AI Engineer / Remote, Bangalore / Apply / …". Greenhouse publishes no
structured data at all, so this is not an edge case — it is one of the most used systems
there is.

What this changes

  • Pick the posting that names the page's own address; fall back to the first only when
    nothing does.
  • Follow itemListElement, mainEntity and item as well as @graph.
  • Read microdata and RDFa into the same shape JSON-LD gives, so one reader serves all three.
    Not a third source — schema.org reads the standard in whichever spelling a board
    used, because it is one vocabulary and a page-metadata result is what it replaces.
  • Fill the field gaps above, and keep a currency only alongside an amount.
  • The fallback reads the page's own landmarks: prefer <main>/<article>, drop nav,
    header, footer, aside, anything with a chrome role, and form controls. Landmarks
    only — never a guess at a class name, so a board that marks nothing is read exactly as
    before. <form> is deliberately not dropped: whole pages are still served wrapped in
    one, and dropping those would empty the description.

And a corpus, which is the actual deliverable

tests/fixtures/postings/ — whole pages with the expected reading beside each, run by
tests/test_source_corpus.py. Six to begin with: We Work Remotely, Greenhouse, a search
page, an ItemList, microdata and RDFa. Every quirk in them was taken from a live page and
each fixture says which; the prose is written rather than saved, because a saved advert is
somebody else's copyright and goes stale the week the posting closes.

The same corpus is in the extension, and that is the point. postulo-chromium reads
these pages through a real DOM and is held to the same expected values, so a page read in
somebody's browser must come out identical to that page read here when it is sent. The two
readers were verified field for field against all six, and against two live pages.

To make that comparable, htmlutil now assembles the tag stream into a small tree before
walking it — still the standard library's parser, still no C extension. Microdata, RDFa and
"is this paragraph inside the navigation" are all questions about nesting, and a streaming
parser cannot answer them without three separate ad-hoc state machines that would have no
hope of matching the DOM the extension walks.

What this does not fix

A board publishing nothing structured is still read by the fallback, and the fallback still
cannot know a company name or a location that the page never marked up. Greenhouse states
both — the company in <title> as "… at GitLab", the location in og:description — in
places no standard says to look. Reading them means either per-board rules or guessing at a
title convention, and both are arguments against the thing this project deliberately does
not do. Worth its own issue and its own decision, not a quiet addition here.

Captures come back wrong often enough to be worth measuring, and until now there was nothing to measure with: the sources were covered by hand-written objects, never by a page. Reading four real adverts found four separate faults, none of which any existing test could have caught. ### What is wrong **A search page returns the wrong advert.** `SchemaOrgSource._find_posting` takes the first `JobPosting` it finds. A results page carries one per hit, so a capture from one brings back whichever the board listed first. Nothing correlates the posting with the address being looked at, even though every posting states its own `url`, `@id` or `identifier`. **A posting inside an `ItemList` is invisible.** `_flatten` follows `@graph` and nothing else. Several large aggregators hang their postings off `itemListElement` → `item`, two levels down, and those pages read as having no structured data at all. **Microdata and RDFa are not read.** Only `<script type="ld+json">`. Older recruitment systems and a good many public-sector boards publish `itemprop` markup and no script, and every one of them falls through to the fallback and comes back with the page's `<title>` as the job title. It is the same schema.org vocabulary, in the spelling the standard also defines. **Fields are read from one place when boards write them in two.** Checked against live pages: - `jobLocation` only, so a board naming the office on `hiringOrganization.address` loses the location entirely — We Work Remotely does this on every advert. - `baseSalary` only, never `estimatedSalary`. - `jobLocationType` only, never `applicantLocationRequirements`, which schema.org defines as where an applicant may be *for a job done remotely* and is therefore a plain statement of remote. - One `jobLocation`, where a posting may name several offices. - A currency and a period are kept with no amount behind them. Boards write a `USD 0–0 YEAR` placeholder into every advert, so captures arrive saying "USD, per year" and nothing else. **The fallback reads the whole page.** `html_to_text(html)` on the body: navigation, cookie banner, "similar jobs", footer, the application form's buttons. On a Greenhouse advert the description began "AI Engineer / Remote, Bangalore / Apply / …". Greenhouse publishes no structured data at all, so this is not an edge case — it is one of the most used systems there is. ### What this changes - Pick the posting that names the page's own address; fall back to the first only when nothing does. - Follow `itemListElement`, `mainEntity` and `item` as well as `@graph`. - Read microdata and RDFa into the same shape JSON-LD gives, so one reader serves all three. **Not a third source** — `schema.org` reads the standard in whichever spelling a board used, because it is one vocabulary and a `page-metadata` result is what it replaces. - Fill the field gaps above, and keep a currency only alongside an amount. - The fallback reads the page's own landmarks: prefer `<main>`/`<article>`, drop `nav`, `header`, `footer`, `aside`, anything with a chrome `role`, and form controls. Landmarks only — never a guess at a class name, so a board that marks nothing is read exactly as before. `<form>` is deliberately *not* dropped: whole pages are still served wrapped in one, and dropping those would empty the description. ### And a corpus, which is the actual deliverable `tests/fixtures/postings/` — whole pages with the expected reading beside each, run by `tests/test_source_corpus.py`. Six to begin with: We Work Remotely, Greenhouse, a search page, an `ItemList`, microdata and RDFa. Every quirk in them was taken from a live page and each fixture says which; the prose is written rather than saved, because a saved advert is somebody else's copyright and goes stale the week the posting closes. **The same corpus is in the extension**, and that is the point. `postulo-chromium` reads these pages through a real DOM and is held to the same expected values, so a page read in somebody's browser must come out identical to that page read here when it is sent. The two readers were verified field for field against all six, and against two live pages. To make that comparable, `htmlutil` now assembles the tag stream into a small tree before walking it — still the standard library's parser, still no C extension. Microdata, RDFa and "is this paragraph inside the navigation" are all questions about nesting, and a streaming parser cannot answer them without three separate ad-hoc state machines that would have no hope of matching the DOM the extension walks. ### What this does not fix A board publishing nothing structured is still read by the fallback, and the fallback still cannot know a company name or a location that the page never marked up. Greenhouse states both — the company in `<title>` as "… at GitLab", the location in `og:description` — in places no standard says to look. Reading them means either per-board rules or guessing at a title convention, and both are arguments against the thing this project deliberately does not do. Worth its own issue and its own decision, not a quiet addition here.
Author
Owner

Board recipes, which is the part this issue said it was not doing

The open question at the bottom of this issue — what to do about a board publishing nothing
a standard can read — is settled. A recipe per board, plugins/builtin/boards/, behind a
third source called board that runs first and lets schema.org and page-metadata
fill whatever it left empty.

That ordering is the whole design. A recipe is a supplement, never a replacement: a board
that starts publishing JSON-LD improves without its recipe being touched, and a recipe that
rots because a board redesigned costs the fields it used to fill rather than the capture. A
recipe that cannot even name the job is treated as not having recognised the page, and the
standards get it whole.

board only claims a host some recipe names. Every other page is read exactly as before —
weworkremotely.com still reads as schema.org, which is the point.

LinkedIn, the case that prompted it

https://www.linkedin.com/jobs/view/4435670736/

Before: title was "Spectrum Dynamics Medical hiring Medical Physicist (EU- Remote) in France
| LinkedIn", company empty, location empty, and an 11,400-character description that opened
with the topcard twice, "188 applicants", and closed with two thousand characters of
"Chemist jobs / 664 open jobs".

After: title "Medical Physicist (EU- Remote)", company "Spectrum Dynamics Medical", location
"France", employment type full_time, and a 4,245-character description that is the advert.

Two traps the recipe is written against, both real and both silent:

  • topcard__flavor--bullet is on the place and on the applicant count. Without
    excluding the --metadata the count also carries, every LinkedIn capture would record its
    location as "188 applicants".
  • job__title on Greenhouse wraps the heading and job__location, so reading the
    container gives "AI Engineer\n\nRemote, Bangalore" as a job title.

What a recipe is allowed to state

Only what it is sure of. LinkedIn shows its date as "2 months ago" — relative, in the
reader's language — so no date is stated at all. Turning that into a date means parsing
39 languages' worth of "month" and guessing what it was relative to; a date somebody can see
on the page and type is better than one Postulo invented.

The employment type is read, and the way it is read is the pattern for this sort of thing.
LinkedIn puts it in a criteria list under a heading in the reader's language. Rather than
keep a table of that heading in 39 languages, every criteria value is offered to Postulo's
own vocabulary and the one it recognises wins — "Full-time" is read, "Mid-Senior level" and
"Health Care Provider" are not, and a page in a language whose words Postulo does not know
leaves the field for a person. EMPLOYMENT_TYPES moved to builtin/vocabulary.py so a
recipe can use it without a recipe ever handing Postulo a value it would refuse.

Two generic fixes that came out of the same page, and help every board

The page's own <h1>. Where a page gives exactly one, outside its furniture, that is a
better job title than og:title — which carries whatever else a board wants a link to read.
Checked against three live adverts: LinkedIn, Greenhouse and We Work Remotely each have
exactly one <h1> and it is exactly the job title, and a search-results page has none, so
the rule cannot misfire there.

Link density. "Similar searches", "people also viewed", "explore top content": every
large board ends an advert with thousands of characters of them and marks them with a class
name and nothing else — LinkedIn's are plain <section>s, so no landmark rule reaches them.
What reaches them is what they are: a block whose text is 70% or more inside links, with at
least six of them, above a floor that keeps it off a short run of links in a real paragraph.
A measurement rather than a guess at markup, and the same one every reading-mode extractor
makes.

Boards with no recipe yet

Asked for: Lever, Ashby, Indeed, Glassdoor, EURES, Arbeitsagentur, StepStone, leboncoin.
None of them was written, deliberately. Indeed returns a security check to anything that
is not a browser, Ashby and leboncoin and EURES render their adverts in script, and the rest
had no reachable advert to read. Selectors written against markup nobody has read are the
exact fault this issue exists to fix — they fail silently, or worse, quietly record the wrong
field.

Each needs one saved advert page to be written against. A capture through the extension is
the easiest way to get one, since it sees the rendered page rather than what a server hands a
script.

Held to a corpus

tests/fixtures/postings/ is now eight pages, linkedin and greenhouse among them, and
the browser extension is held to the same eight. Every field of all eight was verified
identical between the two readers, and against the live LinkedIn, Greenhouse and We Work
Remotely pages.

### Board recipes, which is the part this issue said it was not doing The open question at the bottom of this issue — what to do about a board publishing nothing a standard can read — is settled. A recipe per board, `plugins/builtin/boards/`, behind a third source called `board` that runs **first** and lets `schema.org` and `page-metadata` fill whatever it left empty. That ordering is the whole design. A recipe is a supplement, never a replacement: a board that starts publishing JSON-LD improves without its recipe being touched, and a recipe that rots because a board redesigned costs the fields it used to fill rather than the capture. A recipe that cannot even name the job is treated as not having recognised the page, and the standards get it whole. `board` only claims a host some recipe names. Every other page is read exactly as before — `weworkremotely.com` still reads as `schema.org`, which is the point. ### LinkedIn, the case that prompted it https://www.linkedin.com/jobs/view/4435670736/ Before: title was "Spectrum Dynamics Medical hiring Medical Physicist (EU- Remote) in France | LinkedIn", company empty, location empty, and an 11,400-character description that opened with the topcard twice, "188 applicants", and closed with two thousand characters of "Chemist jobs / 664 open jobs". After: title "Medical Physicist (EU- Remote)", company "Spectrum Dynamics Medical", location "France", employment type `full_time`, and a 4,245-character description that is the advert. Two traps the recipe is written against, both real and both silent: - `topcard__flavor--bullet` is on the place **and** on the applicant count. Without excluding the `--metadata` the count also carries, every LinkedIn capture would record its location as "188 applicants". - `job__title` on Greenhouse wraps the heading **and** `job__location`, so reading the container gives "AI Engineer\n\nRemote, Bangalore" as a job title. ### What a recipe is allowed to state Only what it is sure of. LinkedIn shows its date as "2 months ago" — relative, in the reader's language — so **no date is stated at all**. Turning that into a date means parsing 39 languages' worth of "month" and guessing what it was relative to; a date somebody can see on the page and type is better than one Postulo invented. The employment type is read, and the way it is read is the pattern for this sort of thing. LinkedIn puts it in a criteria list under a heading in the reader's language. Rather than keep a table of that heading in 39 languages, every criteria value is offered to Postulo's own vocabulary and the one it recognises wins — "Full-time" is read, "Mid-Senior level" and "Health Care Provider" are not, and a page in a language whose words Postulo does not know leaves the field for a person. `EMPLOYMENT_TYPES` moved to `builtin/vocabulary.py` so a recipe can use it without a recipe ever handing Postulo a value it would refuse. ### Two generic fixes that came out of the same page, and help every board **The page's own `<h1>`.** Where a page gives exactly one, outside its furniture, that is a better job title than `og:title` — which carries whatever else a board wants a link to read. Checked against three live adverts: LinkedIn, Greenhouse and We Work Remotely each have exactly one `<h1>` and it is exactly the job title, and a search-results page has none, so the rule cannot misfire there. **Link density.** "Similar searches", "people also viewed", "explore top content": every large board ends an advert with thousands of characters of them and marks them with a class name and nothing else — LinkedIn's are plain `<section>`s, so no landmark rule reaches them. What reaches them is what they are: a block whose text is 70% or more inside links, with at least six of them, above a floor that keeps it off a short run of links in a real paragraph. A measurement rather than a guess at markup, and the same one every reading-mode extractor makes. ### Boards with no recipe yet Asked for: Lever, Ashby, Indeed, Glassdoor, EURES, Arbeitsagentur, StepStone, leboncoin. **None of them was written, deliberately.** Indeed returns a security check to anything that is not a browser, Ashby and leboncoin and EURES render their adverts in script, and the rest had no reachable advert to read. Selectors written against markup nobody has read are the exact fault this issue exists to fix — they fail silently, or worse, quietly record the wrong field. Each needs one saved advert page to be written against. A capture through the extension is the easiest way to get one, since it sees the rendered page rather than what a server hands a script. ### Held to a corpus `tests/fixtures/postings/` is now eight pages, `linkedin` and `greenhouse` among them, and the browser extension is held to the same eight. Every field of all eight was verified identical between the two readers, and against the live LinkedIn, Greenhouse and We Work Remotely pages.
Author
Owner

Landed on main as bd6b2b172 — Read the posting a page is showing, however the board wrote it — which said Refs rather than Closes and so left this open. It holds everything the issue and the comment above list: the posting matched to the page's own address, ItemList/mainEntity/item followed, microdata and RDFa read into the JSON-LD shape, the field gaps filled, a currency kept only beside an amount, the fallback reading landmarks, the board recipes under plugins/builtin/boards/ with the <h1> rule and link density, and tests/fixtures/postings/ held by tests/test_source_corpus.py. Closing with it.

Landed on `main` as `bd6b2b172` — *Read the posting a page is showing, however the board wrote it* — which said `Refs` rather than `Closes` and so left this open. It holds everything the issue and the comment above list: the posting matched to the page's own address, `ItemList`/`mainEntity`/`item` followed, microdata and RDFa read into the JSON-LD shape, the field gaps filled, a currency kept only beside an amount, the fallback reading landmarks, the board recipes under `plugins/builtin/boards/` with the `<h1>` rule and link density, and `tests/fixtures/postings/` held by `tests/test_source_corpus.py`. Closing with it.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#176
No description provided.