The sources read the first posting on the page, one spelling of the standard, and the whole page as a description #176
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set.
Reference
Postulo/postulo#176
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Captures come back wrong often enough to be worth measuring, and until now there was
nothing to measure with: the sources were covered by hand-written objects, never by a page.
Reading four real adverts found four separate faults, none of which any existing test could
have caught.
What is wrong
A search page returns the wrong advert.
SchemaOrgSource._find_postingtakes the firstJobPostingit finds. A results page carries one per hit, so a capture from one brings backwhichever the board listed first. Nothing correlates the posting with the address being
looked at, even though every posting states its own
url,@idoridentifier.A posting inside an
ItemListis invisible._flattenfollows@graphand nothingelse. Several large aggregators hang their postings off
itemListElement→item, twolevels down, and those pages read as having no structured data at all.
Microdata and RDFa are not read. Only
<script type="ld+json">. Older recruitmentsystems and a good many public-sector boards publish
itempropmarkup and no script, andevery one of them falls through to the fallback and comes back with the page's
<title>asthe job title. It is the same schema.org vocabulary, in the spelling the standard also
defines.
Fields are read from one place when boards write them in two. Checked against live
pages:
jobLocationonly, so a board naming the office onhiringOrganization.addressloses thelocation entirely — We Work Remotely does this on every advert.
baseSalaryonly, neverestimatedSalary.jobLocationTypeonly, neverapplicantLocationRequirements, which schema.org defines aswhere an applicant may be for a job done remotely and is therefore a plain statement of
remote.
jobLocation, where a posting may name several offices.USD 0–0 YEARplaceholder into every advert, so captures arrive saying "USD, per year" and nothing else.
The fallback reads the whole page.
html_to_text(html)on the body: navigation, cookiebanner, "similar jobs", footer, the application form's buttons. On a Greenhouse advert the
description began "AI Engineer / Remote, Bangalore / Apply / …". Greenhouse publishes no
structured data at all, so this is not an edge case — it is one of the most used systems
there is.
What this changes
nothing does.
itemListElement,mainEntityanditemas well as@graph.Not a third source —
schema.orgreads the standard in whichever spelling a boardused, because it is one vocabulary and a
page-metadataresult is what it replaces.<main>/<article>, dropnav,header,footer,aside, anything with a chromerole, and form controls. Landmarksonly — never a guess at a class name, so a board that marks nothing is read exactly as
before.
<form>is deliberately not dropped: whole pages are still served wrapped inone, and dropping those would empty the description.
And a corpus, which is the actual deliverable
tests/fixtures/postings/— whole pages with the expected reading beside each, run bytests/test_source_corpus.py. Six to begin with: We Work Remotely, Greenhouse, a searchpage, an
ItemList, microdata and RDFa. Every quirk in them was taken from a live page andeach fixture says which; the prose is written rather than saved, because a saved advert is
somebody else's copyright and goes stale the week the posting closes.
The same corpus is in the extension, and that is the point.
postulo-chromiumreadsthese pages through a real DOM and is held to the same expected values, so a page read in
somebody's browser must come out identical to that page read here when it is sent. The two
readers were verified field for field against all six, and against two live pages.
To make that comparable,
htmlutilnow assembles the tag stream into a small tree beforewalking it — still the standard library's parser, still no C extension. Microdata, RDFa and
"is this paragraph inside the navigation" are all questions about nesting, and a streaming
parser cannot answer them without three separate ad-hoc state machines that would have no
hope of matching the DOM the extension walks.
What this does not fix
A board publishing nothing structured is still read by the fallback, and the fallback still
cannot know a company name or a location that the page never marked up. Greenhouse states
both — the company in
<title>as "… at GitLab", the location inog:description— inplaces no standard says to look. Reading them means either per-board rules or guessing at a
title convention, and both are arguments against the thing this project deliberately does
not do. Worth its own issue and its own decision, not a quiet addition here.
Board recipes, which is the part this issue said it was not doing
The open question at the bottom of this issue — what to do about a board publishing nothing
a standard can read — is settled. A recipe per board,
plugins/builtin/boards/, behind athird source called
boardthat runs first and letsschema.organdpage-metadatafill whatever it left empty.
That ordering is the whole design. A recipe is a supplement, never a replacement: a board
that starts publishing JSON-LD improves without its recipe being touched, and a recipe that
rots because a board redesigned costs the fields it used to fill rather than the capture. A
recipe that cannot even name the job is treated as not having recognised the page, and the
standards get it whole.
boardonly claims a host some recipe names. Every other page is read exactly as before —weworkremotely.comstill reads asschema.org, which is the point.LinkedIn, the case that prompted it
Before: title was "Spectrum Dynamics Medical hiring Medical Physicist (EU- Remote) in France
| LinkedIn", company empty, location empty, and an 11,400-character description that opened
with the topcard twice, "188 applicants", and closed with two thousand characters of
"Chemist jobs / 664 open jobs".
After: title "Medical Physicist (EU- Remote)", company "Spectrum Dynamics Medical", location
"France", employment type
full_time, and a 4,245-character description that is the advert.Two traps the recipe is written against, both real and both silent:
topcard__flavor--bulletis on the place and on the applicant count. Withoutexcluding the
--metadatathe count also carries, every LinkedIn capture would record itslocation as "188 applicants".
job__titleon Greenhouse wraps the heading andjob__location, so reading thecontainer gives "AI Engineer\n\nRemote, Bangalore" as a job title.
What a recipe is allowed to state
Only what it is sure of. LinkedIn shows its date as "2 months ago" — relative, in the
reader's language — so no date is stated at all. Turning that into a date means parsing
39 languages' worth of "month" and guessing what it was relative to; a date somebody can see
on the page and type is better than one Postulo invented.
The employment type is read, and the way it is read is the pattern for this sort of thing.
LinkedIn puts it in a criteria list under a heading in the reader's language. Rather than
keep a table of that heading in 39 languages, every criteria value is offered to Postulo's
own vocabulary and the one it recognises wins — "Full-time" is read, "Mid-Senior level" and
"Health Care Provider" are not, and a page in a language whose words Postulo does not know
leaves the field for a person.
EMPLOYMENT_TYPESmoved tobuiltin/vocabulary.pyso arecipe can use it without a recipe ever handing Postulo a value it would refuse.
Two generic fixes that came out of the same page, and help every board
The page's own
<h1>. Where a page gives exactly one, outside its furniture, that is abetter job title than
og:title— which carries whatever else a board wants a link to read.Checked against three live adverts: LinkedIn, Greenhouse and We Work Remotely each have
exactly one
<h1>and it is exactly the job title, and a search-results page has none, sothe rule cannot misfire there.
Link density. "Similar searches", "people also viewed", "explore top content": every
large board ends an advert with thousands of characters of them and marks them with a class
name and nothing else — LinkedIn's are plain
<section>s, so no landmark rule reaches them.What reaches them is what they are: a block whose text is 70% or more inside links, with at
least six of them, above a floor that keeps it off a short run of links in a real paragraph.
A measurement rather than a guess at markup, and the same one every reading-mode extractor
makes.
Boards with no recipe yet
Asked for: Lever, Ashby, Indeed, Glassdoor, EURES, Arbeitsagentur, StepStone, leboncoin.
None of them was written, deliberately. Indeed returns a security check to anything that
is not a browser, Ashby and leboncoin and EURES render their adverts in script, and the rest
had no reachable advert to read. Selectors written against markup nobody has read are the
exact fault this issue exists to fix — they fail silently, or worse, quietly record the wrong
field.
Each needs one saved advert page to be written against. A capture through the extension is
the easiest way to get one, since it sees the rendered page rather than what a server hands a
script.
Held to a corpus
tests/fixtures/postings/is now eight pages,linkedinandgreenhouseamong them, andthe browser extension is held to the same eight. Every field of all eight was verified
identical between the two readers, and against the live LinkedIn, Greenhouse and We Work
Remotely pages.
Landed on
mainasbd6b2b172— Read the posting a page is showing, however the board wrote it — which saidRefsrather thanClosesand so left this open. It holds everything the issue and the comment above list: the posting matched to the page's own address,ItemList/mainEntity/itemfollowed, microdata and RDFa read into the JSON-LD shape, the field gaps filled, a currency kept only beside an amount, the fallback reading landmarks, the board recipes underplugins/builtin/boards/with the<h1>rule and link density, andtests/fixtures/postings/held bytests/test_source_corpus.py. Closing with it.