Occupations and skills should be ESCO, for the reason industries are NACE #266

Open
opened 2026-09-17 19:30:34 +00:00 by tiagoagueda · 0 comments
Owner

A job title is free text everywhere in Postulo. So is a skill. That is correct for what a
person types and wrong for what the application can then do with it: Software Engineer,
Ingénieur logiciel and Engenheiro de software are one job and three strings, and nothing
in the code can tell.

For an application whose whole argument is a job search conducted across European languages,
that is the gap worth closing.

This project has already made this decision once

jobs/industries.py explains why the list of areas of activity is NACE rather than a list
assembled here:

Thirty-two names assembled by hand was not a classification […] the way to fix that is not
to keep adding names, because a hand-made list of two hundred would be two hundred names in
thirty-nine languages, carried by this project for ever (#140).

The translations are the real argument, more than the taxonomy. Eurostat publishes NACE
[…]

Every word of that transfers. ESCO is to occupations and skills what NACE is to
industries
: published by the European Commission, roughly 3,000 occupations and 14,000
skills, in every EU language, mapped to ISCO-08, and maintained by somebody else.

What it would buy

  • Cross-language matching. A listing captured in French and a CV written in Portuguese
    become comparable. Nothing else on the roadmap makes that possible.
  • A skills vocabulary for the CV that is not free text, which the résumé side needs and
    currently invents per entry.
  • Translations that arrive with the data, exactly as NACE's do — which is the argument
    #140 already accepted, and which matters more with each language milestone (#70, #71,
    #72).
  • A ladder to the public services. EURES and the national services already speak ESCO, so
    a source plugin reading one of them (#241) has somewhere to put what it reads.

The judgement to make, and it is the same one as NACE

ESCO's full occupation list is far too deep for a picker, exactly as NACE's 615 classes were
— industries.py chose divisions because they are "the level where the names still mean
something to the person reading them".

The analogue here is probably ISCO-08 unit groups: about 436 four-digit codes, which is
the same order of magnitude as the 87 NACE divisions relative to its classes, and the level
at which a name is still a job somebody recognises. Full ESCO occupations underneath as an
optional refinement, or not at all to begin with.

That choice should be made deliberately and written down in the module the way industries.py
wrote down the NACE one, because it is the decision everything else rests on.

Scope questions this issue has to answer

  • Where it attaches. A title on JobPosting, a skill on the résumé, or both — and
    whether the free-text field stays beside the code, as a person's own wording usually
    should.
  • Licence and distribution. ESCO is published under the EUPL; the data would ship in
    jobs/data/ beside nace-2.1.json with its own LICENCE.md, as NACE already does.
  • Size. NACE at division level is 87 rows. ESCO is a different order of magnitude, and
    what ships in the repository versus what is loaded on demand is a real decision rather than
    a detail.
  • Never a requirement. A person must be able to type a job title that is not in any
    classification. The code is for what Postulo can do across languages; the text is what the
    person meant.

Filed from the same conversation as #241's additions. Not a plugin: this is a core
vocabulary, in the same shape as #140.

A job title is free text everywhere in Postulo. So is a skill. That is correct for what a person types and wrong for what the application can then do with it: *Software Engineer*, *Ingénieur logiciel* and *Engenheiro de software* are one job and three strings, and nothing in the code can tell. For an application whose whole argument is a job search conducted across European languages, that is the gap worth closing. ## This project has already made this decision once `jobs/industries.py` explains why the list of areas of activity is NACE rather than a list assembled here: > Thirty-two names assembled by hand was not a classification […] the way to fix that is not > to keep adding names, because a hand-made list of two hundred would be two hundred names in > thirty-nine languages, carried by this project for ever (#140). > > **The translations are the real argument, more than the taxonomy.** Eurostat publishes NACE > […] Every word of that transfers. **ESCO is to occupations and skills what NACE is to industries**: published by the European Commission, roughly 3,000 occupations and 14,000 skills, in every EU language, mapped to ISCO-08, and maintained by somebody else. ## What it would buy - **Cross-language matching.** A listing captured in French and a CV written in Portuguese become comparable. Nothing else on the roadmap makes that possible. - **A skills vocabulary for the CV that is not free text**, which the résumé side needs and currently invents per entry. - **Translations that arrive with the data**, exactly as NACE's do — which is the argument #140 already accepted, and which matters more with each language milestone (#70, #71, #72). - **A ladder to the public services.** EURES and the national services already speak ESCO, so a source plugin reading one of them (#241) has somewhere to put what it reads. ## The judgement to make, and it is the same one as NACE ESCO's full occupation list is far too deep for a picker, exactly as NACE's 615 classes were — `industries.py` chose divisions because they are "the level where the names still mean something to the person reading them". The analogue here is probably **ISCO-08 unit groups**: about 436 four-digit codes, which is the same order of magnitude as the 87 NACE divisions relative to its classes, and the level at which a name is still a job somebody recognises. Full ESCO occupations underneath as an optional refinement, or not at all to begin with. That choice should be made deliberately and written down in the module the way `industries.py` wrote down the NACE one, because it is the decision everything else rests on. ## Scope questions this issue has to answer - **Where it attaches.** A title on `JobPosting`, a skill on the résumé, or both — and whether the free-text field stays beside the code, as a person's own wording usually should. - **Licence and distribution.** ESCO is published under the EUPL; the data would ship in `jobs/data/` beside `nace-2.1.json` with its own `LICENCE.md`, as NACE already does. - **Size.** NACE at division level is 87 rows. ESCO is a different order of magnitude, and what ships in the repository versus what is loaded on demand is a real decision rather than a detail. - **Never a requirement.** A person must be able to type a job title that is not in any classification. The code is for what Postulo can do across languages; the text is what the person meant. Filed from the same conversation as #241's additions. Not a plugin: this is a core vocabulary, in the same shape as #140.
tiagoagueda added this to the 0.5.0 milestone 2026-09-17 19:30:34 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#266
No description provided.