Where the list of areas of activity comes from #140

Closed
opened 2026-09-09 10:41:56 +00:00 by tiagoagueda · 1 comment
Owner

Observation

you shoud looke for a more extenside and general area of activity list

Prerequisite. Which list, at what depth, and under what licence — answerable on its own, and
the answer decides most of what the company form then looks like.

What exists

jobs/industries.py holds 32 names, assembled by hand:

Broad fields only. A person who works in a niche types the niche, and it joins their own
vocabulary like anything else.

They are gettext strings, so they are translated with the interface — all 39 languages,
as part of Postulo's own catalogues. They are offered as suggestions in a <datalist> and
never as a closed list; the real vocabulary is Industry, one set per person, created by
slug so Fintech and fintech are one thing.

Thirty-two is not a classification. It has Gaming and E-commerce but no Water supply,
no Mining, no Arts, no Public administration as distinct from Public sector, and
nothing at all for a third of the economy.

What this asks for

A real list of areas of activity to seed from, general enough to cover the economy and
recognisable enough that somebody picks their own without reading a manual.

The candidates

NACE — the EU's statistical classification of economic activities, the obvious fit for a
project that has already decided to put Europe first. Rev. 2.1 was adopted in October 2022
and is in use for European statistics from 2025. Rev. 2's structure is four levels: 21
sections
(letters A–U), 88 divisions (two digits), 272 groups, 615 classes. Published
by Eurostat in every official EU language, and available as Linked Open Data.

ISIC — the UN classification NACE derives from, and the honest choice once Postulo is
past Europe (#71, #72, #101). NACE maps onto it, so choosing NACE now does not close it off.

NAICS / SIC — North America. Wrong first audience for this project.

A proprietary list — LinkedIn's ~150 industries and its imitators are the vocabulary
people actually recognise, and none of them come with a licence that permits shipping them.

Worth being careful about

The translations are the argument, more than the taxonomy. A hand-made list of 200 names
would be 200 × 39 = 7,800 translations, carried by this project for ever. NACE is
published in every official EU language by the body that maintains it. That is the single
biggest reason to take a standard rather than write a longer list, and it is worth saying
before anybody starts adding names to industries.py.

Granularity decides whether this is usable at all. 615 classes is a form nobody fills in.
21 sections is coarser than the list we have now — Information and communication covers
software, telecoms, publishing and film in one word. 88 divisions is probably the answer, and
it should be chosen deliberately rather than by taking whatever file is easiest to parse.

NACE classifies the business, not the job, and the difference will bite. Somebody
applying to a bank's software team is applying to K — Financial and insurance activities,
which is true of the employer and useless to the applicant. The current vocabulary is the
applicant's own words about a company, which is a different question wearing the same label.
Whatever is adopted has to survive that: a person who wants to write Fintech must still be
able to.

So this is almost certainly a seed, not a replacement. Industry is per person by
design, with the reason written into the model: "two people on one instance may describe the
same employer differently, and neither should see the other's words."
A closed standard list
would reverse that. Seeding the suggestions from NACE and letting the free vocabulary stand
keeps both — and a stored code beside the name is what makes a report (#56) legible to an
employment office that thinks in NACE.

The licence has to be checked before anything is vendored, not assumed from "it is EU
data". The instruments to read are Regulation (EC) 1893/2006 and Commission Delegated
Regulation 2023/137, and Decision 2011/833/EU on the reuse of Commission documents. An
AGPL project shipping a data file needs the answer in writing.

Whatever is vendored has to be updatable. NACE Rev. 2.1 replaced Rev. 2 after sixteen
years; a list copied into a Python module will be copied again. Where it lives — a data file
with its revision recorded, rather than a tuple in a module — is part of this decision.

People already have industries. Every existing Industry row is somebody's own word, and
seeding a standard list must not duplicate, rename or absorb them. Matching by slug is what
already stops Fintech and fintech from breeding; it will not stop Software from
existing beside J62 Computer programming.

## Observation > you shoud looke for a more extenside and general area of activity list Prerequisite. Which list, at what depth, and under what licence — answerable on its own, and the answer decides most of what the company form then looks like. ## What exists `jobs/industries.py` holds **32 names**, assembled by hand: > Broad fields only. A person who works in a niche types the niche, and it joins their own > vocabulary like anything else. They are `gettext` strings, so they are translated with the interface — all 39 languages, as part of Postulo's own catalogues. They are offered as suggestions in a `<datalist>` and never as a closed list; the real vocabulary is `Industry`, one set per person, created by slug so *Fintech* and *fintech* are one thing. Thirty-two is not a classification. It has *Gaming* and *E-commerce* but no *Water supply*, no *Mining*, no *Arts*, no *Public administration* as distinct from *Public sector*, and nothing at all for a third of the economy. ## What this asks for A real list of areas of activity to seed from, general enough to cover the economy and recognisable enough that somebody picks their own without reading a manual. ## The candidates **NACE** — the EU's statistical classification of economic activities, the obvious fit for a project that has already decided to put Europe first. Rev. 2.1 was adopted in October 2022 and is in use for European statistics from 2025. Rev. 2's structure is four levels: **21 sections** (letters A–U), **88 divisions** (two digits), 272 groups, 615 classes. Published by Eurostat **in every official EU language**, and available as Linked Open Data. **ISIC** — the UN classification NACE derives from, and the honest choice once Postulo is past Europe (#71, #72, #101). NACE maps onto it, so choosing NACE now does not close it off. **NAICS / SIC** — North America. Wrong first audience for this project. **A proprietary list** — LinkedIn's ~150 industries and its imitators are the vocabulary people actually recognise, and none of them come with a licence that permits shipping them. ## Worth being careful about **The translations are the argument, more than the taxonomy.** A hand-made list of 200 names would be 200 × 39 = **7,800 translations**, carried by this project for ever. NACE is published in every official EU language by the body that maintains it. That is the single biggest reason to take a standard rather than write a longer list, and it is worth saying before anybody starts adding names to `industries.py`. **Granularity decides whether this is usable at all.** 615 classes is a form nobody fills in. 21 sections is coarser than the list we have now — *Information and communication* covers software, telecoms, publishing and film in one word. 88 divisions is probably the answer, and it should be chosen deliberately rather than by taking whatever file is easiest to parse. **NACE classifies the business, not the job, and the difference will bite.** Somebody applying to a bank's software team is applying to *K — Financial and insurance activities*, which is true of the employer and useless to the applicant. The current vocabulary is the applicant's own words about a company, which is a different question wearing the same label. Whatever is adopted has to survive that: a person who wants to write *Fintech* must still be able to. **So this is almost certainly a seed, not a replacement.** `Industry` is per person by design, with the reason written into the model: *"two people on one instance may describe the same employer differently, and neither should see the other's words."* A closed standard list would reverse that. Seeding the suggestions from NACE and letting the free vocabulary stand keeps both — and a stored code beside the name is what makes a report (#56) legible to an employment office that thinks in NACE. **The licence has to be checked before anything is vendored**, not assumed from "it is EU data". The instruments to read are Regulation (EC) 1893/2006 and Commission Delegated Regulation 2023/137, and Decision 2011/833/EU on the reuse of Commission documents. An AGPL project shipping a data file needs the answer in writing. **Whatever is vendored has to be updatable.** NACE Rev. 2.1 replaced Rev. 2 after sixteen years; a list copied into a Python module will be copied again. Where it lives — a data file with its revision recorded, rather than a tuple in a module — is part of this decision. **People already have industries.** Every existing `Industry` row is somebody's own word, and seeding a standard list must not duplicate, rename or absorb them. Matching by slug is what already stops *Fintech* and *fintech* from breeding; it will not stop *Software* from existing beside *J62 Computer programming*.
tiagoagueda added this to the 0.3.0 milestone 2026-09-09 10:41:56 +00:00
Author
Owner

Done in 0aff31fd. The three questions, answered.

Which list: NACE

For the reason the issue gives — Europe first, and ISIC stays open because NACE maps onto it. NAICS and SIC are the wrong first audience; LinkedIn's ~150 industries are the vocabulary people recognise and come with no licence that permits shipping them.

At what depth: divisions

87 divisions under 22 sections — Rev. 2.1's shape, which is not Rev. 2's: the restructuring split Information and communication into a publishing/broadcasting section and a telecom/computing one, and there are 22 sections (A–V) rather than 21.

Chosen between three rather than taken from whatever parsed easiest, exactly as asked:

  • 21/22 sections are coarser than the hand-made list ever was.
  • 615 classes are a form nobody fills in.
  • 87 divisions is the level where the names still mean something to the person reading them.

Under what licence: CC BY 4.0, and checked rather than assumed

  • Reuse of Commission documents is authorised under Commission Decision 2011/833/EU.
  • The Commission's default licence for content it owns is CC BY 4.0 — any purpose, commercial included, with credit given and changes indicated.
  • The Decision's exclusions (software, trademarks, logos, names; third-party material) do not apply: this is Commission-owned reference data, no mark travels with it, nothing in it is third-party.

src/postulo/jobs/data/LICENCE.md carries the attribution verbatim, says exactly what was changed (Eurostat prefixes each label with its own code; the code is stored separately and stripped from the name — nothing was translated, renamed, merged or added), and says how to replace the file when Rev. 3 arrives.

Where it came from: the Publications Office's EU Vocabularies copy, concept scheme http://data.europa.eu/ux2/nace2.1/, through the ShowVoc SPARQL endpoint. The harvest script is in the issue rather than in the repository, on purpose: a build step that fetches from the internet is a build step that fails when somebody else's server does.

The translations, which were the real argument

87 × 24 = 2,088 names, written by the body that maintains them. Not one is a gettext string here — they are reference data in the same sense a language's own name is, and messages.py never sees them. One new translatable string went into the catalogues: the field label NACE division.

Fifteen of the thirty-nine languages Postulo speaks are not official EU languages. Those readers get the English division names, and the module says why: inventing NACE names for Catalan or Ukrainian would be this project asserting a classification it does not maintain. pt-br reads the Portuguese names and en-gb the English ones — a variant of a language is the same words.

It bites exactly where you said it would, and survives

Somebody applying to a bank's software team is applying to K — Financial and insurance activities, which is true of the employer and useless to the applicant.

So this is a seed, and three things make that real rather than stated:

  1. Industry is unchanged in kind — still one free vocabulary per person, still matched by slug, still theirs.
  2. Postulo's own 32 short names stay, and stay first. A <datalist> shows its options in order until somebody types, and Software being visible before Manufacture of coke and refined petroleum products is the difference between a helpful list and a statistical yearbook. They are already translated, so they cost nothing. (Education is in both; it is offered once.)
  3. A word somebody made up has no code, and that is not a lesser kind of industry. Fintech works exactly as before.

What the standard buys beyond coverage is Industry.code: a name that matches a division carries it, so a report (#56) can say 62 to an employment office that thinks in NACE while the person goes on reading their own word. The code follows the name rather than sitting beside it — rename an industry to a division and it gains that code; rename it away and it loses it. One invariant, no way for the two to disagree.

Existing rows

Untouched, and tested for it: the migration contains no RunPython. It adds a column and widens name from 60 to 160, because sixty is generous for a word somebody types and too short for Computing infrastructure, data processing, hosting and other information service activities (92). A code arrives when a name is saved that matches a division, never retroactively — a migration that went looking for codes to attach would be this project deciding what somebody meant.

Also

  • The gaps you named are gone: Mining of metal ores, Water collection, treatment and supply, Arts creation and performing arts activities, Public administration and defence; compulsory social security — each a test, because those were the point.
  • The archive needs no format change: industries travel by name, and the code is derived on save.
  • tests/test_nace.py, 23 tests. Suite 4353 passed, 29 skipped; browser suite 54 passed.
Done in `0aff31fd`. The three questions, answered. ## Which list: NACE For the reason the issue gives — Europe first, and ISIC stays open because NACE maps onto it. NAICS and SIC are the wrong first audience; LinkedIn's ~150 industries are the vocabulary people recognise and come with no licence that permits shipping them. ## At what depth: divisions **87 divisions under 22 sections** — Rev. 2.1's shape, which is not Rev. 2's: the restructuring split *Information and communication* into a publishing/broadcasting section and a telecom/computing one, and there are 22 sections (A–V) rather than 21. Chosen between three rather than taken from whatever parsed easiest, exactly as asked: - **21/22 sections** are coarser than the hand-made list ever was. - **615 classes** are a form nobody fills in. - **87 divisions** is the level where the names still mean something to the person reading them. ## Under what licence: CC BY 4.0, and checked rather than assumed - Reuse of Commission documents is authorised under **Commission Decision 2011/833/EU**. - The Commission's default licence for content it owns is **CC BY 4.0** — any purpose, commercial included, with credit given and changes indicated. - The Decision's exclusions (software, trademarks, logos, names; third-party material) do not apply: this is Commission-owned reference data, no mark travels with it, nothing in it is third-party. `src/postulo/jobs/data/LICENCE.md` carries the attribution verbatim, says exactly what was changed (Eurostat prefixes each label with its own code; the code is stored separately and stripped from the name — nothing was translated, renamed, merged or added), and says how to replace the file when Rev. 3 arrives. **Where it came from**: the Publications Office's EU Vocabularies copy, concept scheme `http://data.europa.eu/ux2/nace2.1/`, through the ShowVoc SPARQL endpoint. The harvest script is in the issue rather than in the repository, on purpose: a build step that fetches from the internet is a build step that fails when somebody else's server does. ## The translations, which were the real argument **87 × 24 = 2,088 names, written by the body that maintains them.** Not one is a gettext string here — they are reference data in the same sense a language's own name is, and `messages.py` never sees them. One new translatable string went into the catalogues: the field label *NACE division*. Fifteen of the thirty-nine languages Postulo speaks are not official EU languages. Those readers get the **English** division names, and the module says why: inventing NACE names for Catalan or Ukrainian would be this project asserting a classification it does not maintain. `pt-br` reads the Portuguese names and `en-gb` the English ones — a variant of a language is the same words. ## It bites exactly where you said it would, and survives > Somebody applying to a bank's software team is applying to *K — Financial and insurance activities*, which is true of the employer and useless to the applicant. So this is a **seed**, and three things make that real rather than stated: 1. **`Industry` is unchanged in kind** — still one free vocabulary per person, still matched by slug, still theirs. 2. **Postulo's own 32 short names stay, and stay first.** A `<datalist>` shows its options in order until somebody types, and *Software* being visible before *Manufacture of coke and refined petroleum products* is the difference between a helpful list and a statistical yearbook. They are already translated, so they cost nothing. (*Education* is in both; it is offered once.) 3. **A word somebody made up has no code, and that is not a lesser kind of industry.** *Fintech* works exactly as before. **What the standard buys beyond coverage** is `Industry.code`: a name that matches a division carries it, so a report (#56) can say *62* to an employment office that thinks in NACE while the person goes on reading their own word. The code **follows the name** rather than sitting beside it — rename an industry to a division and it gains that code; rename it away and it loses it. One invariant, no way for the two to disagree. ## Existing rows Untouched, and tested for it: the migration contains no `RunPython`. It adds a column and widens `name` from 60 to 160, because sixty is generous for a word somebody types and too short for *Computing infrastructure, data processing, hosting and other information service activities* (92). A code arrives when a name is saved that matches a division, never retroactively — a migration that went looking for codes to attach would be this project deciding what somebody meant. ## Also - The gaps you named are gone: *Mining of metal ores*, *Water collection, treatment and supply*, *Arts creation and performing arts activities*, *Public administration and defence; compulsory social security* — each a test, because those were the point. - The archive needs no format change: industries travel by name, and the code is derived on save. - `tests/test_nace.py`, 23 tests. Suite 4353 passed, 29 skipped; browser suite 54 passed.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Reference
Postulo/postulo#140
No description provided.