Companies carry external identifiers, starting with a Wikidata id #42

Closed
opened 2026-09-05 18:10:36 +00:00 by tiagoagueda · 0 comments
Owner

Observation

every company can have several unique ids associated, for starting, a wikidata id

What exists today

A company is a name, per person (jobs/models.py, Company, unique on owner and name). Nothing ties "Aperture Science" in one account to the same employer in another, to a public record, or to a future data source. The industries vocabulary (#39) gave companies their first structured attribute; identifiers are the second.

Shape

  1. CompanyIdentifier model: company, scheme, value; unique per (company, scheme) and per (owner, scheme, value), so one account cannot give two companies the same Wikidata id. scheme is a choice with room to grow: Wikidata (Q...) first; then the ones people actually have to hand: a LEI (ISO 17442, 20 characters), a national company register number (VAT / SIREN / NIPC, with the country), a LinkedIn company slug, a Crunchbase slug, an OpenCorporates id. Each scheme knows how to validate its value and how to build a link to it.
  2. Wikidata specifically: a Q followed by digits, linking to https://www.wikidata.org/wiki/Q.... Later, behind a deliberate per-request action and never automatically, a Look up on Wikidata button that searches the label and offers matches, and a Fill from Wikidata that pulls the official website, headquarters and industry (P452) into the company. That is a plugin-shaped feature and a network request, so it waits for the person to ask.
  3. Interface: an Identifiers block on the company page listing each scheme with its value as a link; add and remove on the company form as a small repeating row (scheme select plus value). The companies table gets an optional column per scheme (#20).
  4. Everywhere else: export as a list under each company (format 3 is still unreleased, so no bump), the importer reads it, the API exposes identifiers: [{scheme, value, url}] on companies and accepts them on create and patch, seed_demo gives a few companies an id.
  5. Matching: get_or_create_company keeps matching by name; a Wikidata id, when both sides have one, is a stronger match for the CSV importer (#31) and for capture plugins that know the employer's id.

Classification

Enhancement. Not breaking: a new table, nothing changes for a company without identifiers.

Open questions

  1. A free-text scheme (any key the person likes) or the closed list above with other as an escape? Proposal: the list, with other carrying a label of its own; a free-text key defeats the point of an identifier.
  2. Should a Wikidata id be looked up to validate that the entity is an organisation? Not in the first cut: it is a network call, and the person typing it knows what they typed.
## Observation > every company can have several unique ids associated, for starting, a wikidata id ## What exists today A company is a name, per person (`jobs/models.py`, `Company`, unique on owner and name). Nothing ties "Aperture Science" in one account to the same employer in another, to a public record, or to a future data source. The industries vocabulary (#39) gave companies their first structured attribute; identifiers are the second. ## Shape 1. **`CompanyIdentifier` model**: company, `scheme`, `value`; unique per (company, scheme) and per (owner, scheme, value), so one account cannot give two companies the same Wikidata id. `scheme` is a choice with room to grow: **Wikidata** (`Q...`) first; then the ones people actually have to hand: a **LEI** (ISO 17442, 20 characters), a national company register number (VAT / SIREN / NIPC, with the country), a **LinkedIn** company slug, a **Crunchbase** slug, an **OpenCorporates** id. Each scheme knows how to validate its value and how to build a link to it. 2. **Wikidata specifically**: a `Q` followed by digits, linking to `https://www.wikidata.org/wiki/Q...`. Later, behind a deliberate per-request action and never automatically, a *Look up on Wikidata* button that searches the label and offers matches, and a *Fill from Wikidata* that pulls the official website, headquarters and industry (P452) into the company. That is a plugin-shaped feature and a network request, so it waits for the person to ask. 3. **Interface**: an *Identifiers* block on the company page listing each scheme with its value as a link; add and remove on the company form as a small repeating row (scheme select plus value). The companies table gets an optional column per scheme (#20). 4. **Everywhere else**: export as a list under each company (format 3 is still unreleased, so no bump), the importer reads it, the API exposes `identifiers: [{scheme, value, url}]` on companies and accepts them on create and patch, `seed_demo` gives a few companies an id. 5. **Matching**: `get_or_create_company` keeps matching by name; a Wikidata id, when both sides have one, is a stronger match for the CSV importer (#31) and for capture plugins that know the employer's id. ## Classification Enhancement. Not breaking: a new table, nothing changes for a company without identifiers. ## Open questions 1. A free-text scheme (any key the person likes) or the closed list above with *other* as an escape? Proposal: the list, with *other* carrying a label of its own; a free-text key defeats the point of an identifier. 2. Should a Wikidata id be looked up to validate that the entity is an organisation? Not in the first cut: it is a network call, and the person typing it knows what they typed.
tiagoagueda added this to the 0.2.0 milestone 2026-09-05 18:10:36 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#42
No description provided.