Company identifiers should match regardless of letter case #211

Closed
opened 2026-09-15 19:33:31 +00:00 by tiagoagueda · 0 comments
Owner

Company identifiers should match regardless of letter case. q95 and Q95, Aperture-Science and aperture-science, or a staff number typed ab-12 and AB-12 should be one identifier: found by the same lookups, refused as the same duplicate, and never allowed onto two companies.

How matching works today

Every comparison is an exact value= match on the stored string:

  • Company.by_identifier (jobs/models.py:331), used by the CSV import (core/csv_import.py:787), the importer (core/importer.py:511) and the Wikidata link on capture (applications/services.py:193);
  • the formset's duplicate checks on the company page (jobs/forms.py:309-332): already listed and already carries this identifier;
  • jobs/services.py:40-90, which sets identifiers from elsewhere: the seen set, the clash with another company, and get_or_create;
  • the database constraints on CompanyIdentifier (jobs/models.py:374-389), which PostgreSQL and SQLite both compare case-sensitively.

Where that already works, and where it does not

Named schemes are safe for new values. Scheme.normalise (core/identifiers.py:95) folds case before anything is stored or looked up:

Folded to upper case Folded to lower case
Wikidata, ISNI, LEI, company register number LinkedIn, Crunchbase, OpenCorporates

Other is not folded at all. The Other scheme has neither upper nor lower, so:

  • the same value in two cases on one company passes the formset check (seen_values), the service's seen set, and unique_other_identifier_per_company;
  • Other is skipped entirely by the "another company already carries this" check (scheme != identifiers.OTHER) and by one_company_per_identifier_per_owner. That part may be deliberate, since two companies can share a staff-number format, but it should be stated either way.

Rows saved before folding existed. No migration ever re-normalised stored CompanyIdentifier values. A row saved before a scheme gained upper/lower (for example a lowercase q95 from an early Wikidata import, or a mixed-case LinkedIn slug) is invisible to every exact lookup above, and a correctly folded value can be added beside it on another company.

What a fix has to settle

  • Other's case: fold it for comparison only, and keep displaying it as typed, since a staff number AB-12 should still read AB-12. That means a case-insensitive comparison (value__iexact, or a Lower("value") expression in the constraints) rather than folding the stored value.
  • Constraints: replace the value-based UniqueConstraints with expression constraints on Lower("value") (Django supports UniqueConstraint(Lower("value"), "scheme", …)). Check that the SQLite and PostgreSQL migrations both build.
  • A data migration that re-normalises existing named-scheme values through identifiers.clean. It must report, not silently merge, any pair that becomes a duplicate once folded: two companies may turn out to be one employer, and choosing which to keep is the person's call.
  • Lookups: by_identifier, the formset, and jobs/services.py compare the same way as the constraint, so the form refuses with a readable message before the database raises IntegrityError. This matches what CompanyForm.clean_name already does for names with name__iexact (jobs/forms.py:253-268).
  • Not accent- or width-insensitive: case only. Lower() in PostgreSQL follows the database locale, and in SQLite it folds ASCII only, so a test with a non-ASCII Other value should pin down what is promised.
  • People: PersonIdentifier (accounts/models.py:152) has the same shape. This issue is about companies; decide whether it follows and file it separately.

Tests

  • Adding Q95 to one company and q95 to another is refused, through the company form and through jobs/services.py.
  • An Other identifier AB-12 and ab-12 on the same company is refused, and displays as first typed.
  • Company.by_identifier(owner, "linkedin", "Aperture-Science") finds a company stored as aperture-science.
  • The data migration folds a legacy lowercase Wikidata row and reports a folded duplicate instead of dropping it.
Company identifiers should match regardless of letter case. `q95` and `Q95`, `Aperture-Science` and `aperture-science`, or a staff number typed `ab-12` and `AB-12` should be one identifier: found by the same lookups, refused as the same duplicate, and never allowed onto two companies. ## How matching works today Every comparison is an exact `value=` match on the stored string: - `Company.by_identifier` (`jobs/models.py:331`), used by the CSV import (`core/csv_import.py:787`), the importer (`core/importer.py:511`) and the Wikidata link on capture (`applications/services.py:193`); - the formset's duplicate checks on the company page (`jobs/forms.py:309-332`): *already listed* and *already carries this identifier*; - `jobs/services.py:40-90`, which sets identifiers from elsewhere: the `seen` set, the clash with another company, and `get_or_create`; - the database constraints on `CompanyIdentifier` (`jobs/models.py:374-389`), which PostgreSQL and SQLite both compare case-sensitively. ## Where that already works, and where it does not **Named schemes are safe for new values.** `Scheme.normalise` (`core/identifiers.py:95`) folds case before anything is stored or looked up: | Folded to upper case | Folded to lower case | | --- | --- | | Wikidata, ISNI, LEI, company register number | LinkedIn, Crunchbase, OpenCorporates | **Other is not folded at all.** The *Other* scheme has neither `upper` nor `lower`, so: - the same value in two cases on one company passes the formset check (`seen_values`), the service's `seen` set, and `unique_other_identifier_per_company`; - *Other* is skipped entirely by the "another company already carries this" check (`scheme != identifiers.OTHER`) and by `one_company_per_identifier_per_owner`. That part may be deliberate, since two companies can share a staff-number format, but it should be stated either way. **Rows saved before folding existed.** No migration ever re-normalised stored `CompanyIdentifier` values. A row saved before a scheme gained `upper`/`lower` (for example a lowercase `q95` from an early Wikidata import, or a mixed-case LinkedIn slug) is invisible to every exact lookup above, and a correctly folded value can be added beside it on another company. ## What a fix has to settle - **Other's case:** fold it for comparison only, and keep displaying it as typed, since a staff number `AB-12` should still read `AB-12`. That means a case-insensitive comparison (`value__iexact`, or a `Lower("value")` expression in the constraints) rather than folding the stored value. - **Constraints:** replace the value-based `UniqueConstraint`s with expression constraints on `Lower("value")` (Django supports `UniqueConstraint(Lower("value"), "scheme", …)`). Check that the SQLite and PostgreSQL migrations both build. - **A data migration** that re-normalises existing named-scheme values through `identifiers.clean`. It must report, not silently merge, any pair that becomes a duplicate once folded: two companies may turn out to be one employer, and choosing which to keep is the person's call. - **Lookups:** `by_identifier`, the formset, and `jobs/services.py` compare the same way as the constraint, so the form refuses with a readable message before the database raises `IntegrityError`. This matches what `CompanyForm.clean_name` already does for names with `name__iexact` (`jobs/forms.py:253-268`). - **Not accent- or width-insensitive:** case only. `Lower()` in PostgreSQL follows the database locale, and in SQLite it folds ASCII only, so a test with a non-ASCII *Other* value should pin down what is promised. - **People:** `PersonIdentifier` (`accounts/models.py:152`) has the same shape. This issue is about companies; decide whether it follows and file it separately. ## Tests - Adding `Q95` to one company and `q95` to another is refused, through the company form and through `jobs/services.py`. - An *Other* identifier `AB-12` and `ab-12` on the same company is refused, and displays as first typed. - `Company.by_identifier(owner, "linkedin", "Aperture-Science")` finds a company stored as `aperture-science`. - The data migration folds a legacy lowercase Wikidata row and reports a folded duplicate instead of dropping it.
tiagoagueda added this to the 0.4.0 milestone 2026-09-15 21:33:21 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#211
No description provided.