Company logos: from a URL first, then found on the company's website, then from own media #21

Closed
opened 2026-09-05 12:34:58 +00:00 by tiagoagueda · 0 comments
Owner

Observation

add logo to companies, for now from a url, after that either from a public repository or own media

What exists today

  • Company (jobs/models.py, line 20) has website, careers_url, location, industry and notes. No image. Nothing in Postulo displays an image of anything yet; #7 adds the first (avatars) and this issue reuses its choices rather than making new ones.
  • The production CSP is img-src 'self' data: (prod.py, line 42), so <img src="https://example.com/logo.png"> would not render. "From a URL" therefore cannot mean "show the URL"; it has to mean "fetch it once and keep it".
  • plugins/fetching.py already has what a fetch needs: validate_public_url (resolves the name, refuses private, loopback and link-local addresses) and a capped, timeout-bound httpx fetch. A logo URL is public by nature, so the guard applies unchanged; unlike #11's connections, no operator switch is wanted here.
  • Pillow 12.3 is installed, as WeasyPrint's dependency.

Stage one: from a URL

A Logo URL field on the company form. On save, Postulo fetches it server-side — the guard, a size cap of a couple of megabytes, the response must be an image — re-encodes it with Pillow to square thumbnails (128 and 256 pixels are enough everywhere it shows), stores them under private media and serves them through serve_private_file, like every other file. The URL is kept for Refresh. A failed fetch is shown on the form ("could not fetch: …") and does not stop the company from being saved.

Raster only in this stage: PNG, JPEG, GIF, WebP. SVG is the format company logos most often come in and it is the one that needs care — it can carry scripts and references to other files. As an <img> source a browser will not run a script, but a direct visit to the file's URL is a different context. Accepting SVG means a sanitiser (an allow-list over the tree, dropping script, foreignObject and external references) and a restrictive Content-Security-Policy header on the file response. Worth doing, as its own step, after the raster path exists.

The fallback, for a company without a logo: an initials tile coloured from the name, exactly #7's avatar fallback with a different shape. seed_demo's fictional companies get that and nothing else.

Stage two: found on the company's website

The public repository every company maintains is its own site, and Company.website is already recorded. Find logo on the form fetches the homepage — the same guarded fetch capture uses, with the same courtesy towards robots.txt — and looks, in order, for the largest apple-touch-icon or <link rel="icon">, then og:image or og:logo, then the schema.org Organization logo in JSON-LD (capture already parses JSON-LD with the standard library), then /favicon.ico. Once it proves reliable, it can run automatically when a company is created with a website: one request to a site the person typed, in the same spirit as capture.

Third-party logo services do not belong in the core. Clearbit's free logo API was announced for retirement at the end of 2025; Logo.dev and Brandfetch need a key and carry terms; Wikidata has logos for notable companies under mixed licences; the Google and DuckDuckGo favicon services report every company a person looks up to a third party. Each is a legitimate choice for someone, none for everyone, which is the definition of a plugin: a postulo.logo_sources group later, if anyone wants one, with "from the website" as the built-in source and #11's connections for the keys.

Stage three: from own media

Upload on the company form, through the same validation and re-encoding as stage one — the path #7 builds for avatars. A company holds one logo with a recorded origin (URL, website, upload), with Replace and Remove; the most recent action wins, so there is no precedence to explain.

Where it shows

Small beside the name in the companies table and the applications table (a column for #20), on board cards and on the dashboard's recent items; large on the company page. Not on CVs or letters: a company's logo on a candidate's own document would be odd.

Data

Company.logo (an image under private media), logo_source (URL, website, upload), logo_source_url, logo_fetched_at. Export and import: the file joins the media archive, and the URL is kept so an import without the file can fetch it again.

Classification

Enhancement. Not breaking: nullable fields, nothing changes for a company without a logo.

Open questions

  1. Automatic website lookup on creation from the start, or only on request until it proves reliable? Proposal: on request first.
  2. Accept SVG in stage one? Proposal: no, raster only until the sanitiser exists.
  3. Share one fetched logo between two people who recorded the same company? No: companies are owner-scoped on purpose (the model's docstring says why), and a shared file would reveal who else applied there.
## Observation > add logo to companies, for now from a url, after that either from a public repository or own media ## What exists today - `Company` (`jobs/models.py`, line 20) has `website`, `careers_url`, location, industry and notes. No image. Nothing in Postulo displays an image of anything yet; #7 adds the first (avatars) and this issue reuses its choices rather than making new ones. - **The production CSP is `img-src 'self' data:`** (`prod.py`, line 42), so `<img src="https://example.com/logo.png">` would not render. "From a URL" therefore cannot mean "show the URL"; it has to mean "fetch it once and keep it". - `plugins/fetching.py` already has what a fetch needs: `validate_public_url` (resolves the name, refuses private, loopback and link-local addresses) and a capped, timeout-bound `httpx` fetch. A logo URL is public by nature, so the guard applies unchanged; unlike #11's connections, no operator switch is wanted here. - Pillow 12.3 is installed, as WeasyPrint's dependency. ## Stage one: from a URL A *Logo URL* field on the company form. On save, Postulo fetches it server-side — the guard, a size cap of a couple of megabytes, the response must be an image — re-encodes it with Pillow to square thumbnails (128 and 256 pixels are enough everywhere it shows), stores them under private media and serves them through `serve_private_file`, like every other file. The URL is kept for *Refresh*. A failed fetch is shown on the form ("could not fetch: …") and does not stop the company from being saved. **Raster only in this stage: PNG, JPEG, GIF, WebP.** SVG is the format company logos most often come in and it is the one that needs care — it can carry scripts and references to other files. As an `<img>` source a browser will not run a script, but a direct visit to the file's URL is a different context. Accepting SVG means a sanitiser (an allow-list over the tree, dropping `script`, `foreignObject` and external references) and a restrictive `Content-Security-Policy` header on the file response. Worth doing, as its own step, after the raster path exists. **The fallback**, for a company without a logo: an initials tile coloured from the name, exactly #7's avatar fallback with a different shape. `seed_demo`'s fictional companies get that and nothing else. ## Stage two: found on the company's website The public repository every company maintains is its own site, and `Company.website` is already recorded. *Find logo* on the form fetches the homepage — the same guarded fetch capture uses, with the same courtesy towards `robots.txt` — and looks, in order, for the largest `apple-touch-icon` or `<link rel="icon">`, then `og:image` or `og:logo`, then the schema.org `Organization` logo in JSON-LD (capture already parses JSON-LD with the standard library), then `/favicon.ico`. Once it proves reliable, it can run automatically when a company is created with a website: one request to a site the person typed, in the same spirit as capture. **Third-party logo services do not belong in the core.** Clearbit's free logo API was announced for retirement at the end of 2025; Logo.dev and Brandfetch need a key and carry terms; Wikidata has logos for notable companies under mixed licences; the Google and DuckDuckGo favicon services report every company a person looks up to a third party. Each is a legitimate choice for someone, none for everyone, which is the definition of a plugin: a `postulo.logo_sources` group later, if anyone wants one, with "from the website" as the built-in source and #11's connections for the keys. ## Stage three: from own media Upload on the company form, through the same validation and re-encoding as stage one — the path #7 builds for avatars. A company holds one logo with a recorded origin (URL, website, upload), with *Replace* and *Remove*; the most recent action wins, so there is no precedence to explain. ## Where it shows Small beside the name in the companies table and the applications table (a column for #20), on board cards and on the dashboard's recent items; large on the company page. Not on CVs or letters: a company's logo on a candidate's own document would be odd. ## Data `Company.logo` (an image under private media), `logo_source` (URL, website, upload), `logo_source_url`, `logo_fetched_at`. Export and import: the file joins the media archive, and the URL is kept so an import without the file can fetch it again. ## Classification Enhancement. Not breaking: nullable fields, nothing changes for a company without a logo. ## Open questions 1. Automatic website lookup on creation from the start, or only on request until it proves reliable? Proposal: on request first. 2. Accept SVG in stage one? Proposal: no, raster only until the sanitiser exists. 3. Share one fetched logo between two people who recorded the same company? No: companies are owner-scoped on purpose (the model's docstring says why), and a shared file would reveal who else applied there.
tiagoagueda added this to the 0.2.0 milestone 2026-09-05 12:34:58 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#21
No description provided.