Import applications from a spreadsheet (CSV), with a mapping step #31

Closed
opened 2026-09-05 14:04:12 +00:00 by tiagoagueda · 0 comments
Owner

Why

The only import Postulo has is of its own export format (core/importer.py, import_data: "Creates records; never merges"). Everyone arriving from a spreadsheet — which is how most people track a job search until it hurts — or from another tracker's export has to retype their history or start from zero. Onboarding decides whether they stay.

Shape

Settings → Your data → Import from a spreadsheet (#22), four steps:

  1. Upload a CSV. Delimiter and encoding detected (csv.Sniffer; UTF-8 with a BOM, then Latin-1 as the fallback, because Excel). A downloadable template CSV with Postulo's columns for people who prefer to start clean.
  2. Map columns to fields. Company, role, URL, location, status, applied date, deadline, salary, channel, notes, tags, source. Guessed from the header names in English, French and Portuguese ("Entreprise", "Empresa" → company); every guess editable; unmapped columns can go into notes or be ignored.
  3. Preview the first rows as Postulo will read them: dates parsed (day-first or month-first chosen once, per file), statuses mapped through a small table from whatever the spreadsheet said to Postulo's statuses (unknown → applied, with the original in a note), companies matched by name the way the importer already does (name__iexact, line 187) or created.
  4. Import in one transaction, with a report: rows imported, companies created, rows skipped and why. Every imported application gets a timeline event "imported from file.csv", so provenance is never in doubt. Rows with no applied date become listings (#25) rather than applications.
  • Deduplication by URL when present, otherwise by (company, role, applied date); duplicates are reported, not created.
  • Command-line twin: manage.py import_csv <email> <file> --mapping <json>, for the person with a 2,000-row history who would rather script it.
  • Not Excel. Asking for Save as CSV avoids a dependency on openpyxl for a one-time operation. Revisit if people stumble.

Classification

Enhancement. Not breaking: a new import path; the existing importer is untouched.

Open questions

  1. Recognise the export formats of specific trackers by their headers (Huntr, Teal, a Notion table) as ready-made mappings? Only once someone brings a real file; guessing at their columns is how mappings rot.
  2. Import companies and contacts from a second CSV, or applications only? Applications only first; a companies import is the same machinery with a different field list.
## Why The only import Postulo has is of its own export format (`core/importer.py`, `import_data`: "Creates records; never merges"). Everyone arriving from a spreadsheet — which is how most people track a job search until it hurts — or from another tracker's export has to retype their history or start from zero. Onboarding decides whether they stay. ## Shape Settings → Your data → **Import from a spreadsheet** (#22), four steps: 1. **Upload a CSV.** Delimiter and encoding detected (`csv.Sniffer`; UTF-8 with a BOM, then Latin-1 as the fallback, because Excel). A downloadable **template CSV** with Postulo's columns for people who prefer to start clean. 2. **Map columns to fields.** Company, role, URL, location, status, applied date, deadline, salary, channel, notes, tags, source. Guessed from the header names in English, French and Portuguese ("Entreprise", "Empresa" → company); every guess editable; unmapped columns can go into notes or be ignored. 3. **Preview** the first rows as Postulo will read them: dates parsed (day-first or month-first chosen once, per file), statuses mapped through a small table from whatever the spreadsheet said to Postulo's statuses (*unknown → applied, with the original in a note*), companies matched by name the way the importer already does (`name__iexact`, line 187) or created. 4. **Import** in one transaction, with a report: rows imported, companies created, rows skipped and why. Every imported application gets a timeline event "imported from *file.csv*", so provenance is never in doubt. Rows with no applied date become **listings** (#25) rather than applications. - **Deduplication** by URL when present, otherwise by (company, role, applied date); duplicates are reported, not created. - **Command-line twin**: `manage.py import_csv <email> <file> --mapping <json>`, for the person with a 2,000-row history who would rather script it. - **Not Excel.** Asking for *Save as CSV* avoids a dependency on `openpyxl` for a one-time operation. Revisit if people stumble. ## Classification Enhancement. Not breaking: a new import path; the existing importer is untouched. ## Open questions 1. Recognise the export formats of specific trackers by their headers (Huntr, Teal, a Notion table) as ready-made mappings? Only once someone brings a real file; guessing at their columns is how mappings rot. 2. Import companies and contacts from a second CSV, or applications only? Applications only first; a companies import is the same machinery with a different field list.
tiagoagueda added this to the 0.2.0 milestone 2026-09-05 14:04:12 +00:00
tiagoagueda referenced this issue from a commit 2026-09-05 21:24:06 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#31
No description provided.