Languages: the rest of Europe, which finishes the continent #118

Closed
opened 2026-09-08 09:32:41 +00:00 by tiagoagueda · 0 comments
Owner

Decision

on issue #70 push african languages to target 0.4.0
the remaining european languages will target 0.3.0

#70 asked this as its own open question — the rest of the European continent was phase 2 in
#43, was named in no milestone, and would have fallen behind Asia and South America by
default. It is answered here: Europe finishes in 0.3.0, Africa moves to 0.4.0.

The languages

Ukrainian, Turkish, Norwegian, Icelandic, Serbian, Bosnian, Albanian, Macedonian, Georgian,
Armenian, Basque, Catalan, Galician, Welsh, Luxembourgish.

With the 24 from 0.2.0 and Brazilian Portuguese from #110, that is the continent.

Why this one is cheaper than it looks

Several are a short step from a catalogue that already exists, and #110 proved the
method on Brazilian Portuguese: seed from the sibling rather than translate from English
again, because every string has already been translated once by somebody thinking about
this application, and what a near neighbour wants is that work carried across.

New Seed from Why
Bosnian, Serbian (Latin) Croatian mutually intelligible; the differences are lexical and known
Catalan, Galician Spanish, Portuguese Galician is closer to Portuguese than to Spanish
Norwegian (bokmål) Danish written Danish and bokmål are very close

The rest are genuine first translations.

And #110 also proved the hazard, which is worth repeating here rather than rediscovering:
a mechanical pass over a seeded catalogue must be word-boundary anchored. Substituting
without one produced "Contmouse a termo" from contrato, and that is now a test.

What has to be checked per language, not assumed

  1. Plural forms. Welsh has six, Ukrainian, Serbian and Bosnian have three, Icelandic's
    rule is not n != 1. scripts/messages.py writes the slots from its own table and check
    insists they are filled, so the table is what has to be right — and it is the one place
    where copying a neighbour's line without looking makes every count on every page
    ungrammatical, as nearly happened for pt-br.
  2. Scripts and fonts. Ukrainian and Macedonian are Cyrillic, Georgian is Mkhedruli,
    Armenian its own. Serbian is written in both Cyrillic and Latin, which is a decision
    about which one sr means before it is a translation. The bundled font stack has to draw
    all of them, in the browser and in the PDF WeasyPrint renders: a missing glyph is a box,
    and a box on somebody's CV is worse than English.
  3. Dates, numbers, the first day of the week — supplied by USE_L10N, and nobody looks
    at them until they are wrong.
  4. The document themes. A CV is rendered in the document's language, not the interface's,
    so every built-in theme carries the section titles for every language here.
  5. The language picker's own list. Each entry names itself in its own language and is
    marked with the language it is in, so a screen reader pronounces it properly rather than
    reading Georgian with English rules.

What it does not cost

Every user-facing string is already wrapped, and scripts/messages.py extract writes a slot
in every catalogue. A language is a catalogue and nothing else — no code path changes. First
drafts are machine-assisted and flagged draft until a speaker has read them
(docs/TRANSLATING.md), on the principle that a wrong translation somebody can fix beats an
English gap nobody notices.

Right-to-left is not needed here: every language on this list is written left to right.
#67 is closed regardless, so Arabic is not blocked when it arrives with #70 in 0.4.0.

Classification

Enhancement. Not breaking: more catalogues, no code path changes.

## Decision > on issue #70 push african languages to target 0.4.0 > the remaining european languages will target 0.3.0 #70 asked this as its own open question — the rest of the European continent was phase 2 in #43, was named in no milestone, and would have fallen behind Asia and South America by default. It is answered here: Europe finishes in 0.3.0, Africa moves to 0.4.0. ## The languages Ukrainian, Turkish, Norwegian, Icelandic, Serbian, Bosnian, Albanian, Macedonian, Georgian, Armenian, Basque, Catalan, Galician, Welsh, Luxembourgish. With the 24 from 0.2.0 and Brazilian Portuguese from #110, that is the continent. ## Why this one is cheaper than it looks **Several are a short step from a catalogue that already exists**, and #110 proved the method on Brazilian Portuguese: seed from the sibling rather than translate from English again, because every string has already been translated once by somebody thinking about *this* application, and what a near neighbour wants is that work carried across. | New | Seed from | Why | | --- | --- | --- | | Bosnian, Serbian (Latin) | Croatian | mutually intelligible; the differences are lexical and known | | Catalan, Galician | Spanish, Portuguese | Galician is closer to Portuguese than to Spanish | | Norwegian (bokmål) | Danish | written Danish and bokmål are very close | The rest are genuine first translations. **And #110 also proved the hazard**, which is worth repeating here rather than rediscovering: a mechanical pass over a seeded catalogue must be **word-boundary anchored**. Substituting without one produced "Cont**mouse** a termo" from *contrato*, and that is now a test. ## What has to be checked per language, not assumed 1. **Plural forms.** Welsh has **six**, Ukrainian, Serbian and Bosnian have three, Icelandic's rule is not `n != 1`. `scripts/messages.py` writes the slots from its own table and `check` insists they are filled, so the table is what has to be right — and it is the one place where copying a neighbour's line without looking makes every count on every page ungrammatical, as nearly happened for `pt-br`. 2. **Scripts and fonts.** Ukrainian and Macedonian are Cyrillic, Georgian is Mkhedruli, Armenian its own. Serbian is written in **both** Cyrillic and Latin, which is a decision about which one `sr` means before it is a translation. The bundled font stack has to draw all of them, in the browser *and* in the PDF WeasyPrint renders: a missing glyph is a box, and a box on somebody's CV is worse than English. 3. **Dates, numbers, the first day of the week** — supplied by `USE_L10N`, and nobody looks at them until they are wrong. 4. **The document themes.** A CV is rendered in the document's language, not the interface's, so every built-in theme carries the section titles for every language here. 5. **The language picker's own list.** Each entry names itself in its own language and is marked with the language it is in, so a screen reader pronounces it properly rather than reading Georgian with English rules. ## What it does not cost Every user-facing string is already wrapped, and `scripts/messages.py extract` writes a slot in every catalogue. A language is a catalogue and nothing else — no code path changes. First drafts are machine-assisted and flagged `draft` until a speaker has read them (`docs/TRANSLATING.md`), on the principle that a wrong translation somebody can fix beats an English gap nobody notices. Right-to-left is **not** needed here: every language on this list is written left to right. #67 is closed regardless, so Arabic is not blocked when it arrives with #70 in 0.4.0. ## Classification Enhancement. Not breaking: more catalogues, no code path changes.
tiagoagueda added this to the 0.3.0 milestone 2026-09-08 09:32:41 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#118
No description provided.