Choose the next language by how many people speak it, not by which continent it is on #269

Open
opened 2026-09-17 19:43:28 +00:00 by tiagoagueda · 0 comments
Owner

Decision

Which language Postulo learns next is decided by how many people speak it, not by which
continent it is on. #70, #71 and #72 divided the world into Africa, Asia and South America,
and the rest; that division is replaced by this ordering, and those three issues are
re-scoped below rather than closed.

Why geography failed, concretely

Russian has no catalogue. Roughly 255 million speakers, and it is named in none of the
three issues — not in Africa, not in "Asia and South America", not in "the rest of the
world". Russia straddles two continents and each list assumed the other carried it. It is not
in src/postulo/locale/ at all.

That is not an oversight to patch; it is what happens when the organising principle is a map.
A continent is also not a unit of work: #70 is 29 catalogues at about 1,790 strings each —
roughly 52,000 translations — for a milestone due in twenty-five days that holds
twenty-four other issues, while the single largest sweep this project has ever done was 110
strings across 36 languages.

Ordering by speakers spends the same effort on more people, and it cannot lose a language
between two lists.

Where the catalogues actually stand

  • 39 filled, all of Europe and the Caucasus: bg bs ca cs cy da de el es et eu fi fr_FR ga gl hr hu hy is it ka lb lt lv mk mt nb nl pl pt_BR pt_PT ro sk sl sq sr sv tr uk
  • 29 scaffolded and empty, the African set from #70: af ak am ar bm ee ff ha ig kab ln mg nr ny om rw sn so ss st sw ti tn ts ve wo xh yo zu
  • Everything else has no catalogue at all — no ru, zh, hi, bn, ja, ko, vi,
    id, ur, fa, th.

An empty catalogue costs nothing, because accounts/forms.py::language_choices does not
offer a language until somebody starts it. That is what makes this re-ordering free: no
half-finished language is ever visible to anybody.

Two gates that outrank the speaker count

Speaker order sets the priority. It does not override these, and the current plan violates
both.

Fonts (#74). Han, Devanagari, Bengali, Tamil, Thai, Khmer and Ge'ez need glyphs that ship
— #74's title says it: a declared package is not a drawn glyph. Mandarin has more speakers
than anything else here and still waits, because a page of tofu boxes is worse than the same
page in English.

Right-to-left (#73, #67). Arabic, Urdu, Persian and Hebrew need the layout before the
catalogue. ar is currently in #70 at 0.4.0, while the RTL layout work is at 0.6.0 — the
first right-to-left language is scheduled two milestones before the thing it depends on. That
is a bug in the plan regardless of which ordering wins.

The tiers

Speaker figures are approximate and include second-language speakers; they order the list,
they are not a promise.

Tier 1 — 0.4.0. Nothing blocks these.

Latin or Cyrillic, no new font work, no RTL.

  • Russian ~255M — new catalogue. Cyrillic is already proven by bg, mk, sr, uk.
  • Indonesian ~200M — new.
  • Vietnamese ~85M — new; Latin, heavy diacritics.
  • Filipino / Tagalog ~83M — new.
  • Hausa ~80M and Swahili ~72M — ha and sw already scaffolded.
  • The rest of the Latin-script African set already scaffolded: Yoruba, Igbo, Zulu, Xhosa,
    Afrikaans, Somali, Oromo, Kinyarwanda, Shona, Sesotho, Setswana, Wolof, Malagasy and the
    others.

Tier 2 — 0.5.0, gated on #74.

  • Mandarin Chinese ~1.1B — simplified and traditional are two catalogues.
  • Hindi ~610M, Bengali ~280M.
  • Japanese ~125M, Korean ~82M.
  • Marathi, Telugu, Tamil, Punjabi, Gujarati, Kannada.
  • Thai ~60M.
  • Amharic and Tigrinya — Ge'ez script; carved out of #70, which cannot reach them
    before the fonts land.

Tier 3 — 0.6.0, gated on #73 and #67.

  • Arabic ~400M — carved out of #70, where it is currently scheduled ahead of its layout.
  • Urdu ~230M, Persian ~80M, Hebrew.
  • Tamazight / Kabyle (kab) — also carved out of #70.

1.0.0

  • #101, the PALOP national languages, unchanged — it is a commitment about honesty regarding
    machine translation, not a place in this ordering.
  • The long tail as a process rather than a list: #72's second half, which is the part of
    that issue worth keeping.

What changes on the three existing issues

  • #70 stays at 0.4.0 and narrows to the Latin-script African languages. ar, kab,
    am and ti leave it for Tiers 2 and 3, so it can close when the ones that are not
    blocked are done.
  • #71 moves to 0.5.0 — its Asian catalogues are the font-gated tier. Its RTL members
    (ur, fa, he) move to Tier 3. Quechua, Guaraní and Aymara are Latin-script and may be
    pulled into Tier 1 by whoever starts it.
  • #72 moves to 0.6.0, and should be re-read as the process issue it half already is:
    a speaker adding a language without waiting for a developer. That is what makes the tail
    reachable at all, and Weblate is already running for it.

What does not change

  • A language is draft until a speaker reviews it, and stays out of language_choices until
    somebody starts it.
  • The working set while developing is still en / fr-fr / pt-pt, with the full sweep on a
    release commit.
  • tests/test_translations.py::test_every_european_union_language_stays_complete still hard-
    gates the 24 EU languages on every commit, and its conflict with the three-language working
    rule is still unresolved. This issue does not fix that and should not be read as having
    done so.
## Decision Which language Postulo learns next is decided by **how many people speak it**, not by which continent it is on. #70, #71 and #72 divided the world into Africa, Asia and South America, and the rest; that division is replaced by this ordering, and those three issues are re-scoped below rather than closed. ## Why geography failed, concretely **Russian has no catalogue.** Roughly 255 million speakers, and it is named in none of the three issues — not in Africa, not in "Asia and South America", not in "the rest of the world". Russia straddles two continents and each list assumed the other carried it. It is not in `src/postulo/locale/` at all. That is not an oversight to patch; it is what happens when the organising principle is a map. A continent is also not a unit of work: #70 is 29 catalogues at about 1,790 strings each — roughly **52,000 translations** — for a milestone due in twenty-five days that holds twenty-four other issues, while the single largest sweep this project has ever done was 110 strings across 36 languages. Ordering by speakers spends the same effort on more people, and it cannot lose a language between two lists. ## Where the catalogues actually stand - **39 filled**, all of Europe and the Caucasus: `bg bs ca cs cy da de el es et eu fi fr_FR ga gl hr hu hy is it ka lb lt lv mk mt nb nl pl pt_BR pt_PT ro sk sl sq sr sv tr uk` - **29 scaffolded and empty**, the African set from #70: `af ak am ar bm ee ff ha ig kab ln mg nr ny om rw sn so ss st sw ti tn ts ve wo xh yo zu` - **Everything else has no catalogue at all** — no `ru`, `zh`, `hi`, `bn`, `ja`, `ko`, `vi`, `id`, `ur`, `fa`, `th`. An empty catalogue costs nothing, because `accounts/forms.py::language_choices` does not offer a language until somebody starts it. That is what makes this re-ordering free: no half-finished language is ever visible to anybody. ## Two gates that outrank the speaker count Speaker order sets the priority. It does not override these, and the current plan violates both. **Fonts (#74).** Han, Devanagari, Bengali, Tamil, Thai, Khmer and Ge'ez need glyphs that ship — #74's title says it: *a declared package is not a drawn glyph*. Mandarin has more speakers than anything else here and still waits, because a page of tofu boxes is worse than the same page in English. **Right-to-left (#73, #67).** Arabic, Urdu, Persian and Hebrew need the layout before the catalogue. **`ar` is currently in #70 at 0.4.0, while the RTL layout work is at 0.6.0** — the first right-to-left language is scheduled two milestones before the thing it depends on. That is a bug in the plan regardless of which ordering wins. ## The tiers Speaker figures are approximate and include second-language speakers; they order the list, they are not a promise. ### Tier 1 — 0.4.0. Nothing blocks these. Latin or Cyrillic, no new font work, no RTL. - **Russian** ~255M — new catalogue. Cyrillic is already proven by `bg`, `mk`, `sr`, `uk`. - **Indonesian** ~200M — new. - **Vietnamese** ~85M — new; Latin, heavy diacritics. - **Filipino / Tagalog** ~83M — new. - **Hausa** ~80M and **Swahili** ~72M — `ha` and `sw` already scaffolded. - The rest of the Latin-script African set already scaffolded: Yoruba, Igbo, Zulu, Xhosa, Afrikaans, Somali, Oromo, Kinyarwanda, Shona, Sesotho, Setswana, Wolof, Malagasy and the others. ### Tier 2 — 0.5.0, gated on #74. - **Mandarin Chinese** ~1.1B — simplified and traditional are two catalogues. - **Hindi** ~610M, **Bengali** ~280M. - **Japanese** ~125M, **Korean** ~82M. - **Marathi, Telugu, Tamil, Punjabi, Gujarati, Kannada.** - **Thai** ~60M. - **Amharic** and **Tigrinya** — Ge'ez script; carved out of #70, which cannot reach them before the fonts land. ### Tier 3 — 0.6.0, gated on #73 and #67. - **Arabic** ~400M — carved out of #70, where it is currently scheduled ahead of its layout. - **Urdu** ~230M, **Persian** ~80M, **Hebrew**. - **Tamazight / Kabyle** (`kab`) — also carved out of #70. ### 1.0.0 - #101, the PALOP national languages, unchanged — it is a commitment about honesty regarding machine translation, not a place in this ordering. - The long tail as a *process* rather than a list: #72's second half, which is the part of that issue worth keeping. ## What changes on the three existing issues - **#70** stays at 0.4.0 and narrows to the **Latin-script** African languages. `ar`, `kab`, `am` and `ti` leave it for Tiers 2 and 3, so it can close when the ones that are not blocked are done. - **#71** moves to **0.5.0** — its Asian catalogues are the font-gated tier. Its RTL members (`ur`, `fa`, `he`) move to Tier 3. Quechua, Guaraní and Aymara are Latin-script and may be pulled into Tier 1 by whoever starts it. - **#72** moves to **0.6.0**, and should be re-read as the *process* issue it half already is: a speaker adding a language without waiting for a developer. That is what makes the tail reachable at all, and Weblate is already running for it. ## What does not change - A language is `draft` until a speaker reviews it, and stays out of `language_choices` until somebody starts it. - The working set while developing is still en / fr-fr / pt-pt, with the full sweep on a release commit. - `tests/test_translations.py::test_every_european_union_language_stays_complete` still hard- gates the 24 EU languages on every commit, and its conflict with the three-language working rule is still unresolved. This issue does not fix that and should not be read as having done so.
tiagoagueda added this to the 0.4.0 milestone 2026-09-17 19:43:28 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
Postulo/postulo#269
No description provided.