All PALOP national languages — and an honest account of what a machine can translate #101
Labels
No labels
accessibility
authentication
breaking change
bug
documentation
enhancement
interface
internationalisation
observability
security
tier
1
tier
2
tier
3
tier/4
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Depends on
#70 Languages: every language of Africa
Postulo/postulo
Reference
Postulo/postulo#101
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observation
What "all languages from PALOP" has to mean
The PALOP states are Angola, Cabo Verde, Guinea-Bissau, Mozambique and Sao Tome and
Principe. The official language of all five is Portuguese, and Postulo has spoken that
since 0.2.0 -- so read literally the request is already done, which is plainly not what is
being asked.
What is being asked for is the national and vernacular languages of those five countries:
the creoles and the Bantu languages people actually speak at home, none of which any
job-search tool ever offers.
Five of them are already there
#70 added twenty-nine African languages to the
0.3.0branch, and five happen to be PALOPnational languages:
ffPulaar (Guinea-Bissau),tsXitsonga (Mozambique),nyChichewa(Mozambique),
snchiShona (Mozambique),swKiswahili (northern Mozambique).What is missing
umbkmbkoncjkmckkuakeapovmnkblevmwsehnglchwkderngtscccecriaoapreTwenty-one, which is scopeable. The flag column applies the rule
languages.pyalreadystates -- "a language with no uncontested home gets no flag at all. No flag beats a wrong
flag" -- and it lands well here: most of these belong to exactly one country, which is
unusual and convenient.
One scope question. Equatorial Guinea has Portuguese as an official language and joined
the CPLP in 2014, but is not conventionally PALOP. Including it adds Fang, Bube and
Annobonese. Say which.
The part that needs saying plainly: an LLM cannot do most of this
The project already permits machine-assisted translation --
docs/TRANSLATING.mdgives everyunreviewed string a
draftflag until a speaker sees it -- so this is not a policy change.It is a question of what a machine will actually produce, and the answer differs enormously
across that table:
translation systems. A draft is worth having.
Tshwa, Copi have very little written text online. A model asked for these will not say
it does not know: it will return Portuguese with altered morphology, or invented
words, fluently and confidently.
critically endangered, with almost nothing written. There is no honest machine output for
these at all.
This matters more here than in most software. A wrong label on Delete account or Withdraw
application in a tool holding somebody's job search is not a cosmetic bug, and a fluent
wrong translation is harder to spot than an obvious one.
languages.pyalready reasons thisway about flags; the same principle applies to words.
The proposal, which gets the request without the harm
Add all twenty-one to the picker. This part is cheap, correct, and most of the point:
somebody finding Umbundu in a list is being told this software considers their language
real. #70 has already established that a language may be offered with an empty catalogue.
Fill only what a machine can honestly do, every string flagged
draft, and record perlanguage which of the three tiers above it fell into -- so a reviewer knows whether they
are checking a draft or writing the first version.
Fall back to Portuguese, not English. This is the most valuable part and it is not
currently possible. An untranslated string falls back to
en-gb, the source language. Butevery one of these languages is spoken in a country whose official language is
Portuguese, which Postulo already speaks fully. A Cape Verdean seeing Portuguese where
Kabuverdianu is missing is far better served than one seeing English. A per-language
fallback chain would make a half-filled catalogue genuinely usable instead of a patchwork
of two languages neither of which is the one asked for.
Django does not do arbitrary fallback chains --
get_supported_language_variantonlywalks to the generic form -- so this is real work, and it is the piece worth building even
if the catalogues stay empty for years.
Plural rules need care. Every language needs a
PLURAL_FORMSentry, and a wrong onemakes every count in the interface ungrammatical. Most Bantu languages take
nplurals=2; plural=(n != 1);and the creoles likewise, but this should come from CLDRwhere CLDR has it rather than from a guess, and be marked as unverified where it does not.
Scope
NATIVE_NAMES,FLAG_COUNTRIESandPLURAL_FORMS, on the0.3.0branch where #70's work lives.
draft.machine draft for a reviewed translation.
Classification
Enhancement, internationalisation. Depends on #70 for the machinery and shares its branch.
Moved to 1.0.0. The scale is the reason: twenty-one national languages, most with little or no digital corpus, and three of them — Angolar, Principense, and arguably Cokwe — with no honest machine output at all. That is a long piece of work with real speakers in the loop, and pinning it to 0.3.0 would either delay that release or produce catalogues nobody has read.
The fallback work should not wait for it. The most valuable part of this issue is that an untranslated string in a PALOP language should fall back to Portuguese rather than English — every one of those countries has Portuguese as an official language, and Postulo speaks it fully. That is worth building even while the catalogues stay empty, and it is what makes a half-filled catalogue usable instead of a patchwork of two languages neither of which was asked for.
Brazilian Portuguese is filed separately for 0.3.0: it is a different kind of job entirely — a variant of a language Postulo already speaks completely.