FontGenerator research · Version 2026-10-02-v1

Unicode Fonts: A 54-Language Character Reference

A disclosed character-inventory reference for checking language requirements. No actual font file was tested in this audit.

Start with the text you need to set

Choose the languages and real text you intend to publish. Check the required characters against a font’s character map, then inspect shaping and positioning in the application that will render the text. A character count alone cannot certify usable language support.

The Unicode font FAQ discusses Unicode and font coverage. These measurements concern this project’s disclosed inventory and baseline, not a catalogue of font recommendations.

Main characters, auxiliary characters and shaping

CLDR distinguishes main exemplars from auxiliary exemplars. FontGenerator’s pinned data applies case closure and keeps those inventories separate. In this checker, full means all main and auxiliary code points are present; core means all main code points are present but an auxiliary character is missing; and missing means a main character is absent.

These are set-comparison labels, not quality grades. Shaping is a separate step: code-point presence does not prove glyph design, GSUB, GPOS, mark placement, joining, ligatures or browser layout.

The ASCII baseline

We supplied U+0020–U+007E: 95 code points including ASCII letters, numbers, punctuation and space. We did not load a font. For every language assessed by the pinned checker, we independently subtracted that inventory from the main and auxiliary sets, then compared the results with the checker function.

Results of the synthetic ASCII character-inventory baseline
Checker outcomeLanguagesInterpretation
Full2: Malay ms, Swahili swMain and auxiliary sets are present
Core3: English en, Indonesian id, Zulu zuMain set is present; auxiliary gaps remain
Missing49At least one main-set code point is absent
Not assessed2: Arabic ar, Hindi hiExcluded from this character check

Citable finding. Against U+0020–U+007E, the pinned FontGenerator checker assesses 54 language inventories: 2 meet its main-plus-auxiliary criterion, 3 meet its main-only criterion, and 49 miss at least one main-set code point. This is a synthetic character-inventory test, not an evaluation of actual fonts.

English has 52 pinned main and 76 auxiliary code points, so ASCII meets its main set but misses auxiliary characters. German has seven non-ASCII main characters (Ä Ö Ü ß ä ö ü); Spanish misses 14 main characters under this baseline and French misses 32.

Worked inventory examples

Worked inventory examples from the disclosed language data
LanguageMainAuxiliaryMain missing from ASCIILabel
en52760core
de59767missing
es666614missing
fr844432missing
sw4840full
ms5200full

These are inventory counts and synthetic-baseline outcomes, not measured character counts from a font file.

What this check covers

The local inventory contains 56 language entries derived from CLDR 48.2.0, commit bb334e8d6250c9363e957e131bf7e6d08ec72f91. The pinned character checker returns results for 54. Arabic and Hindi remain in the disclosed dataset but are marked not assessed, so their character lists are not silently lost.

Citable finding. The assessed main-character union contains 611 distinct code points; including auxiliary characters raises that union to 817. These counts describe the 54 assessed entries after the project’s case closure, not the whole CLDR database or the glyph inventory of a font.

The two exclusions do not certify the other 54 languages’ shaping. Precomposed characters and decomposed sequences must not be assumed interchangeable merely because they look similar in a particular application.

Use the reference with an actual font

  1. Find your language’s disclosed row and inspect the main and auxiliary characters.
  2. Use the existing language check tool when checking character presence.
  3. Test the real words, names and punctuation needed by your document in its target application.
  4. Record the font version and rendering environment separately from this inventory version.

A complete map is a reason to proceed to rendering checks, not a guarantee that those checks will pass. This reference does not choose fallback fonts or score typographic quality.

Downloads and reproduction

Save the JSON files and script together, then run python3 reproduce.py. It independently recalculates every disclosed ASCII gap, status and union count using the Python standard library. Neither this operation nor the original run renders a font.

CLDR-derived character data is attributed to Unicode, Inc.; the supplied licence accompanies this reference. This first-party audit is limited to a curated local inventory and a pinned working-tree snapshot; a later checker version can differ. For citation, name FontGenerator, this page title, version 2026-10-02-v1 and the relevant result section, and describe the baseline as synthetic.