Virgula de sub Ș

75 websites · 87 font families · July 2026

The Comma Below the S

Romanian writes ș and ț with a comma below. In 1987 the encoding standards handed it the cedilla instead, the hook used by French and Turkish, and in 1999 Unicode corrected the error and left a one-sentence instruction next to the old character.

Twenty-seven years later, 2,981 of the 26,861 letters in that family served by 75 Romanian-language pages are still the Turkish letter. In central government the proportion rises to 13.2 per cent. The Senate spells its own Constituţia with a cedilla eight times on its front page, Bucharest City Hall spells Bucureşti nine times, and the National Institute of Statistics writes Naţional with a cedilla twelve times.

The uncomfortable half of the measurement comes next. On most of those pages the error cannot be seen. The font repairs it, every time the page declares its language, and nobody has any way of noticing what the memory actually holds.

11.1%of the Ș and Ț letters served by the 75 pages measured are the cedilla form
13.2%in central public administration, the highest of the four groups
1.0%in the independent newsrooms, the group with the cleanest encoding
59 of 84font families that hide the error completely on a page that declares its language

Two letters that look the same

A cedilla grows out of the foot of the letter and stays attached to it, as in the French ç. The comma below is a separate mark, set under the letter without touching it. At newspaper size the difference is a few hundredths of a millimetre. At headline size you can see it from across the room.

For a program the distance between them has nothing to do with the drawing. They are two different numbers, two different characters, as unrelated as S and 5.

latin small letter s with comma below
code
U+0219
mark
U+0326 combining comma below
languages
Romanian, Latvian, Livonian
added
Unicode 3.0, 1999
latin small letter s with cedilla
code
U+015F
mark
U+0327 combining cedilla
languages
Turkish, Azerbaijani, Gagauz
added
Unicode 1.1, 1993

The same letter, the same size, two characters

the character 0219 should be used instead for Romanian

The annotation at U+015F, Unicode Character Database

The sentence has been in the standard since 1999, three lines from the character that 32 of the 36 central government sites still serve.

How the cedilla got into Romanian

The error began as a saving of space in 1987, was confirmed by documentation in the nineties, corrected in 1999 and delivered to the user in 2007. Every step took longer than the one before it.

  1. 1987Latin-2

    Romanian is given a character set that does not contain it

    ISO/IEC 8859-2, known as Latin-2, divides 96 positions among the languages of Central Europe. Romanian is on the list, and for the sounds written ș and ț it is handed the positions occupied by S and T with cedilla, the hooked mark used by French and Turkish. Cedilla and comma end up in the same drawer because they look alike at small print sizes and because space is expensive.

  2. 1993Unicode 1.1

    The four wrong letters get permanent addresses

    Unicode 1.1 fixes Ş, ş, Ţ and ţ at U+015E, U+015F, U+0162 and U+0163. The same edition encodes the comma below at U+0326 as a combining mark, so the mark Romanian needs enters the standard at the outset. The letters carrying it do not, and the cedilla forms remain the only thing a system can type, store and transmit.

  3. 1995The documentation

    The standard states in writing that the cedilla will do for Romanian

    In the Unicode 1.1.5 documentation the four cedilla characters are described as suitable for both Turkish and Romanian. Font vendors and operating system vendors read exactly what suits them: one form, one drawing, no extra work. Romanian orthography, which has required a comma since the nineteenth century, has no standing to contest a technical note.

  4. 1998The rejected revision

    The request to correct Latin-2 closes with a footnote

    The revision of ISO/IEC 8859-2 passes without the amendment the Romanian side requested. The final text merely notes that the cedilla letters may be used as a substitute for the comma letters. That sentence turns the substitute into the official form for another decade of documents.

  5. 1999The correction

    Unicode 3.0 adds the correct letters and says which they are

    The block U+0218 through U+021B enters the standard: Ș, ș, Ț, ț, letters with a comma below. Next to the old U+015F, Unicode attaches a one-sentence annotation: the character U+0219 should be used instead for Romanian. In the same year the Romanian standard SR 13411 fixes the comma form for the national character set.

  6. 2001Latin-10

    A European character set adopts the correct form

    ISO/IEC 8859-16, Latin-10, includes the comma letters. It arrives at a point when the world is already moving to Unicode, so its practical reach is small. Its value lies elsewhere: from 2001 nobody can claim that the international standards have nowhere to put the letter.

  7. 2003The Academy

    The Institute of Linguistics puts the orthographic rule in writing

    At the request of the group revising the keyboard standard, the Institute of Linguistics of the Romanian Academy confirms in writing that the mark below S and T is a comma. The document closes the linguistic question. The technical one remains, and that depends on vendors.

  8. 2004The keyboard

    The Romanian keyboard standard moves to the correct codes

    SR 13392:2004 defines two layouts and ties them explicitly to the comma-below Unicode codes. A keyboard standard only takes effect if the operating system ships it, and the dominant operating system still did not.

  9. 2007The delivery

    The fonts and the keyboard finally reach the user

    Microsoft releases the European Union Expansion Font Update, adding the comma letters to Times New Roman, Arial, Trebuchet and Verdana for Windows XP SP2, Windows Server 2003 and Windows Vista. Vista makes the correct layout the default and renames the old one Romanian (Legacy). Romania had joined the European Union eight months earlier.

  10. 2026The measurement

    The Turkish letter is still in circulation

    Twenty-seven years after the correction, 2,981 of the 26,861 letters in the Ș and Ț family served by 75 Romanian sites are still the cedilla form. In central government the proportion rises to 13.2 per cent. On most of those pages the error is invisible.

The bench

Three switches decide what the reader sees: the character in memory, the language the page declares for itself, and the font family. Only one of them has anything to do with what the page actually says.

Two wrong letters, three switches

The bench starts with “Şedinţe publice” misspelled twice, once with Ş and once with ţ, in a very widely used font, on a page that declares its language as Romanian. Look at both letters before you touch anything.

Şedinţe publice

What the memory holdsU+015EU+0065U+0064U+0069U+006EU+0163U+0065U+0070U+0075U+0062U+006CU+0069U+0063U+0065

The characters in memory
The language the page declares
Font family

distance to the letter 51 thousandths of the em · repairs only some of the letters

Result

The font repairs Ş and leaves ţ as it is. The same error, in the same word, hides on one letter and shows on the other. Eleven of the 84 families behave this way, Roboto among them.

Seventy of the 84 families that have the letter declare this substitution, but only 59 carry it through. Eleven repair S and leave T untouched, so a word like “Şedinţe” comes out half corrected. It is a well-meant typographic courtesy, and it is the reason the error survived a quarter of a century without inconveniencing anyone.

What the sites serve

On 29 July 2026 I requested the front page of 82 websites publishing in Romanian once each and counted, in their text, how many letters of the Ș and Ț family arrive with a comma and how many with a cedilla. Seventy-five answered with a page.

The bar shows the proportion for each. To the right is the cedilla word that site repeats most often, which is usually the way the institution spells its own name.

commacedillacomma share · most frequent cedilla word

Public administration

ministries, autonomous authorities, courts, city halls of county seats 13.2% cedilla overall · 36
Ministerul Afacerilor Externe 51.0% Evitaţi ×16
ANAF 60.3% şi ×7
Ministerul Agriculturii 77.4% şi ×7
Curtea Constituțională 79.6% Curţilor ×6
Primăria Capitalei 81.0% Administraţia ×9
Ministerul Transporturilor 82.6% Anunţuri ×19
Avocatul Poporului 85.5% şi ×8
Ministerul Muncii 86.3% şi ×7
Poliția Română 86.4% Citeşte ×14
Autoritatea Electorală Permanentă 88.6% cunoştinţa ×6
Ministerul Culturii 91.3% şi ×4
Guvernul României 91.4% INFORMAŢIE ×4
Ministerul Finanțelor 91.5% Noutăţi ×3
Autoritatea de Supraveghere Financiară 91.5% Autorităţii ×3
Ministerul Sănătății 92.4% şi ×2
ANPC 92.4% DETERGENŢI ×4
Senatul României 92.8% şi ×18
ANCOM 93.3% Legislaţie ×4
Primăria Iași 94.4% Iaşi ×3
Ministerul Mediului 94.8% situaţii ×4
Banca Națională a României 96.8% Bucureşti ×1
Ministerul Educației 97.3% învăţământul ×2
Monitorul Oficial 97.4% Alimentaţia ×1
ANCPI 97.8% Bucureşti ×1
Ministerul Economiei 98.0% foloseşte ×1
Primăria Cluj-Napoca 98.7% şi ×4
IGSU 100.0%

General-interest press

wire services, broadcasters, dailies and business publications with national reach 11.6% cedilla overall · 22
Ziarul Financiar 10.1% şi ×121
News.ro 28.9% şi ×30
TVR Info 66.0% şi ×11
Digi24 76.2% şi ×6
G4Media 80.7% şi ×16
Antena 3 CNN 82.5% şi ×10
Dilema veche 82.8% şi ×7
Agerpres 87.9% Ştirile ×3
Știrile ProTV 88.5% şi ×4
Cotidianul 88.6% aşteptări ×6
Economedia 92.0% Plăcuţele ×2
Observator cultural 94.1% şi ×83
HotNews 94.4% Aşa ×2
Adevărul 94.5% şi ×5
Gândul 96.8% Curţii ×2
Libertatea 97.4% şi ×3
Mediafax 98.3% Justiţie ×4
Evenimentul Zilei 98.4% eşuat ×1
Profit.ro 98.9% Piaţa ×2
Biziday 99.2% foloseşte ×1
Spotmedia 99.3% Naţional ×1
Ziare.com 99.9% ştiri ×1

Independent newsrooms

investigative publications funded by donations and grants 1.0% cedilla overall · 9
Context.ro 93.4% şi ×4
DoR 97.5% condiţii ×1
Buletin de București 98.8% sancţiuni ×3
PressOne 99.3% Şercan ×3
Misreport 100.0%
Panorama 100.0%
Recorder 100.0%
Scena9 100.0%
Snoop 100.0%

Control group

publishers based outside Romania who publish in Romanian 10.1% cedilla overall · 8
RFI România 50.9% şi ×10
Europa Liberă România 99.2% Roşia ×1
dexonline 100.0%

Where the cedilla retreated to

The word most often written with a cedilla is “şi”, the commonest conjunction in the language: it appears that way on 46 of the 75 pages and is the most repeated cedilla word on 29 of them. After it come the fittings of a website: “Anunţuri” in the menu of five institutions, “Citeşte” on the read-more button of four publications, “Bucureşti” and “Administraţia” in the masthead of Bucharest City Hall. On gov.ro, madr.ro and economie.gov.ro the very same pair turns up, “foloseşte” and “exprimaţi”, the text of a single cookie notice copied from one place to another. The wrong letter has left today’s article and settled in the labels nobody ever rewrites.

On the Senate front page, two of the 93 cedillas are legitimate: the name of Numan Kurtulmuş, a Turkish politician. One more foreign cedilla word exists in the whole sample, “Tolışi” on Wikipedia. The rest are Romanian words. The 1987 collision shows up twice, exactly as it was designed.

Seven sites did not answer

They appear in no total. The Chamber of Deputies, the High Court, the Ministry of Justice, the Ministry of Development and the legislative portal answered no request at all, browser included. The Trade Register rejected the request, and Casa Jurnalistului never finishes loading its content.

The fonts fixed themselves

I downloaded the files Google Fonts serves for 87 families and read them from the inside: whether the letter exists, what it is built from, how far the mark falls, and whether the family carries a Romanian localisation.

The result is unexpectedly good. All 84 families that have the letter draw it with a real comma, built from the mark at U+0326, rather than a cedilla moved further down. The typographic layer of the problem is solved. What is not solved is the text.

3 of 87

do not have the letter at all

Lato loses it completely: the latin-ext subset Google serves has 47 glyphs and no Romanian diacritics whatsoever. Righteous has ş with a cedilla and lacks ș with a comma, so it rewards precisely the error. Prata has no latin-ext subset, so all its diacritics come from another font.

87 of 87

lack the letter in the latin subset

The latin subset stops at U+00FF in every family measured. The letter arrives only from latin-ext, and a page that subsets fonts itself to save weight ends up without it.

59 of 84

silently repair the wrong text

The family declares a localised substitution for Romanian and swaps the cedilla glyph for the comma glyph when the page says the text is Romanian. Seventy families declare the feature, but only 59 cover all four letters. Eleven repair S and leave T untouched, Roboto among them, the most used font on the web: on a Romanian page “şi” corrects itself and “ţara” does not.

How far the mark falls

The distance between the foot of the letter and the top of the comma, in thousandths of the em. The median across the 84 families is 46, which is four and a half pixels in a 96-pixel headline. At 94 the comma starts reading as punctuation that wandered onto the line below.

median 46 Roboto: 51Open Sans: 51Noto Sans: 51Montserrat: 89Poppins: 40Inter: 46Roboto Condensed: 51Oswald: 55Raleway: 39Nunito: 66Nunito Sans: 71Ubuntu: 38Rubik: 40Mulish: 71Work Sans: 66Source Sans 3: 46PT Sans: 43Barlow: 62Karla: 36Fira Sans: 65DM Sans: 38Manrope: 94Figtree: 48Outfit: 69Plus Jakarta Sans: 58Space Grotesk: 47Archivo: 75Public Sans: 38Cabin: 81Josefin Sans: 34Quicksand: 22Titillium Web: 68Heebo: 76Exo 2: 48Asap: 42Chivo: 59Overpass: 33Sora: 34Urbanist: 40Bricolage Grotesque: 45Instrument Sans: 62Geist: 41Onest: 64Playfair Display: 55Merriweather: 50Lora: 24PT Serif: 35Noto Serif: 51Libre Baskerville: 43EB Garamond: 43Cormorant Garamond: 74Crimson Text: 34Bitter: 61Literata: 42Source Serif 4: 39Fraunces: 53Bodoni Moda: 94Libre Bodoni: 67DM Serif Display: 38Instrument Serif: 58Newsreader: 42Petrona: 68Spectral: 65Old Standard TT: 28Playfair: 45Bebas Neue: 33Anton: 45Lobster: 58Pacifico: 45Dancing Script: 39Caveat: 30Comfortaa: 34Alfa Slab One: 46Abril Fatface: 44Archivo Black: 47Teko: 50Big Shoulders Display: 33Unbounded: 33Roboto Mono: 22JetBrains Mono: 87IBM Plex Mono: 50Space Mono: 32Source Code Pro: 46Inconsolata: 21 Inconsolata 21 Bodoni Moda, Manrope 94 20406080100 thousandths of the em

Four of the families, Roboto among them, still build the letter from a glyph named uniF6C3. The name comes from the private use area of the Adobe glyph list of the nineties, put there because the list refused the Unicode assignment. To this day the correct comma travels under a code invented to route around an error.

Three layers, three answers

The same pair of words, Constituția spelled correctly and Constituţia spelled wrongly, passes through three systems that judge it differently. All three answers were verified in the browser you are reading this in.

The eyealmost nothing

If the page declares its language as Romanian and the font carries a complete localisation, the two words look identical. Fifty-nine of the 84 families measured behave this way, Open Sans, Montserrat and Playfair Display among them. Another eleven repair some of the letters, so they leave a visible trace at random.

The sorted listsees nothing

Romanian collation treats cedilla and comma as equivalent, at every level of sensitivity. A correctly ordered list of names gives no sign that half of it is spelled with the wrong letter.

The comparisonsees two words

String equality does not use collation, it uses characters. U+021B is not U+0163, and canonical decomposition does not bring them closer: one yields s followed by U+0326, the other s followed by U+0327. The compatibility form does not unify them either.

checkresult
'și' === 'şi'false
'și'.normalize('NFD')U+0073 U+0326 U+0069
'şi'.normalize('NFD')U+0073 U+0327 U+0069
NFKD('și') === NFKD('şi')false
new Intl.Collator('ro').compare('și', 'şi')0
new Intl.Collator('ro', { sensitivity: 'variant' })0
new Intl.Collator('tr').compare('și', 'şi')-1
/și/i.test('şi')false

A state keeps its records in the third layer. That is where exact filters live, and joins between two registers on the name field, and deduplication passes, and the regular expressions in forms, and the comparisons inside database migrations. A citizen whom one institution spelled with a cedilla and another with a comma is, to any system that puts the two together without collation, two people with the same name.

That is why the error held. It produces no ugly page, raises no alert and troubles no reader. It costs only in the places nobody looks at.