Romanian diacritic marks

 · 
March 23, 2023
 · 
30 min read

By Cristian Kit Paul, co-founder and Creative Director of Brandient. Originally published 27 October 2008 · Revised and extended July 2026.

Note to the reader (2026). This article was first published on October 27, 2008, and its original timeline stopped in 2013. It is preserved below as written, with corrections marked, because it remains an accurate forensic record of how Romania's letters got mangled between 1987 and 2013. What has changed is the conclusion. The encoding war described here has been won on the technical side: every major platform now inputs and renders ș and ț correctly by default. The 2019–2026 extension of the timeline, and the new sections that follow it, tell the second act—in which the enemy is no longer the lack of core standards and technical limitations but the delivery pipeline: fonts subset carelessly for the web, autocorrect on phones, identity documents that strip diacritics by design, and artificial intelligence that restores them unreliably.

And a word on why I keep returning to this, eighteen years on: design is a gateway profession—whatever we let pass becomes the culture—and for two decades this story had a technical excuse. The excuse is gone; what remains is Papanek's definition of the job itself: "Design is a conscious and intuitive effort to impose meaningful order." This article supplies the knowledge that makes the conscious effort possible. Imposing the order is the designer's social responsibility—my peers' and mine.

In brief: Romania's five special letters—ă, â, î, ș, ț—spent three decades mangled by wrong standards (cedilla instead of comma below), missing fonts, and absent keyboards; the 1987–2026 timeline below reconstructs that failure actor by actor, from ISO and Unicode to Microsoft, Apple, and Adobe. The technical war is over: every major platform now types and renders the correct letters by default. What remains broken is the delivery pipeline—web fonts subset without latin-ext, identity documents that strip diacritics by ICAO design, autocorrect habits, and AI models that restore diacritics imperfectly because they were trained on three decades of unaccented Romanian text. The rules for writing Romanian correctly, and a checklist for implementers, close the article.


How did we end up looking uncivilized and half-illiterate?

Lack of standards, wrong standards, confusing workarounds, and then a very slow adoption of good standards—no wonder the diacritics turned into a puzzle that people are too tired to try solving by hacking fonts to add the missing glyphs or hunting obscure code points in the character map. Magazine headlines, television supers, and advertisements obliviously used incorrect Romanian spelling, or non-Romanian letters.

The situation was so bad that even on national banknotes the spelling was bastardized—if you know who designed this, please encourage the person(s) to quit design and pursue a career in agriculture:

The banknote outrage above dates from the 2005 polymer series.1 On December 1, 2021, the National Bank of Romania put into circulation a new 20-lei note featuring Ecaterina Teodoroiu—the first circulating Romanian banknote to depict a woman—and the caron is still there, in the denomination itself: the note reads LEI DOUǍZECI (A-caron) where Romanian requires DOUĂZECI (A-breve). Two decades, a full redesign opportunity, a historic first, and the wrong letter spelling out the note's own face value. The career in agriculture remains available—these people are clearly not concerned with design matters.

What happened? How did we end up looking like idiots?

An introduction to diacritics

Romanian glyphs

In the sense of diacritics as signs added to letters to alter their pronunciation or to distinguish between words, the Romanian alphabet does not have diacritics. There are, however, five special letters in the Romanian alphabet (two of them associated with the same sound), formed by modifying other Latin letters. Strictly speaking they are not diacritics but distinct letters of the alphabet—with their own code points, their own place in collation, and their own legal standing in orthography. They are, nonetheless, generally referred to as diacritics, and this article follows that usage:

Anatomical diagram of the letters Ă and ă, highlighting the breve above.

Ă / ă · Upper: U+0102 · Lower: U+0103 · Unicode name (capital): LATIN CAPITAL LETTER A WITH BREVE · Block: Latin Extended-A · Since: Unicode 1.1 (1993) · NFD decomposition: A/a + U+0306 COMBINING BREVE · UTF-8 (hex): C4 82 / C4 83 · AGL glyph name: Abreve / abreve

Anatomical diagram of the letters  and â, highlighting the circumflex above.

 / â · Upper: U+00C2 · Lower: U+00E2 · Unicode name (capital): LATIN CAPITAL LETTER A WITH CIRCUMFLEX · Block: Latin-1 Supplement · Since: Unicode 1.0 (1991) · NFD decomposition: A/a + U+0302 COMBINING CIRCUMFLEX ACCENT · UTF-8 (hex): C3 82 / C3 A2 · AGL glyph name: Acircumflex / acircumflex

Anatomical diagram of the letters Î and î, highlighting the circumflex above.

Î / î · Upper: U+00CE · Lower: U+00EE · Unicode name (capital): LATIN CAPITAL LETTER I WITH CIRCUMFLEX · Block: Latin-1 Supplement · Since: Unicode 1.0 (1991) · NFD decomposition: I/i + U+0302 COMBINING CIRCUMFLEX ACCENT · UTF-8 (hex): C3 8E / C3 AE · AGL glyph name: Icircumflex / icircumflex

Anatomical diagram of the letters Ș and ș, highlighting the comma below.

Ș / ș · Upper: U+0218 · Lower: U+0219 · Unicode name (capital): LATIN CAPITAL LETTER S WITH COMMA BELOW · Block: Latin Extended-B ("Additions for Romanian") · Since: Unicode 3.0 (1999) · NFD decomposition: S/s + U+0326 COMBINING COMMA BELOW · UTF-8 (hex): C8 98 / C8 99 · AGL glyph name: Scommaaccent / scommaaccent

Anatomical diagram of the letters Ț and ț, highlighting the comma below.

Ț / ț · Upper: U+021A · Lower: U+021B · Unicode name (capital): LATIN CAPITAL LETTER T WITH COMMA BELOW · Block: Latin Extended-B ("Additions for Romanian") · Since: Unicode 3.0 (1999) · NFD decomposition: T/t + U+0326 COMBINING COMMA BELOW · UTF-8 (hex): C8 9A / C8 9B · AGL glyph name: none since AGL 2.0—use uni021A / uni021B

Although we only have five diacritics (Czech has fifteen), we sometimes manage to get three of the five wrong. Most of the time we get two of those three wrong: Ș and Ț.

Common mistakes

The most common mistakes plaguing written Romanian are the following glyph substitutions:

  • Letter A with tilde (ã) or letter A with caron (ǎ) instead of letter A with breve (ă).
  • Letter S with cedilla (ş) instead of letter S with comma below (ș). S with cedilla is a Turkish letter.
  • Letter T with cedilla (ţ) instead of letter T with comma below (ț). [Corrected 2026:] T with cedilla is not used by any major living language and no language needs it—though, for the record, it does appear in the General Alphabet of Cameroon Languages, in some Gagauz orthographies, and in local Kabyle usage, alongside its original role in Semitic transliteration. The point stands: it has no business anywhere near Romanian.
Large à and ã glyphs — the letter A with a tilde.
Large Ǎ and ǎ glyphs — the letter A with a caron.
Large Ş and ş glyphs — the letter S with a cedilla.
Large Ţ and ţ glyphs — the letter T with a cedilla.

The one thing to remember: Ş and ş, Ţ and ţ are not styling variants of the Romanian letters—they are different characters—one Turkish, one belonging to no language at all!—with different code points (U+015E–0163 versus U+0218–021B). Using the cedilla forms does not merely look wrong; it breaks search, sorting, spell-checking, and matching in every system that reads the text. Wrong letter and no letter are both data corruption—one substitutes a foreign character, the other amputates a Romanian one.

While the first mistake is caused mainly by indolence, the second and third have an epic story behind them and deserve a closer look.

But first,

A note on â and î (the reform nobody explains)

The pair â/î deserves its own footnote, because it is the one place where the difficulty is orthographic convention rather than typography: both letters denote the same sound, the close central vowel /ɨ/, which has no English equivalent—so a foreign reader should watch position, not pronunciation. The rule is purely geographic: î at word edges, â everywhere inside.

The 1953 reform of the Communist era—decreed that September, in force from April 1954—abolished â almost entirely, imposing î everywhere; the spelling sînt ("I am"/"they are"), though older than the reform, became its emblem.2 A 1965 concession restored â in român ("Romanian") and its derivatives, for obvious reasons of national vanity—the regime could stomach sînt but not Romînia.

On February 17, 1993, the Romanian Academy reversed course: â inside words (român; mână, "hand"), î at word edges (început, "beginning"; a coborî, "to descend"), and the forms of the verb "to be" restored to sunt, suntem, sunteți ("I am/they are," "we are," "you are"). The boundary rule survives prefixation intact: împăcat ("reconciled") keeps its î in the compound neîmpăcat ("unreconciled"), because the î still marks the seam of the original word—which is exactly the sort of thing one must know the language's morphology to get right. DOOM2 (2005) and DOOM3 (2021) reaffirmed the rule, and it has been stable ever since—which makes it all the more telling that, as we shall see, misplacing â versus î is today the single largest error class when artificial intelligence restores Romanian diacritics. The machines fail exactly where the rule is arbitrary or obscure.

The story of Ș and Ț: an epic clusterfuck

When I say epic, I really mean exactly that. Imagine everything going atrociously wrong—for thirty years! Here is what happened.

The timeline, part one: 1987–2013

ISO 1987 — Romanian is associated with ISO 8859-2 (Latin 2)—the international standard stipulates S-cedilla and T-cedilla glyphs. Romanian officials are oblivious to the matter. Very, very bad.

UNICODE 1995 — The Unicode consortium specifies in version 1.1.5 the code points U+015E (Latin Capital Letter S with Cedilla), U+015F, U+0162 (Latin Capital Letter T with Cedilla), and U+0163 as suitable for both Turkish and Romanian, defining them as carrying the cedilla accent. Turkish indeed uses U+015E/F but makes no use of U+0162/3. Romanian uses none of them. Very bad.

MICROSOFT 1995 — Windows 95 launches with no Romanian support by default. Support arrives on the CD-ROM Extras for the Windows 95 Upgrade. The typeface ILP Rumanian B100 substitutes Q/q with Ă/ă. No, read that again: a “Rumanian” typeface substitutes Q with Ă. Dark ages! Moronically bad.

APPLE 1997 — Apple's Mac OS 7.6.1 honors Romanian S/s and T/t with comma below via MacRomanian (ten years before Microsoft). Interestingly, its tables do not resolve U+015E/F or U+0162/3—wow!—no cedilla forms at all. Good.

ADOBE 1997 — The Adobe Glyph List (AGL 1.0, July 17, 1997; AGL 1.1, November 24, 1997)3 maps U+0162/3 to Tcommaaccent/tcommaaccent and U+015E/F to Scommaaccent/scommaaccent, with Scedilla/scedilla parked in the private-use area. Both letters carry comma-below names, so Adobe fonts render the legacy code points consistently—comma below on both. Good—for now.

ASRO 1997 — It takes ten years for ASRO, Romania's national standardization body, to react. In 1997 the association complains to ISO about the cedilla standardization, requesting an amendment. Good.

ISO 1998 — The revised ISO/IEC 8859-2 is approved without the requested amendment (published as ISO/IEC 8859-2:1999). A note concedes—paraphrasing—that the cedilla letters of Part 2 may be used to substitute for Romania's comma-below letters, "subject to the agreement of originator and receiver." Very bad.

ADOBE 1998 — AGL 1.2 (October 22, 1998) reverses only the S: U+015E/F flips from Scommaaccent back to Scedilla, while U+0162/3 stays Tcommaaccent. Now S renders with a cedilla and T with a comma below—the mismatched pair that plagues Romanian text at the legacy code points. In the same release Adobe pre-empts Unicode, assigning U+0218/9 to Scommaaccent and U+021A/B to Tcommaaccent—eleven months before Unicode 3.0, and well after Mac OS 7.6.1. Good going bad.

UNICODE 1999 — In its 3.0 release, Unicode adds U+0218/U+0219 (S with comma below) and U+021A/U+021B (T with comma below), defined as carrying a comma accent. Great.

ASRO 1999 — The Romanian Standards Association adopts SR 13411, stipulating S/s-comma and T/t-comma as official Romanian letters. Good.

ISO 2001 — On July 15, ISO publishes ISO/IEC 8859-16 (Latin-10, "South-Eastern European"), edited by Michael Everson and incorporating Romania's SR 14111:1998 character set (ISO-IR 226)4, despite negative votes from the United States, the Netherlands, and Sweden—the FDIS passed 16–35. Finally Romanian's standard form is also the correct one. Good.

MICROSOFT 2001 — Microsoft Office v. X for Mac OS X ships crippled, without support for Unicode font display or input. Office documents with diacritics created on Windows won't display properly on the Macintosh. Bad.

APPLE 2001 — Apple immediately aligns OS X to ISO/IEC 8859-16. Good, but…

APPLE 2001 — Mac OS X does not recognize the "*commaaccent" glyph names defined by Adobe for Romanian and Baltic languages (Tcommaaccent, Rcommaaccent, Kcommaaccent, Ncommaaccent), only the "*cedilla" names or the "uni****" names. OS X thus fails to map the comma-accent glyphs to their Unicode points. [Adam Twardoch] Bad.

MICROSOFT 2001 — Microsoft, along with other software vendors, disregards ISO/IEC 8859-16. Fugly.

MICROSOFT 2001 — Windows XP launches with no fix in the box for S-comma and T-comma: no official way to type them—only third-party keyboards or the Character Map—and the font relief will not arrive until the European Union Expansion Font Update, six years later (see 2007). Bad.

ADOBE 2002 — AGL 2.0 (September 20, 2002) finally restores Tcedilla/tcedilla at U+0162/3 and deletes the U+021A/B name mappings altogether, leaving the comma-below T code points with no AGL name at all. The mismatch is now policy: consistency beats correctness. Fixed, sort of.

MACROMEDIA 2003 — Macromedia FreeHand MX (11) is released without OpenType support. Bad.

ADOBE 2003 — Adobe releases Creative Suite 1 with Unicode support. Designers can produce cross-platform Romanian typography without hacking fonts. Great.

MICROSOFT 2003 — People protest against Microsoft's practices—most notably Mr. Cristian Secară with his open letter to Microsoft Romania. Good.

ACADEMIA 2003 — The dormant Linguistic Institute of the Romanian Academy finally rules on the exact form of the marks under S and T: it must be a comma. Very late, still good.

MICROSOFT 2004 — Microsoft Office 2004 for Mac is released with Unicode support. Good.

BNR 2005 — On July 1, the leu is redenominated—1 new leu (RON) for 10,000 old lei (ROL)—and the BNR issues the new polymer series (Circulara BNR nr. 14/2005). On the fresh banknotes, the Ă is printed as A-caron—a Czech letter on Romanian legal tender, in the state's most-reproduced piece of graphic design. The error propagates unchanged through the 200-lei note of 2006 (Circulara nr. 23/2006) and the redesigned 10-lei of 2008 (Circulara nr. 37/2008). Bad, at face value.

MICROSOFT 2007 — Six years late, and five months after Romania (and Bulgaria) joined the EU, Microsoft releases updated fonts including all official glyphs of the Romanian alphabet, targeting Windows XP SP2, Server 2003, and Vista. [2026 note:] Vista did more than fonts—it shipped the Romanian (Standard) and Romanian (Programmers) keyboard layouts conforming to SR 13392:2004, both producing comma-below forms, relegating the old cedilla layout to "Romanian (Legacy)." This, not the font update alone, is the moment the Windows platform problem was actually solved. Good, at last.

APPLE 2007 — Mac OS X ignores the glyph-to-Unicode mapping in the cmap table of OpenType PS (CFF/.otf) fonts, using glyph names instead. Bad.

MICROSOFT 2007 — Windows Glyph List 4 still does not include the comma-below variants. [2026 note:] WGL4 was never updated—and it stopped mattering; Windows font coverage moved far past that baseline years ago. Bad then, moot now.

MICROSOFT 2008 — Some Adobe OpenType fonts and all Vista C-series fonts implement the optional OpenType feature GSUB/latn/ROM/locl, which forces S-cedilla to render with the comma-below glyph. When this remapping occurs, Romanian text renders correctly regardless of code-point variant. Good. [2026 note:] The refinement that later became standard practice: type designers now deliberately draw U+015E–0163 as cedilla forms and rely on locl to localize them for Romanian—because when locl fails, Romanian readers prefer two consistent cedillas to a mismatched pair.

MICROSOFT 2008 — Very few Windows applications support the locl feature; from Adobe CS3, only InDesign. Bad.

APPLE 2008 — Apple updates iPhone OS to 2.1, adding a Romanian keyboard and correct glyphs. Good.

NOKIA 2008 — Nokia phones still use cedilla glyphs. Bad.

GOOGLE 2013 — Google fixes the Roboto font in Android 4.3. Android now has proper Romanian support in both keyboard and fonts. [reported by readers Cristian and Mihai in the original comment thread] Good.

The timeline, part two: 2015–2026

The character-level question stayed closed after 2013—a review of Unicode and CLDR release notes turns up no Romanian-specific change in Unicode versions 12 through 176. What follows is a different kind of history: whether real systems carry ș and ț through intact.

ROTLD 2015 — On June 2, ROTLD enables internationalized .ro domain names with ă, â, î, ș, and ț, conforming to ISO/IEC 8859-16 and explicitly refusing the cedilla forms. The Romanian internet's front door finally speaks Romanian. Good.

BNR 2017 — The BNR reissues the entire banknote and coin series bearing the country's new coat of arms (Circulara nr. 25/2017)—every denomination reopened, redrawn, reprinted. The A-caron survives the heraldic surgery untouched. Twelve years in, the error is no longer an accident; it is house style. Bad.

RESEARCH 2020 — RoBERT, the first monolingual Romanian BERT model, is published at COLING 2020—and is benchmarked, among other tasks, on diacritic restoration. RoGPT2 follows in 2021. The machines begin learning to fix what Romanians will not type. Good.

ROMANIA 2021 — On August 2, the electronic identity card (Cartea Electronică de Identitate) pilot launches in Cluj-Napoca under EU Regulation 2019/1157. The printed card carries full diacritics; the machine-readable zone, per ICAO Doc 93037, strips them to bare ASCII—ș becomes S, ț becomes T, ă becomes A. Good and bad in the same laminate.

ACADEMIA 2021 — The Romanian Academy publishes DOOM3, the third edition of the orthographic dictionary. It treats diacritics as basic orthographic norm, reaffirms the 1993 rule on â and î, and changes nothing about glyph forms—the encoding question stays settled. Good.

BNR 2021 — On December 1, the National Bank issues the 20-lei Ecaterina Teodoroiu banknote (Circulara nr. 28/2021)—the first circulating Romanian note to depict a woman, and still carrying the A-caron in place of Ă, in the denomination itself: LEI DOUǍZECI. The first woman on a Romanian banknote deserved a real Ă. Bad, as usual.

OPENAI 2022 — On September 21, OpenAI open-sources Whisper. Robust Romanian speech recognition that outputs correct diacritics makes dictation a practical input path—typing them is no longer even required. Good.

ANAF 2022 — In October—developer forums date it to the 18th8—ANAF's RO e-Factura validation begins rejecting invoice XML that is not UTF-8, enforced by the official RO_CIUS Schematron validator under the rule ERR_SCHEMATRON: "Codificarea caracterelor pentru facturile XML TREBUIE sa fie UTF-8." A Romanian government system enforcing the encoding that carries ș and ț correctly—read that twice. Good.

OPENAI 2022 — On November 30, ChatGPT launches. Within a year, asking an AI to "add diacritics" becomes the mass-market restoration tool that two decades of utilities never were. The catch: the models are inconsistent about which diacritics—comma below or cedilla—reviving inside generative AI the exact ambiguity this article documented in 1998. Good and bad, at scale.

ACADEMIA 2023 — On August 1, DOOM3 goes fully online, free, at doom.lingv.ro. The normative reference is now one search away for everyone. Good.

ANAF 2024 — On January 1 and July 1, RO e-Factura becomes mandatory for B2B invoicing (reporting, then full clearance): near-universal, state-validated UTF-8 XML. The largest forced-correct-encoding event in Romanian history, and nobody framed it that way. Good.

EU 2024 — On May 20, the EU Digital Identity Regulation (eIDAS 2.0) enters into force, setting an end-of-2026 deadline for Romania's digital identity wallet—which will define how Romanian names, diacritics included or excluded, travel across European systems. Good, pending.

ROMANIA 2024 — On September 4, the fourth generation of electronic passports enters circulation, per the Direcția Generală de Pașapoarte communiqué: new blue-green and orange chromatics, an embossed peony—the national flower—on the back cover, production by Imprimeria Națională. Unannounced among the announced features: the cover's PAŞAPORT—S-cedilla, worn by the 2011 and 2019 designs alike, a Turkish letter on the document certifying Romanian citizenship—finally becomes PAȘAPORT. Thirteen years late, silently. And the communiqué announcing it carries cedillas in its own letterhead: DIRECŢIA GENERALĂ DE PAŞAPOARTE, Bucureşti. The directorate corrected the passport before its stationery. Good, sneaked in.

RESEARCH 2024 — OpenLLM-Ro releases RoLlama2 (May 14) and RoMistral (October 9), the first open Romanian LLMs with published recipes. Good.

ROMANIA 2025 — On March 20, the electronic ID card rollout goes national, with five million cards funded through the PNRR. Every Romanian's official machine-readable identity is now, by ICAO design, diacritic-free. Bad—by design and not our design.

RESEARCH 2025 — In November, a Babeș-Bolyai University study (Nadăș and Dioșan, arXiv:2511.13182) benchmarks a dozen LLMs on Romanian diacritic restoration. GPT-4o performs best; open models vary widely; the dominant error class is the â/î distinction; and the models waver between comma-below and cedilla output. The clusterfuck of 1998, measured with 2025 instruments. Good that it is measured; bad what the measurement shows.

There’s room for some more bad news.

More bad news: the keyboards

The Romanian national standard SR 13392:2004 establishes two layouts: a primary one and a secondary one.

The Romanian SR 13392:2004 keyboard layout, with the ă â î ș ț keys marked.
Romanian SR 13392:2004 primary layout. Source: Wikipedia.

The primary layout is intended for traditional users who learned to type on older, Microsoft-style Romanian keyboards. The secondary layout is used mainly by programmers and does not contradict the physical arrangement of a US-style keyboard; it is the default on most GNU/Linux distributions.

Apple is indeed the only company which sells localized physical keyboards on the Romanian market, but must now re-locate the specific Romanian letters on the physical keyboard according to the Romanian standard. Sooner or later Apple must do that, but the sooner the better.

—Sorin Paliga, author of Romanian keylayouts for Mac OS

It turned out that the localized keyboards Apple shipped to Romania—although functioning perfectly—were not standard compliant. And that's not all.

Physical keyboard engraving, 2008

Even though Apple's OS X was ahead of the diacritics adoption curve, and Apple was the first manufacturer to ship localized keyboards, those keyboards carried a glaring bug for years: the Romanian keyboard was engraved with the wrong glyphs—the S-comma key marked with S-cedilla, the T-comma key with T-cedilla. Un-fșșțțing-believable! I filed this with Apple's bug tracker: bug ID 6287188.

Close-up of an Apple keyboard key engraved with S-cedilla (ş) instead of S-comma (ș).

Apple keyboards erroneously inscribed—here is the S-cedilla engraving error.

Close-up of an Apple keyboard key engraved with T-cedilla (ţ) instead of T-comma (ț).

Apple keyboards erroneously inscribed—here is the T-cedilla engraving error.

[2026 note:] Closure, of a sort: the engraving flaw has since been corrected—Apple laptops with Romanian keyboards now carry the proper comma-below glyphs on their keycaps. Whether bug 6287188 played any part, Cupertino never said.

And some good news: Romanian keyboard on iPhone

The iPhone gained a Romanian keyboard with correct diacritics in firmware 2.1, released in 2008. In order to use them back then, you switched the Romanian keyboard on (Settings → General → International → Keyboards → Romanian → On), then pressed the globe key, and the space bar read “Spațiu” instead of “Space”. Then you tapped and held one of the keys (A, I, S, or T), and a row of additional letters unfolded, containing the diacritic marks.

iPhone Romanian keyboard showing the correct diacritic letters.

Android caught up—after five years!—in 2013.

Today every major mobile keyboard exposes ș and ț correctly (verified July 2026). The live problem has migrated from the keyboard to the fingers: most Romanians simply do not type diacritics on phones, and autocorrect quietly normalizes the omission. One structural pressure did dissolve along the way—SMS messages containing diacritics used to force Unicode encoding, cutting a message from 160 to 70 characters and making correct Romanian literally cost more; the migration to WhatsApp, iMessage, and RCS removed that tax. The excuse is gone. The habit remains.

The new failure modes (2019–2026)

Web-font subsetting: the new hollow document

Modern web performance practice splits fonts into Unicode-range subsets. In the Google Fonts pipeline, ș and ț (U+0218–021B) live in the latin-ext subset, not the default latin subset. A site that loads only latin—a common misconfiguration, documented in bug trackers across the WordPress and page-builder ecosystem—silently drops the comma-below letters or swaps in a mismatched fallback font. Cristian Secară described the "hollow document" of 2005, where missing glyphs rendered as empty boxes; this is its 2026 reincarnation, hidden inside a performance optimization. The fix is one line: always include latin-ext for Romanian, and preserve the locl feature when subsetting.

Artificial intelligence: restoration and relapse

Diacritic restoration—recovering ă, â, î, ș, and ț from stripped text—is now a mainstream NLP task, and large language models are startlingly good at it: the 2025 Babeș-Bolyai benchmark puts GPT-4o's restoration accuracy far above the do-nothing baseline. This is the practical answer to thirty years of diacritic-less typing, and it works in both directions: dictation via Whisper-class models produces correct diacritics without typing at all.

But the relapse is instructive. The same benchmark found the models' dominant error is the â/î distribution—the one rule that is orthographic convention rather than sound—and that models are inconsistent about emitting comma-below versus cedilla forms. A corpus study of crawled Romanian web text (Horia Cristescu's OpenCrawl-based analysis, github.com/horiacristescu) found that only 81% of it carries diacritics at all; the models were trained on our indolence, and they reproduce both it and the 1990s encoding confusion. The machines are a mirror.

The passport: a Turkish letter on the cover

For at least thirteen years—the 2011 design and its 2019 successor—the Romanian passport carried PAŞAPORT on its cover: S-cedilla, the Turkish letter, embossed in gold on the document whose entire purpose is to certify that its bearer is Romanian.

Close-up of a Romanian passport cover embossed PAŞAPORT — an S-cedilla (ş), not the correct S-comma (ș).
Burgundy pre-EU Romanian passport cover, gilt coat of arms, embossed ROMANIA and PASAPORT (no diacritics).

Romanian Passport issued June 1994–January 2002: uses no diacritics.

Burgundy Romanian passport cover (second pre-EU design), embossed ROMANIA and PASAPORT (no diacritics).

Romanian Passport issued January 2002–December 2008: uses no diacritics.

Dark-burgundy EU Romanian biometric passport cover — UNIUNEA EUROPEANA · ROMANIA · PASAPORT (no diacritics).

EU Romanian Biometric Passport issued 2008–2011: uses no diacritics.

Dark-purple EU Romanian biometric passport cover reading ROMÂNIA and PAŞAPORT — a cedilla ş, not the correct comma ș.

EU Romanian Biometric Passport issued 2011–2019: mistakenly uses wrong diacritics—letter S with cedilla (ş) instead of letter S with comma (ș).

Purple EU Romanian biometric passport cover reading ROMÂNIA and PAŞAPORT — a cedilla ş, not the correct comma ș.

EU Romanian Biometric Passport issued 2019–2024: mistakenly uses wrong diacritics—letter S with cedilla (ş) instead of letter S with comma (ș).

Red EU Romanian biometric passport cover reading ROMÂNIA and PAȘAPORT — the correct comma-below ș.

EU Romanian Biometric Passport issued 2024 and on: finally uses the correct diacritics.

Nobody seems to have protested; nobody seems to have noticed. The correction came with the fourth generation of electronic passports, put into circulation on September 4, 2024. The official communiqué proudly enumerates the new features—the blue-green chromatics, the embossed peony, the advanced security elements—and says not one word about the corrected letter; the fix shipped anonymously, folded into a security refresh. Better still, the communiqué itself is a diptych of the whole problem: its letterhead reads DIRECŢIA GENERALĂ DE PAŞAPOARTE, Str. Nicolae Iorga, Bucureşti—cedillas throughout, the institution misspelling its own name—while the body text below flows in impeccable comma-below Romanian. One document, both eras, no one looking.

Which makes the passport the middle panel of a neat official-documents triptych: the banknote, where the error was never corrected; the passport, where it was corrected silently; and the identity card, below, where the correct letters are stripped out on purpose.

The identity card: correct on the face, stripped in the zone

The electronic identity card prints your name with its diacritics and stores it correctly on the chip—and then, in the machine-readable zone that actually gets scanned at borders and by banks, transliterates it to bare ASCII, because ICAO Doc 9303 permits nothing else. Every Ștefan is STEFAN where it counts computationally. This is not Romanian incompetence—it is international standardization choosing interoperability over fidelity—but it means the diacritic-stripped double of every Romanian name is now a permanent, official artifact, and every KYC system in Europe has to engineer around the mismatch. The 2026 digital identity wallet mandated by eIDAS 2.0 will decide whether the next layer preserves names or launders them.

Bottom line

Current status: embarrassment, retired

In 2008 this article ended in embarrassment: computers in Romania could not reliably process Romanian text. That sentence is no longer true, and it deserves to be retired with honors. Windows, macOS, iOS, Android, and Linux all input and render the correct letters by default (verified July 2026). Nearly all modern fonts carry the comma-below glyphs. The standards, the code points, the keyboards, the fonts—solved.

What was not solved is us. In two decades indolence became a de facto standard, and in the decade since it became heritage: 19% of Romanian web text still carries no diacritics (the Cristescu corpus study, cited above), phones are typed diacritic-free by default of habit rather than of hardware, and the texts we feed our machines teach them the same sloppiness back. The remaining failure modes—subset fonts, ASCII identity zones, wavering AI output—are pipeline problems, and pipelines can be fixed. Habits are harder.

The only officially responsible institution to set things right from the very beginning was and is the Romanian Academy (Academia Română) via the Institute for Linguistics (Inst. de lingvistică), nobody else. This institution was and is the only responsible for this remarkable mess.

—Sorin Paliga, author of Romanian keylayouts for Mac OS, in the comments, reply no. 81, July 23, 2013.

The Academy did, eventually, do its part—the 2003 ruling, DOOM3, the online dictionary. The institutions came around. The question this article asked in 2008—how did we end up looking like idiots?—now has a sharper 2026 sequel: the tools are here, so what is our excuse now?

Take a stand

How can we improve the situation? By using the correct diacritics—there is no longer a second-best.

In 2008 I recommended dropping the Romanian diacritics when they were not available, rather than replacing them with bastardized substitutes. That is no longer good advice: they are now available—always, everywhere, with minimum effort. Not using them no longer has any excuse.

So the revised stand has three rules, in order:

  1. Use the correct letters: ă, â, î, ș, ț—comma below, never cedilla, never tilde, never caron.
  2. If your device seems unable to produce them, fix the device, not the language: the correct layout is two taps away on every platform now, and dictation will type the diacritics for you.
  3. Never substitute and never strip. The wrong letter bastardizes the language; the absent letter amputates it. Because substitutions mislead those who don't know better. Because stripped text is a backwards-compatibility baggage we hand to every parser, index, and model that reads it. Because it means you're a shitty designer. And because, well, in the end, it's just bad taste.

Short explanation, updated: wearing no underwear is not, in fact, preferable to wearing it over your trousers—both remain wardrobe failures.

© Cristian Kit Paul. Originally published October 27, 2008. Revised and extended July 2026.


For implementers: the Romanian text-integrity checklist

  • Encode ș and ț as U+0218–U+021B (comma below), never U+015E–U+0163 (cedilla); ă is U+0102/U+0103 (breve), never caron or tilde.
  • Serve UTF-8 end to end, and reject or normalize legacy encodings at the boundary.
  • Web fonts: include the latin-ext coverage (or the equivalent unicode-range) for Romanian, and preserve the OpenType locl feature when subsetting—see the hollow-document section above.
  • Normalize legacy cedilla input to comma-below on ingestion; the mapping is one-to-one: U+015E→U+0218, U+015F→U+0219, U+0162→U+021A, U+0163→U+021B.
  • Validate â/î placement against DOOM3: â inside words, î at word edges; the boundary rule survives prefixation (neîmpăcat keeps its î).
  • Test with all five letters in both cases before shipping: ĂÂÎȘȚ / ăâîșț — "În școală, țânțarul învață." A missing or substituted glyph anywhere in that string is a failed pipeline.

This checklist may be reused freely with attribution (CC BY 4.0). For NLP researchers: the measurable error classes in machine diacritic restoration are the â/î distribution and comma-versus-cedilla inconsistency (Nadăș and Dioșan, 2025); corpora should have legacy cedilla code points normalized before training or evaluation.


Quick answers

Comma below or cedilla for Romanian? Comma below—ș (U+0219) and ț (U+021B) are the Romanian letters. The cedilla forms ş and ţ are different characters (ş is Turkish; ţ belongs to no major living language), and using them in Romanian is an encoding error, not a font style.

What actually breaks with the wrong variant? Search, sorting, spell-checking, and matching—the wrong letter substitutes a foreign character and the absent letter amputates a Romanian one; both are data corruption.

Why do AI models restore Romanian diacritics imperfectly? They were trained on Romanian web text of which only 81% carries diacritics at all, and their dominant error is the â/î placement—the one rule that is orthographic convention rather than sound.

How to cite this article: Cristian Kit Paul, "Romanian diacritic marks," brandient.com, first published 27 October 2008, revised July 2026 — https://brandient.com/kit-on-romanian-diacritics.


Further reading

  • Cristian Secară's "ș-uri și ț-uri" (secarica.ro), the Romanian-language primary source on the encoding saga, still maintained.
  • Filip Blažek's Diacritics Project (diacritics.typo.cz), the typographer's reference on drawing diacritics correctly.
  • The Glyphs tutorial "Localize your font: Romanian and Moldovan", the canonical guide to the locl substitution for type designers.
  • Sorin Paliga's Romanian keylayouts for Mac OS.
  • DOOM3 online (doom.lingv.ro), the normative orthographic reference, free since 2023.
  • Decât o Revistă's feature on the diacritics war, "Diacritice: căciuliță, coif, virguliță"—the piece that wrote that this article's author "declared war on the misuse of Romanian diacritical marks."
  • And Nadăș and Dioșan, "Evaluating Large Language Models for Diacritic Restoration in Romanian Texts" (arXiv:2511.13182, 2025), the study that closes—or reopens—this article's loop.

Footnotes

  1. On the banknote Ă. The error has named a public witness: Ana Iorga, of the Academy's Institutul de Lingvistică "Iorgu Iordan – Al. Rosetti," confirmed the caron for the press, noting it appears across circulating denominations. The banknote images reproduced in this article are the National Bank of Romania's own official images (bnr.ro), where the glyph can be verified directly. ↩
  2. On the â/î reforms. The 1953 norms were approved by Hotărârea Consiliului de Miniștri nr. 3135 of September 16, 1953, and entered force on April 1, 1954; the spelling sînt predates them. The 1993 reversal is Hotărârea Adunării Generale a Academiei Române of February 17, 1993, published in Monitorul Oficial nr. 51 of March 8, 1993, with implementation rules in Monitorul Oficial nr. 59 of March 22, 1993. ↩
  3. On the Adobe Glyph List entries (1997, 1998, 2002). The version history is reconstructed from the changelog preserved in the AGL 1.2 glyph-list DATA FILE and the current file in Adobe's agl-aglfn repository; the separate AGL specification document's changelog dates versions differently (e.g., a v1.1 dated December 17, 1998), so the dates here follow the data-file record. The 1.2 changelog's own comparison table shows U+0162/3 anchored to Tcommaaccent across versions 1.0–1.2, the S reversal in 1.2, and the pre-Unicode assignment of U+0218–021B; AGL 2.0 (2002) restored Tcedilla and deleted the U+021A/B names. John Hudson's TypeDrawers account corroborates the later industry practice of drawing consistent cedillas and letting locl localize them for Romanian. ↩
  4. On the standard numbers. ISO/IEC 8859-16 incorporates SR 14111:1998, Romania's character-set standard, registered internationally as ISO-IR 226 on August 30, 1999—distinct from SR 13411:1999, cited in the 1999 entry. ↩
  5. On the 8859-16 vote (2001). The opposition wasn't about ș and ț. The US ballot comment argued that 8-bit character sets were obsolete beside Unicode and UTF-8, called 8859-16 "nothing more than a political response by a standards body to a perceived lack of sufficient stature for Romanian," and hoped it would be "widely ignored in implementation." The Netherlands, through J. W. van Wingen—himself the convener of the very working group—held that 8859-2 already served Eastern Europe, that a Romania-only standard would create "an unwanted trade barrier," and that EU accession would likely force its withdrawal. Sweden objected on principle to new 8859 parts. Sources: the FCD Summary of Voting (SC2 N3419, March 24, 2000) and the FDIS Table of Replies (SC2 N3523, May 23, 2001). ↩
  6. On "no Unicode or CLDR change, 2019–2026." A negative finding from a review of the release notes for Unicode 12 through 17 and CLDR 35 through 48, not an official Unicode statement. The only Romanian-adjacent change in the window is administrative: CLDR 35 (2019) remapped the deprecated language code mo from ro_MD to plain ro. ↩
  7. On the MRZ. ICAO Doc 9303 permits in the machine-readable zone only the twenty-six letters A–Z, the digits, and the filler "<"; national characters must be transliterated or truncated. The diacritic stripping on the identity card is therefore ICAO's design, applied worldwide. ↩
  8. On the e-Factura UTF-8 date. October 18, 2022 is the date reported by accounting-software developer forums (profox.ro, sagasoft.ro); the nearest officially documented validator update is October 25, 2022. The UTF-8 requirement is enforced at the RO_CIUS Schematron layer under the identifier ERR_SCHEMATRON; it is not a numbered BR-RO rule in OMF nr. 1.366/2021. ↩

About Brandient

Brandient is an independent brand strategy and design consultancy founded in Bucharest in 2002 by Aneta Bogdan, Cristian Kit Paul and Mihai Bogdan — never part of a global network, working directly for the owners, CEOs and boards of its clients, with an all-senior team and category exclusivity. In over two decades the firm has created or rebuilt more than 300 brands, 279 of them published as case studies at brandient.com/work — from Romania’s national leaders to brands carried worldwide. The firm has been active in Southeast Asia for more than a decade. The work has earned 84 international awards and distinctions — most recently a Red Dot (2025) and two Graphis Golds (2026) — 64 documented mentions in international design books and publications, and, in 2015, a place among the inaugural inductees of the REBRAND Hall of Fame, alongside Interbrand and MetaDesign.

get in touch

Client? contact us

5 Mendeleev Street, 3rd floor,
Bucharest 010361

Expert? Join Us

Are you brimming with talent and unshaken by challenges? Then, consider this your invitation to join our elite team, dive into complex problems, and have the opportunity to see your work make a mark far and wide.

Please submit your cover letter and relevant credentials:

friend? follow us

Copyright © Brandient 2002–2026. All Rights Reserved.