Skip to main content
Text Processing

Case conversion is more than CSS text-transform

The fastest way to make a list of names look uniform is text-transform: uppercase in CSS. It works right up until you need those names in a CSV export, an email merge, or a database dedupe — at which point you discover the display was never the data. "JOHN SMITH" is still stored as "JOHN SMITH", and every downstream consumer of that field now has to clean it up again.

Case conversion belongs at the data layer when the data itself should change, and in CSS only when the change is purely presentational. Sorting the two out early saves the classic bug where a search for "McDonald" finds nothing because everything was normalized to "mcdonald" months ago.

Where plain uppercase() is not enough

Three traps show up constantly:

  • Names and titles. "McDonald", "van der Berg", "iPhone", "III" all get mangled by naive capitalization. Lowercase-everything destroys them permanently, and re-capitalizing guesses wrong.
  • Locales. Turkish distinguishes dotted and dotless i — strtoupper('istanbul') in a Turkish locale produces "İSTANBUL", while the same code in English gives "ISTANBUL". PHP's mb_convert_case with an explicit locale handles this; raw strtoupper does not.
  • Format conversions. camelCase to snake_case is not a character transform, it is a parse: you need word boundaries, acronyms kept intact ("parseHTML" should become "parse_html", not "parse_h_t_m_l"), and leading uppercase handled. Doing this by hand for a list of fifty fields is where the typos come from.

The data-cleaning side

Most case bugs I have debugged started as a copy-paste job. Someone pastes a column of mixed-case emails from a spreadsheet, dedupes in the app, and "Anna@x.com" and "anna@x.com" coexist forever because the comparison was case-sensitive. Lowercasing emails and usernames at ingestion is boring and works.

For larger pastes the workflow matters more than the code. Get the text into one case, then deal with the other problems hiding underneath it: duplicated lines that differed only by case, stray whitespace, inconsistent separators. A case converter handles the first step, and its neighbors cover the rest — removing duplicate lines after a case fold, or normalizing Unicode so that two visually identical characters stop comparing as different strings.

That last one deserves a paragraph. Case-insensitive comparison in JavaScript usually calls toLowerCase(), which is fine until a German sharp s ("ß") meets "SS", or a character exists in two Unicode forms. If your dedupe "sometimes misses", the invisible cause is usually normalization, not case.

Choosing the transform on purpose

The list of transforms looks trivial — upper, lower, title, sentence, camel, snake, kebab — and that is exactly why nobody thinks about it. The right question is what the target system expects. HTTP headers are case-insensitive by spec but lowercase by convention. Environment variables shout. Database table names follow whatever the team chose in year one. A JSON API that returns snake_case keys consumed by a JavaScript frontend will generate a camelCase conversion on every response, forever, so automate it once.

And when the goal is purely visual, keep it in CSS. text-transform changes what the reader sees without corrupting the string for the next system that reads it. The mistake is not using CSS — it is using CSS for data that later needs to be real.

A quick test

Take the string that keeps annoying you and run it through the transform you have in production code. If the result surprises you — a lost accent, a mangled acronym, a Turkish i — you have found your next bug before a user did. Two minutes with a slug generator or a converter beats a ticket about "the search does not find anything".