Clean Scraped Text Before It Breaks Your Script
Text you copy from a web page, a PDF, or a spreadsheet is rarely the text it appears to be. It looks clean in the editor and then a comparison fails, a join returns nothing, or a regex refuses to match, and you spend twenty minutes blaming your code when the actual culprit is an invisible character you pasted in with the data.
The invisible troublemakers
A few characters cause most of the pain. The non-breaking space (U+00A0) looks exactly like a space and breaks explode(' ', $line). Zero-width spaces and joiners come along from web text and are completely invisible in a terminal. Windows line endings (CRLF) hide at the end of every line until a strict comparison or a CSV parser trips over them. Different Unicode forms of the same letter, like é composed versus decomposed, make identical-looking strings compare as different.
A Unicode normalizer flushes all of this out in one pass. Run the text through it and the characters you could not see are gone or converted to their plain equivalents. If you want to see what you were dealing with, paste the same text in before and after and compare the byte counts.
Structure problems
Once the characters are clean, the structure usually needs the same treatment. Scraped lists come with blank lines between every entry, duplicated rows from overlapping page sections, and trailing whitespace that nobody meant to include. A duplicate line remover plus a pass to strip the line breaks turns a messy paste into a list you can split and join without surprises.
If the source is HTML rather than rendered text, strip the tags first. A tag remover gets the content out from between the markup, and normalizing the result takes care of the entities that survive it.
Make it a habit
The pattern I settled on for anything pasted from outside: normalize Unicode, cut line endings down to one kind, drop duplicate and empty lines, then start working with the data. Thirty seconds of cleanup up front, versus debugging why two strings that look identical do not match. The comparison always loses to the invisible character.
Where this actually bites
The classic victim is the comparison. You fetch a value from an API, compare it to a value from the database, and they are never equal, because one carries a zero-width joiner and the other does not. Hash comparisons behave the same way. The second classic is the import script that worked perfectly in testing and then produced one garbage row per hundred in production, because some rows came from Excel and some came from a web form, and the two sources disagree about what a space is.
Another flavor is the CSV that will not import. A field contains an embedded line break, the parser starts a new record halfway through, and everything downstream shifts by one column. Normalizing the raw text before it reaches the parser removes most of these surprises, and converting the CSV to JSON in one step is often easier to debug than eyeballing the raw file.
Normalize once, at the border
Cleanup is cheapest at the boundary. Whatever enters your system, from a paste, an upload, a scrape, or a partner API, gets normalized before it is stored or compared. Cleanup inside the business logic tends to spread: every endpoint grows its own idea of what clean means, and the code becomes a museum of slightly different string operations. One pass at the door, stored clean, and the rest of the code can trust its own data.