URL encoding: the characters that quietly break your links
Somebody sends you a link with a UTM tag on it. You click it, land on the page, and the analytics show a source of "newsletter" instead of "newsletter&medium=email". Or worse, the page just ignores everything after the ampersand. Nobody broke the link on purpose. It was never encoded.
URLs have a small alphabet of reserved characters. Ampersand splits query parameters, equals sign assigns values, hash jumps to an anchor, plus means space in some contexts. The moment one of these characters appears inside a value instead of doing its structural job, parsers get confused and your data gets truncated at whatever character the parser treats as a boundary.
The everyday cases where it bites
The classic is an email subject line with an ampersand: mailto:me@example.com?subject=Q&A session. Unencoded, the mail client sees a second parameter named "A session" and your subject becomes "Q". The fix is subject=Q%26A%20session, and once you have typed it wrong twice, you start encoding it with a tool instead.
UTM links are the other repeat offender. Campaign names with spaces or ampersands silently split into phantom parameters, so your reports fragment into several fake sources. Building them in a UTM builder that encodes values keeps the tracking data in one piece. Same story for mailto links, pre-filled forms, and any redirect URL you stuff inside another URL's query string.
Which characters need encoding
&,=,?,#— reserved by the URL structure itself- Space, quotes, angle brackets — invalid in URLs and inconsistent across clients
- Non-ASCII text like Cyrillic or accented letters — encoded as UTF-8 byte sequences
- Everything else can stay as-is; encoding it also works, just makes URLs longer
The difference between %20 and + trips people regularly. In a query string both usually mean a space, but in the path portion + is a literal plus. If a backend decodes + as a space in a path segment, encoded spaces work and literal pluses stop meaning addition. When in doubt, decode the URL and look at what you actually have.
Encoding twice is its own bug
Double-encoding is the mirror problem. Encode a URL, then encode it again, and %26 becomes %2526. The server decodes once, sees %26, and treats it as literal text instead of the parameter separator it was meant to be. This mostly happens in redirect chains: an app builds a return URL, a middleware encodes it again, and the user lands somewhere with %2520 visible in the address bar.
The rule is simple: encode at the last possible moment, in exactly one place, and treat encoded values as opaque strings you do not touch again. If you inherit a system where URLs seem to degrade over time, decode one sample all the way down and count how many layers of encoding you find. More than one layer is the bug.
A two-minute habit
Before you ship any link that contains user input, a campaign tag, or a sentence with punctuation in it: decode it once and read it as plain text. It should look like what you meant. If it does not, encode the values and move on. That habit catches the broken mailto links, the fragmented UTM data, and the double-encoded redirects before anyone clicks them.