What Your Random String Choice Says About Your Security
Every developer has shipped something that "just needed a random string" and picked the first method that came to mind. Most of the time nothing breaks. Sometimes two users get the same identifier and the support ticket takes a week to make sense of.
The odd part is that the actual randomness is rarely the problem. JavaScript's Math.random() is not cryptographically secure, true — but for a session ID or an API key that weakness is only one of several ways the system fails. Collisions, encoding, and human error usually get there first.
Three failure modes, three fixes
1. Collisions from short strings
A 6-character alphanumeric string has roughly 56 billion combinations, which sounds enormous until you generate thousands per day and remember the birthday paradox. Short tokens are fine for short-lived things — a one-time link that expires in an hour does not need more. Long-lived identifiers need length: 32+ characters makes collisions practically impossible.
When you need to check how much room you actually have, a random string generator that lets you set length and alphabet explicitly beats guessing. Generate a few, look at them, sanity-check the length your storage column actually allows.
2. Characters that break somewhere else
The string survives your database and then breaks in a URL, an Excel sheet, or a phone system. Symbols like + and = are the offenders. For anything that ends up in a link, restrict the alphabet to letters and digits, or use base62. If you are debugging why a token got mangled in transit, a URL decoder shows you exactly what a middle layer did to it.
3. UUIDs treated as strings
UUIDs solve the collision problem for good — generating one gives you an identifier unique enough to use across systems without coordination. The mistake is treating them as opaque strings everywhere: they index poorly in databases at scale, and they leak the generation method in their version nibble. Use them for cross-system identity, not as a lookup key you sort on.
What to actually use
- Database identifiers and API keys: long random strings from a proper alphabet, 32 characters and up.
- Anything user-visible: human-readable parts where possible. A support ticket labeled
K7M-2XQbeats9f8b2c1e-4d5aon every phone call. - Security-sensitive tokens: the crypto-random source of your language, not a hand-rolled combination of timestamp and username. A strength checker gives you a rough feel for what entropy looks like in practice.
None of this requires you to become a cryptography expert. It requires deciding, once per project, which failure you can afford: a rare collision you can retry, or a predictable token someone can guess. Pick the generator to match that decision, write it down in the codebase, and stop re-deciding it every sprint.
The next time something "random" misbehaves, check the alphabet and the length before you blame the entropy. That is where these bugs usually live.