Skip to main content
Data Conversion

UUIDs: which version to pick and when v7 actually matters

If you have ever typed Str::uuid() and moved on, this is for you. There are eight UUID versions, and picking the wrong one does not usually break anything. It just makes your database slower and your logs harder to sort. Here is the short version of which one does what.

The versions that actually get used

v4 is random. Every bit that is not part of the version header comes from a random source. No ordering, no meaning, no collisions in practice. This is the default in most libraries and the right choice most of the time.

v1 is timestamp plus MAC address. It is sortable, but it leaks the MAC of the machine that generated it and some drivers generate them in a way that leaks timing details. Mostly a legacy format now.

v5 is a hash. You feed it a name and a namespace, and the same input always gives the same UUID. Useful when you need the same entity to produce the same ID across systems, for example mapping an external feed item to a stable local key.

v7 is a timestamp in the high bits, random in the low bits. Standardized in RFC 9562. It sorts correctly by creation time and still has enough randomness to avoid collisions between machines.

Why sorting matters more than it sounds

Random v4 keys land in random places in a B-tree index. On InnoDB or Postgres, an insert into the middle of an index means page splits, a fatter buffer pool, and write amplification you pay for on every insert. With v7, new rows always land at the right edge of the index, which is the same behavior you get from a plain auto-increment integer.

The other payoff is boring and practical: ORDER BY id on a v7 column gives you creation order. That makes "show me the last 50 events" a free index scan instead of a sort, and makes log triage much easier when you are staring at a wall of IDs.

When v4 is still the right answer

If the ID is not stored in a large table, if it never appears in an ORDER BY, or if it identifies something whose creation order must stay private, keep v4. Sorting by ID leaks creation time, and for things like invitation tokens or order references that can matter. v4 also wins when IDs are generated offline and merged later without a clock you can trust.

Where v5 fits

Use v5 when the ID has to be reproducible from a natural key. Import pipelines are the classic case: if a feed re-sends the same item, a v5 derived from the item URL means your importer can upsert instead of dedupe by extra columns.

Checking a UUID you already have

Before you commit to a version in code, it is worth knowing what you already ship. Paste an ID into the UUID generator and it breaks the ID into version, variant, and timestamp fields where applicable. Handy when a third-party API hands you a UUID and you need to know whether it is sortable before you index it.

For the rest of your data hygiene, the conversion tools cover the other formats that end up next to UUIDs in most systems.

The short version

Default to v4. Switch to v7 when the table is big, write-heavy, or you need creation-ordered IDs. Reach for v5 when the ID must be derivable from a natural key. Skip v1 for anything new.