GENERATORS
UUID v4: The 122 Random Bits, and What They Cost a Database
8 min read · ToolsBay editorial · Published · Updated
Just want to do it now?
Generate cryptographically random v4 UUIDs in bulk.
The most copy-pasted UUID on the internet is 123e4567-e89b-12d3-a456-426614174000. It shows up in API docs, in test fixtures, in tutorials about version 4 UUIDs. It is not a version 4 UUID. Look at the first character of the third group: 12d3. That 1 is the version field, and it says version 1 — a timestamp-and-MAC-address UUID from the 1990s.
That single digit is the whole trick to reading these things, and almost nobody is taught it.
128 bits, of which 122 are yours
A UUID is 128 bits — 16 bytes — written as 32 hexadecimal digits in five hyphenated groups. In version 4, six of those bits are not random. They are stamped with fixed values so any parser can tell what it is holding.
Here is a real one, generated on the machine that wrote this paragraph, alongside its 16 bytes:
79c18e88-3d3f-4821-9af1-186ffc4669c4
79 c1 8e 88 3d 3f 48 21 9a f1 18 6f fc 46 69 c4
^^ ^^
byte 6 byte 8Byte 6 is 0x48. Its high nibble is 4 — the version. Byte 8 is 0x9a, and its top two bits are 10 — the variant, which says "this follows the RFC layout" rather than the older Microsoft GUID or NCS orderings. Four bits plus two is six, and 128 minus 6 leaves 122 random bits.
In the string those land at fixed offsets: the 13th hex digit is always 4, and the 17th is always 8, 9, a or b — the only four nibbles whose top two bits are 10. Generate a thousand values in the UUID generator and check the first character of the third and fourth groups on every line. It will be 4 and one of those four, on every line, with no exceptions.
That is also the fastest validity check available without a library. If the 13th digit is not 4, it is not a v4, whatever the variable is named.
The collision numbers, worked out
"Practically unique" is where most articles stop. The arithmetic is short enough to just do.
With 122 random bits there are 2^122 possible values, about 5.32e36. Collisions follow the birthday problem, not intuition: the chance that some pair matches grows with the square of how many you draw. For a space of size N and n values drawn, the probability of at least one collision is close to n^2 / (2N) while that number stays small.
Put real numbers in:
- A billion rows (
n = 1e9): about9.4e-20. Roughly one chance in a hundred quintillion. - A trillion rows (
n = 1e12): about9.4e-14. Still one in ten trillion. - To reach a one-in-a-billion chance of a single duplicate anywhere, you need about 103 trillion UUIDs. That is
sqrt(2N * 1e-9), and it checks out on a calculator.
The fifty-fifty point is sqrt(2 * ln2 * N), which comes to 2.7e18 values. Generate a billion per second and you arrive there in about 86 years. That is where the widely quoted figure comes from, and it is worth being precise about what it means: at coin-flip odds after 86 years of a billion a second, you are not going to see a collision. But "impossible" is the wrong word, and a post claiming the probability is zero has stopped one step early.
Every number above assumes the 122 bits are genuinely unpredictable. If they are not, none of it holds — and that is where real duplicates actually come from.
Where the bits come from
crypto.randomUUID() returns a v4 built from the operating system's entropy source. One call, present in every current browser, and Node exposes it too.
The alternative you will find on Stack Overflow builds the string from Math.random() — the template full of xs with a hard-coded 4, replaced digit by digit. The output is indistinguishable from a real v4 by eye. It passes any regex you write. It is still the wrong thing, because Math.random is not a random source. In V8 it is xorshift128+, a generator with 128 bits of internal state and no cryptographic claim whatsoever. Given a run of consecutive outputs you can solve for that state, and once you have it, every past and future value falls out. There is published work doing exactly that.
For a primary key, nobody can do much with it. For a password reset link, an invitation URL, or an unverified-signup token — all things routinely represented as a UUID — an attacker who can mint one identifier of their own and read it back has a foothold for predicting yours. That is the distinction the UUID generator is pointing at when it says values come from crypto.getRandomValues rather than Math.random, and the same one behind the secure password generator, where it bites harder still.
One deployment detail catches people out: crypto.randomUUID() is restricted to secure contexts. On a page served over plain http:// — anything other than localhost — it is simply absent, and you get a TypeError rather than a warning. crypto.getRandomValues() carries no such restriction. That is why this site's own generator falls back to filling 16 bytes with getRandomValues and setting the two special nibbles by hand:
const bytes = new Uint8Array(16);
crypto.getRandomValues(bytes);
bytes[6] = (bytes[6] & 0x0f) | 0x40; // version 4
bytes[8] = (bytes[8] & 0x3f) | 0x80; // variant 10xxThose two assignments are the entire specification of a v4 UUID. Everything else is formatting.
A UUID is not an access control
One claim gets repeated whenever UUIDs come up: sequential integer IDs let an attacker walk /users/45 to /users/46, and UUIDs fix it.
They do not fix it. They make the door harder to find; they do not lock it. If /users/<uuid> returns data to anyone who knows the UUID, that is an authorization bug and the UUID has only hidden it. Identifiers leak — into referrer headers, support tickets, screenshots, browser history, third-party analytics, the shared spreadsheet someone pasted them into. The fix is checking on every request that the caller is allowed to see the row. Opaque identifiers are a reasonable second layer. They are not the first one.
The cost the database pays
This is the part most write-ups skip, and it is the part that shows up in a graph.
In MySQL, InnoDB stores the table itself in primary key order. A sequential key means every insert lands at the right-hand edge of the tree, in a page already sitting in memory. A random v4 key means every insert lands in an arbitrary leaf. Once the index stops fitting in RAM, that becomes a disk read before it can be a write, and pages split at random fill levels instead of packing tight — so the index grows larger than the data warrants.
Then there is width. InnoDB stores a copy of the primary key inside every secondary index entry. Store UUIDs as CHAR(36) rather than BINARY(16) and you pay the extra 20 bytes per row in the clustered index and again in every secondary index. Ten million rows with five secondary indexes: 20 bytes times 5 times 10 million is a full gigabyte of pure formatting, before any data.
PostgreSQL is a softer story, and people over-apply the MySQL warning to it. Postgres heaps are not clustered on the primary key, so inserts do not scatter the table itself — but the B-tree on the uuid column still has no hot edge, so its cache behaviour degrades the same way. At least the uuid type is 16 bytes natively there, so you are not paying for hex.
UUIDv7, and what it gives back
RFC 9562, published in 2024, standardised version 7: a 48-bit Unix millisecond timestamp in the high bits, then the version nibble, twelve more bits, the variant, and 62 random bits. Values generated later sort after values generated earlier, so inserts return to the right-hand edge of the index.
The timestamp reads straight out of the string. Take the first twelve hex digits of this v7:
01a08ab1-7e2d-727f-abe8-5954b620ccd2
^^^^^^^^ ^^^^
0x01a08ab17e2d = 1789033283117 msPut 1789033283117 into the timestamp converter and it resolves to 10 September 2026, 09:41:23.117 UTC. Note the 7 in the 13th position — same rule as before.
Because lowercase hex digits sort in ASCII order, a v7 UUID sorts correctly as plain text, not only as bytes. ULID makes the same trade with a different encoding: 48 bits of timestamp plus 80 random bits, written as 26 Crockford base32 characters instead of 36.
Two things to weigh. First, v7 carries at most 74 random bits, not 122 — an enormous per-millisecond space still, but a third of the entropy gone. Second, and more important: a v7 UUID publishes its own creation time to the millisecond. For a row ID that is often useful. For an invite link or a reset token it means you have handed out a precise timestamp and narrowed the guessing space, and anyone holding two of them can work out the order and the interval between the records.
So: v4 when the identifier should reveal nothing, v7 when it is the key of a table you insert into constantly. They are not rival versions. They answer different questions, and the honest reason to reach for either is the same one — you wanted to generate an identifier without asking anything else for permission first.