ENCODING / SECURITY
Base64 Encoding Explained: What Actually Happens to Your Bytes
8 min read · ToolsBay editorial · Published · Updated
Just want to do it now?
Encode text to Base64 or decode it back, with full Unicode support.
Every JWT you have ever seen starts with the letters eyJ. Every one. That is not a version marker or a convention — it falls out of the encoding. A JSON object starts with {", and Base64 turns those two bytes plus whatever comes next into eyJ followed by one variable character. Once you can see why, you can read a fair amount of the internet's plumbing by eye.
Base64 is a way of writing arbitrary bytes using 64 characters that survive being pasted into a stylesheet, an HTTP header, a query string or an email body without anything along the way rewriting them. It is not encryption, it is not compression, and it costs you a third of your bytes. Here is exactly what happens to them.
Three bytes in, four characters out
The alphabet is A–Z (values 0–25), a–z (26–51), 0–9 (52–61), then + (62) and / (63). Sixty-four values, so each character carries exactly 6 bits.
Bytes are 8 bits. Base64 characters are 6 bits. The smallest number both divide into is 24 — 3 bytes, or 4 characters — so the encoder chews through your data three bytes at a time and emits four characters.
Take the word Man, the canonical example:
Text M a n
ASCII 77 97 110
Bits 0100 1101 0110 0001 0110 1110
Regrouped 010011 010110 000101 101110
Value 19 22 5 46
Base64 T W F uMan becomes TWFu. Three bytes, four characters. Nothing is added and nothing is lost — the same 24 bits are simply cut in a different place. That regrouping is the whole algorithm. Everything else is bookkeeping.
The 4/3 ratio is where the famous 33% overhead comes from. It is not an approximation. Encode this site's own 6,004-byte icon-192.png and you get exactly 8,008 characters, because 6,004 bytes is 2,002 groups of three once you round up.
Where the = signs come from
Your data is rarely a neat multiple of three bytes, so the last group is short and the encoder has to say so.
Drop the n from Man and you have two bytes, 16 bits — enough for two full 6-bit values with 4 bits left over. The encoder pads those 4 bits out to 6 with zeros, emits a third character, and appends one = to mark that the final group was one byte short:
Man -> TWFu (3 bytes, no padding)
Ma -> TWE= (2 bytes, one pad)
M -> TQ== (1 byte, two pads)The = is not data and carries no bits. Its job is to keep every encoded chunk a multiple of four characters, so concatenated Base64 streams can still be split apart correctly. When Base64 is the only thing in the field, a decoder can work the length out without the padding — which is why one of the two common variants throws it away entirely.
What the encoding was actually built for
The usual story is that old routers could not cope with binary. Routers move packets and never look at your bytes. The real constraint was in mail.
SMTP was specified to carry 7-bit US-ASCII. A message body is terminated by CRLF.CRLF — a line containing a single dot — and lines are capped at 1,000 octets. Push raw binary through that and you get three separate failure modes: a gateway strips the high bit off every byte and silently corrupts the file; a byte sequence that happens to spell a lone dot on its own line ends the message early; and a long run without a 0x0A blows the line limit. None of this involves anything panicking. It is a text protocol doing exactly what it was told to do with data that was never text.
Base64 removes all three problems at once, because none of its characters can trigger any of them. It is also why MIME caps encoded lines at 76 characters. With the CRLF on each line, an email attachment is not 33% larger than the file: 57 source bytes become 78 transmitted bytes, so it is closer to 37%.
Two alphabets, and what a JWT really is
+ and / are both meaningful in a URL. / is a path separator and + is read as a space in form-encoded data, so standard Base64 mangles on contact with a query string.
RFC 4648 defines a second alphabet for exactly this: base64url, where value 62 becomes - and value 63 becomes _, and the = padding is normally dropped. The three bytes FB EF BE encode as ++++ in standard Base64 and ---- in base64url. Same bits, different spelling. The Base64 encoder has a URL-safe toggle on the encode side and accepts either alphabet when decoding.
This is the variant JWTs use, and it is worth being precise about what a JWT holds, because the two-thirds-right version is everywhere. A token is three base64url segments joined by dots:
eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiIxMjM0NTY3ODkwIiwibmFtZSI6IkFkYSIsImV4cCI6MTg5MzQ1NjAwMH0.drqZS3BzLggYBY6iMM489OP3Q8s2id0Q4hMHQCTOmu0The first two segments are JSON: {"alg":"HS256","typ":"JWT"} and the claims. The third is not. It is the raw output of the signing algorithm — for HS256, the 32 bytes of an HMAC-SHA256 digest, which is why it is always 43 characters once the padding is stripped. Decode it as text and you get binary, correctly.
That distinction is the point of the format. The header and payload are readable by anyone holding the token; Base64 hides nothing. The signature is what proves the claims have not been edited, and checking it needs the secret. Reading a token in a JWT decoder tells you what it says, not whether it is genuine.
data: URIs, and the advice that breaks images
A Base64 string on its own is not an image source. Paste one straight into src and the browser goes looking for a file with a very long name. It needs the data: scheme from RFC 2397 — a media type, the word base64, a comma, then the payload:
<img src="data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…" alt="">.icon { background-image: url("data:image/svg+xml;base64,PHN2ZyB4bWxu…"); }Get the media type wrong and browsers disagree about what to do with it, which is the usual reason a hand-written data URI renders in one browser and not another. The image to Base64 converter reads the type from the file's actual bytes and writes the prefix for you.
The trade-off nobody quotes
Inlining is sold as removing an HTTP request. Measure it before you believe it.
That 6,004-byte icon becomes 8,008 characters of Base64. Compressed for transport the picture changes completely. Run both through gzip at level 9 and the PNG lands at 5,697 bytes while the Base64 version lands at 5,882 — so the 33.4% you started with is 3.2% by the time it reaches the browser. Brotli, which is what most CDNs actually serve, gives 5,654 against 5,856, or 3.6%. The real cost is elsewhere. An inlined asset cannot be cached on its own. It is re-downloaded with every revision of the file it lives in, and two pages that both use it cannot share one copy. And HTTP/2 multiplexes requests over a single connection, so the request you saved was never as expensive as the advice assumes.
There is one case where inlining is clearly wrong, and it is the one people reach for most. This site's 810-byte icon.svg compresses to 284 bytes with brotli on its own. Base64 it into a data URI and the same artwork compresses to 579 — 104% worse, because Base64 shreds the repeated byte patterns the compressor was feeding on. Percent-encode it instead, leaving it as text, and it compresses to 329. Base64-ing an SVG more than doubles what you ship. For SVG, skip Base64.
Every figure in this section is from the files in this site's own public/ directory, measured with Node's zlib at gzip level 9 and brotli's default. Compression numbers move with the level, so a figure quoted without one cannot be checked — which is the whole point of quoting them.
One trap in the browser
If you are encoding text yourself, btoa is not quite the function you think it is. It takes a string of bytes, not a string of characters, and reads each code unit as one byte. Anything above U+00FF throws InvalidCharacterError, so btoa('café ☕') fails outright. The quieter case is worse: btoa('café') succeeds and returns Y2Fm6Q==, because é is U+00E9 and fits in a single byte. Those are Latin-1 bytes. A UTF-8 decoder on the other end gives you caf and a replacement character.
The fix is to encode to UTF-8 bytes first, then Base64 those bytes:
const bytes = new TextEncoder().encode(text);
const b64 = btoa(String.fromCharCode(...bytes));
// 'café' -> 'Y2Fmw6k='Compare the two results: Y2Fm6Q== against Y2Fmw6k=. One byte versus two, for the same character. Encoders that skip the UTF-8 step round-trip English perfectly and quietly corrupt everything else — a bug that shows up in production and never in testing.
And it is still not encryption
Nothing above involved a key, and no step of it is hard to reverse. Base64 in an HTTP Basic auth header is username:password in public view; the site's encoder will read it back in one keystroke, and so will anyone watching an unencrypted connection. Base64 in a config file is a value someone chose not to read, not a value they cannot read. It is a transport format. Treat anything it wraps as fully visible.