ENCODING / SECURITY
Why URL Parameters Break: %20, +, and %2520
8 min read · ToolsBay editorial · Published · Updated
Just want to do it now?
Percent-encode or decode URLs and query string values.
You send someone a link. It arrives in two pieces: a clickable part that 404s, and a few words of plain grey text trailing after it. The usual explanation is that the URL specification is strict about which characters are allowed to travel. That is not what happened here.
Plain-text email has no markup. There is no anchor tag, nowhere to record where the URL ends and the sentence resumes. So the mail client guesses. Its linkifier scans forward from https:// and stops at the first character it treats as a terminator, and whitespace is always a terminator. Put a space in the middle of a URL and the link ends at the space, because nothing in the text says that space is yours rather than the sentence's.
The URL did not break. It was never one URL as far as the client was concerned. Percent-encoding fixes it by removing the ambiguity — but it is worth knowing precisely what percent-encoding is doing, because almost every other bug in this area comes from applying it in the wrong amount.
What a URL actually reserves
RFC 3986 divides the characters three ways.
Unreserved — safe anywhere, never need escaping: A-Z, a-z, 0-9, and - . _ ~. That is the complete list.
Reserved — the punctuation a URL is built from. The gen-delims : / ? # [ ] @ separate the major parts. The sub-delims ! $ & ' ( ) * + , ; = are available to individual components for structure of their own.
"Reserved" does not mean forbidden. It means the character has a job somewhere, and if you want it as data rather than as punctuation you have to escape it. Which component you are in decides. A ? inside a fragment is just a question mark — everything after the first # is fragment, so there is nothing left for it to delimit. A / in a query string is legal and routine; look at any OAuth redirect parameter. The same / in a path is a separator and can never be data.
One thing worth stating flatly, because a lot of writing on this gets it backwards: RFC 3986 does not define a query string as key/value pairs. Its grammar for the query component is, roughly, "any path character, plus / and ?". The key=value&key=value shape is a separate convention called application/x-www-form-urlencoded, defined by HTML rather than by the URL spec. That is why & and = are sub-delims instead of hard delimiters — and it is why the rules below are shaped the way they are.
encodeURI and encodeURIComponent, and how each ruins the other's job
JavaScript ships two escaping functions. Choosing the wrong one is the second most common encoding bug there is.
encodeURIComponent escapes everything except the unreserved set plus ! ' ( ) *. Slashes, colons, question marks, ampersands, hashes — all escaped.
encodeURI assumes you handed it a finished address and leaves every reserved character alone, # included. In practice it escapes spaces, non-ASCII, and a few odds like <, >, ", {, }.
Each one destroys the other's use case. Feed a complete URL to the component encoder:
encodeURIComponent('https://example.com/a?b=c d')
// 'https%3A%2F%2Fexample.com%2Fa%3Fb%3Dc%20d'That is not a link any more. It is a string that spells one.
Now the reverse — a parameter value containing an ampersand, escaped with the whole-URL rule:
encodeURI('name=Tom & Jerry')
// 'name=Tom%20&%20Jerry'The space is handled; the ampersand is not, because to encodeURI an ampersand is structure. Send that and the server reads two parameters: name with the value Tom (trailing space), and a second parameter whose name is Jerry and whose value is empty. Nothing throws. Half the data quietly went missing.
The rule that falls out: escape each value with the component rule, then join the pieces with & and = yourself. Never escape an assembled query string in one pass. The URL encoder offers both modes side by side because there is no single correct answer — it depends on which of the two you are holding, a value or an address.
The five characters JavaScript leaves alone
encodeURIComponent does not escape ! ' ( ) *. All five are sub-delims in RFC 3986 — reserved. The exemption is inherited from the older RFC 2396, which classed them as "marks" and treated them as safe.
Usually harmless. It stops being harmless when something downstream disagrees about the reserved set. OAuth 1.0 signatures (RFC 5849) require those five to be percent-encoded, so a signature computed over an unescaped apostrophe will not match the one the server computes. The standard patch:
const strict = (s) =>
encodeURIComponent(s).replace(/[!'()*]/g, (c) =>
'%' + c.charCodeAt(0).toString(16).toUpperCase());%20 or +, and why a plus sign comes back as a space
Both encode a space, in different contexts.
%20 is percent-encoding — the RFC 3986 answer, valid in any component. + is form-urlencoding: the HTML form serializer maps a space to +, browsers use it for every GET form submission, and URLSearchParams follows the same rules.
The consequence catches people constantly: inside a query string, a literal + means a space. Search for C++ the naive way and the term never arrives.
new URL('https://example.com/search?q=C++').searchParams.get('q')
// 'C ' ← two spaces, no plus signsThe correct query string is ?q=C%2B%2B, and you rarely need to write that by hand — new URLSearchParams({ q: 'C++ tutorial' }).toString() returns q=C%2B%2B+tutorial. Plus escaped where it is data, plus used where it is a space, in the same string.
The ambiguity is baked into the ecosystem, not into any one library. Java's URLEncoder.encode emits + for a space. PHP's urlencode emits +; its rawurlencode emits %20. Python's urllib.parse.quote gives %20, while quote_plus and urlencode give +. None of them is wrong. They are answering different questions.
It also means percent-decoding is not one operation. Decoding query-string data should turn + into a space; decoding a path segment should not, because there a + is only ever a plus. And if you paste standard Base64 into a query-string decoder, every + in it becomes a space and the data is ruined — which is the whole reason base64url exists, swapping + and / for - and _.
Non-ASCII is bytes, not characters
Percent-encoding escapes bytes. Text is encoded to UTF-8 first, then each unsafe byte becomes % plus two hex digits. So the cost of a character depends on how many bytes it takes.
éis two bytes:%C3%A9日is three:%E6%97%A5🚀is four:%F0%9F%9A%80— one emoji, four triplets, twelve characters in the URL
Because it is bytes and not characters, the same letter produces different escapes under different encodings. In Latin-1, é is the single byte 0xE9, so an old form might submit %E9. Decode that as UTF-8 and you do not get a wrong character, you get an error: 0xE9 announces a three-byte sequence whose remaining bytes never arrive. That is what "URI malformed" means when decodeURIComponent throws, and it is why a decoder that guesses rather than failing is doing you no favours.
%2520, the bug you will actually hit
%25 is a percent sign. So %2520 is the escaped form of %20, which is the escaped form of a space. Decode once and you get %20. Decode twice and you get the space.
Double-encoding happens whenever an already-escaped URL is escaped again — most often a redirect parameter that gets built correctly, then passed through the encoder a second time on its way into a link.
encodeURIComponent(encodeURIComponent('https://example.com/a b'))
// 'https%253A%252F%252Fexample.com%252Fa%2520b'The symptoms are recognisable once you know the shape: a filename displayed with a literal %20 in it, a redirect landing on a URL with %2F where the slashes should be, a search box that returns nothing for a term containing an apostrophe. Paste the URL into the URL parser and it shows up straight away — the tool decodes each parameter once, so any value still carrying a % triplet afterwards was encoded twice.
While you are in there: an encoded slash inside a path is legal per the spec but not universally accepted. Apache httpd rejects those requests with a 404 by default, because AllowEncodedSlashes is off unless someone turned it on. A path segment carrying a %2F can work locally and 404 in production.
Putting a link somewhere it survives
Back to the broken email. Two things keep it in one piece.
Encode the spaces before you paste, so the linkifier has nothing to stop at: /reports/Q3%20final.pdf rather than /reports/Q3 final.pdf. And wrap the whole address in angle brackets — <https://example.com/reports/Q3%20final.pdf> — which RFC 3986 Appendix C recommends for exactly this situation and most clients honour. It also keeps a full stop at the end of a sentence from being pulled into the link.
The same question applies elsewhere: who reads this string next, and what do they think the punctuation means? A URL going into an HTML attribute needs two escapes at once — percent-encoding for the URL layer, and & for the ampersand, because href="?a=1&b=2" hands the HTML parser something that begins like a character reference. Percent-encoding does not help there; that is a job for the HTML entity encoder, applied after the URL is already correct.
None of this is deep. It is one question asked over and over, at every boundary the string crosses.