URL Encoding vs Base64: When to Use Which
2026-09-30 · 1276 words
Both percent-encoding and Base64 turn input into a restricted set of ASCII characters, both are reversible, and both are routinely described as "making a string safe". They are also answers to completely different questions, which is why choosing between them by feel produces data that survives the round trip in testing and falls apart in a log file three weeks later.
The distinction is one sentence long: percent-encoding protects characters that have syntactic meaning inside a URL, and Base64 re-encodes bytes that are not text at all. Get that backwards and nothing throws an exception — the value simply arrives different, or half of it arrives.
The rule, stated once
- If what you have is text that must sit inside a URL component (a search term, a filename, a redirect target), percent-encode it for that component.
- If what you have is bytes (an image, a hash, an encryption key, a protobuf), Base64 them — and then percent-encode the Base64 if it has to travel in a URL.
The second clause is the one people skip, and it is the source of the classic bug covered below. The two transforms compose; they do not compete.
What each one actually guarantees
Percent-encoding is defined per component of the URL grammar. It says: replace the bytes that are not allowed here with % followed by two hex digits. Its output is still text with the same meaning — a human can read an encoded path and see what it was. Its cost is unbounded in the worst case: an ASCII character costs one %XX triplet if it needs escaping (three characters), and a character outside ASCII costs one triplet per UTF-8 byte, so a single CJK character is nine characters.
Base64 is defined over bytes, with no knowledge of URLs. Its output uses 64 characters — A–Z, a–z, 0–9, +, / — plus = padding, and it expands by exactly one third. All of those characters are legal in a URL as characters, which is exactly why Base64 breaks there: + means something else, / means something else, and = is used by other parts of the URL syntax. Legal, and load-bearing, are not the same property.
The failure everyone hits: Base64 in a query string
Take an image encoded as standard Base64 and dropped into a query parameter:
https://example.test/img?data=iVBORw0KGgoAAAANSUhEUg...+/8AAAAASUVORK5CYII=
Every character there is legal in a URL, so the browser sends it and the server receives it. Then the server parses the query string with application/x-www-form-urlencoded rules — which is what most frameworks do by default — and in those rules + means a space. The Base64 that arrives is not the Base64 that was sent. If it decodes at all, it decodes to different bytes; more often it fails at a length that is not a multiple of four, at a totally different character position from the one that was altered.
There are two correct fixes, and they are not interchangeable:
- Percent-encode the Base64 before appending it, and decode in the reverse order on the way back.
+becomes%2B, which form parsing leaves alone. - Use the URL-safe alphabet —
-and_instead of+and/— which removes the ambiguity at the source rather than escaping around it. That alphabet, and where it comes from, is covered in URL-safe Base64 vs standard Base64.
The order of operations is worth writing down, because it is easy to get one step wrong and hard to see: bytes → Base64 → percent-encode → query string, and in reverse query string → percent-decode → Base64-decode → bytes. A percent-decode after a Base64-decode is a bug that only shows up on inputs containing %.
encodeURI vs encodeURIComponent
JavaScript offers both, and the difference is exactly the "per component" rule.
encodeURIComponent('a&b=c d'); // 'a%26b%3Dc%20d'
encodeURI('https://x.test/?a&b=c d'); // 'https://x.test/?a&b=c%20d'
encodeURI is designed to take a whole URL you already assembled. It therefore leaves &, =, ?, / and : alone — they are the URL's structure, and encoding them would destroy it. Used on a value that is then pasted into a query parameter, it leaves the & alone, and the & splits your value into a second parameter. That is a straightforward injection: user text becomes URL syntax.
encodeURIComponent is designed for one piece of a URL, so it encodes everything structural. It is the function to reach for by default, and one of the few places a rule of thumb is genuinely correct: if you are not certain, encode harder — an over-encoded unreserved character is harmless, an under-encoded & is a parameter injection.
Note also what encodeURIComponent does not encode: ! ' ( ) * - . _ ~ and alphanumerics are left as-is, because RFC 3986 marks them as unreserved. That is correct, and it is why two different producers can emit different-looking but equivalent strings — a fact that matters later, when a signature or cache key is computed over one of them.
Plus and space, the rule behind the folklore
+ meaning space is not a URL rule. It is a rule of application/x-www-form-urlencoded, the format used by HTML forms and, by convention, by query strings on many servers. Consequences worth knowing:
- In a query string,
+is usually read as a space. In a path, it is a literal plus. So?q=C%2B%2Band?q=C++are different on a standards-compliant parser, and%2Bis the only reliable way to send a literal plus in a value. encodeURIComponentproduces%20for a space, not+.new URLSearchParams({ q: 'a b' }).toString()producesq=a+b— same meaning, different spelling. Both decode correctly under a form parser; neither decodes correctly under a strict URL parser that treats+literally. Which one you get is decided by the library, not by you, which is why round-tripping a value through two libraries is a reliable way to grow an extra space.- When a value is already encoded and you encode it again, spaces and
+do not survive a naive double transform. That is the topic of percent-encoding: the reserved characters that break URLs, which goes through the character-level rules this section summarises.
When neither belongs in a URL
Both transform encodings cost something, and neither is a solution to "the URL is too long". Percent-encoding can triple the length of ASCII input and worse for non-ASCII; Base64 adds a flat 33% and then gets percent-encoded, so a binary payload in a URL can end up nearly twice its size in characters before anything else is counted — see why Base64 costs 33% for the arithmetic.
Practical limits are the reason this matters: older browsers capped URLs near 2,083 characters, many servers and proxies cap the request line at 8 KB, and every URL ends up in logs, analytics and referrer headers, where a megabyte of encoded data is not something anyone wants to retain. If the payload is more than a token or a short parameter, it belongs in the request body, not in the query string. If it is binary and it must be in a URL — a signed thumbnail, an inline data URI — that is a design decision with a cost, and Base64 plus percent-encoding is that cost.
The decision list
- User text going into a URL component →
encodeURIComponent, or the equivalent in your language. - An assembled URL that needs non-ASCII or spaces escaped →
encodeURI, knowing it leaves structure intact. - Binary going through a text channel → Base64; if the destination is a URL, use the URL-safe alphabet or percent-encode the result.
- Text going into a form field or a body → percent-encode only if the field is defined as URL-encoded; otherwise send it as-is.
- Anything large → the body, not the URL.
- Base64 inside a JSON string inside a URL — the case that catches everyone twice — encode the Base64 first, then the JSON, then the URL component, and decode in exact reverse. The URL encoder and decoder shows the character-by-character result of each step, which is the fastest way to see which layer is eating your
+.