Percent-Encoding: The Reserved Characters That Break URLs

2026-09-30 · 1208 words

A URL is not a string with escapes in it. It is a small grammar with several interchangeable parts — scheme, authority, path, query, fragment — and a character's meaning depends on which part it appears in. Percent-encoding exists so that data can be carried inside those parts without being mistaken for the grammar itself. When that goes wrong, nothing reports an error: the value simply arrives shorter, split in two, or pointing somewhere the server never sees.

The two sets, and one detail about them

RFC 3986 divides characters into two groups.

Unreserved — A–Z, a–z, 0–9, -, ., _, ~. These never need encoding, and encoding them anyway creates a second spelling of the same URL. That looks harmless until two spellings reach a cache, a signature calculation or a URL-based access rule, and are treated as two different resources.

Reserved, in two flavours:

Reserved does not mean "always encode". It means "this character has a job here". : is a delimiter in the authority (host:port) and perfectly ordinary data in a path segment. & is a separator between query parameters and perfectly ordinary data inside a path. The rule is always relative to the component you are writing.

One detail that trips parsers: percent-encoded triplets are case-insensitive — %2F and %2f mean the same octet — but string equality in your code is not. Canonicalise to uppercase hex before comparing, hashing or signing, or two identical URLs will produce two different signatures.

The characters that cause incidents

& in a query value. The value ends at the &; everything after it becomes another parameter. Symptom: a long token arrives truncated at an arbitrary point, and an extra parameter appears in the request that no client intended to send.

# anywhere in a query value. Everything from the # onward becomes the fragment, and a browser never sends the fragment to the server. So the server sees the value cut at that point, and the browser's own address bar shows the rest — the single most confusing version of this bug, because the URL looks complete to everyone who inspects it locally.

+ in a query value. Under application/x-www-form-urlencoded rules it means a space, so C++ becomes C . In a path it is a literal plus. Send a literal plus as %2B if you want it to survive being handled by two different libraries.

% on its own. A literal percent sign must be %25. A single % followed by something that is not two hex digits makes the URL invalid, and strict parsers reject the whole thing rather than the character — which is why an input like 100% or %AB in a template can break an otherwise fine request.

Space. %20 in a path and in a standards-compliant query; + in form-encoded data. Both are widely accepted on the way in, and that mutual tolerance is precisely why the bug survives so long: it works with the parser you tested and fails with the one in production.

/ inside a path segment value. Encode it as %2F and the value no longer looks like two segments — but many servers decode the path before routing, so %2F is turned back into a separator and your value splits anyway. Some servers reject encoded slashes outright, and the reason is security: %2F and %2E are the building blocks of path traversal attempts, so proxies and application servers have strong opinions about them. The honest conclusion is that a value containing a slash does not belong in a single path segment.

? in a query value. Legal in a query by the grammar, and best encoded anyway, since a second ? invites a second round of parsing in intermediaries.

Non-ASCII is bytes, not characters

This is where percent-encoding stops looking like an escaping scheme and starts looking like an encoding.

A %XX triplet carries one byte, not one character. A character outside ASCII has to be turned into bytes first, and the only interoperable choice is UTF-8:

If the producer used a different legacy encoding, the triplets are still perfectly well-formed and still decode to something — just not to the text that was sent. A GBK-encoded Chinese string arriving at a UTF-8 decoder produces a page of replacement characters, and nothing in the request looks wrong. So "the encoding is valid" and "the characters are correct" are separate questions, and the second one needs the producer's encoding to be known.

The related rule for implementers: always percent-encode bytes, never code points and never characters. In languages where a string is a sequence of UTF-16 code units, encoding a character directly is how astral-plane characters get written as surrogate halves, which no decoder will accept.

A decoder should refuse, not guess

Decoding is where the damage becomes permanent, because a decoder that is generous with invalid input turns a visible bug into a plausible value.

A decoder that does the opposite — normalising, repairing and reporting success — is the reason two systems can disagree about the same URL while both claim to have decoded it. The strict-but-explicit behaviour is what the URL encoder and decoder implements: it shows which characters it changed, keeps UTF-8 byte handling separate from the text, and points at the exact position of a malformed sequence instead of quietly dropping it.

The rules worth keeping

  1. Encode per component, not per URL. &, =, # and + are data in one part and syntax in another.
  2. Never decode twice. %2520 is a percent sign followed by 20, not a space; if a second decode "fixes" it, the producer double-encoded, and that is their bug to fix.
  3. Compare canonical forms: uppercase hex triplets, no needlessly encoded unreserved characters, one spelling per resource.
  4. Treat non-ASCII as UTF-8 bytes and say so explicitly; where the sender's encoding is unknown, ask rather than assume.
  5. Prefer refusing malformed input to repairing it — this is the difference between a decoder and a data corruption device.

The comparison with Base64 covers which of the two layers a given payload needs; when the payload is a token that must survive a query string untouched, the URL-safe alphabet is usually the better starting point than escaping your way out of an alphabet that already contains + and /.

Related reading

Try the Base64 to Image converter →