URL Encoding Guide: Percent-Encoding, encodeURIComponent and Form Traps
You add a search filter to a URL and the page stops loading. You pass a query parameter containing an email address and the API returns a validation error about the @ sign. You build a redirect URL by concatenating strings, and it works in every test until a customer's name contains an ampersand.
Every one of those failures is the same mistake: putting a value that has never been encoded into a place that expects an encoded value. URL encoding looks like a one-line utility call, which is exactly why it is worth understanding rather than memorising — the rules differ between encoding layers that all look identical in the URL bar.
The problem: a URL has structure, and unencoded values destroy it
A URL is not an opaque string. It has a grammar — scheme, host, path, query, fragment — and the delimiters expressing that grammar (/, ?, &, =, #, +) are only meaningful in their own layer. The moment a user-supplied value crosses into a URL unencoded, its delimiters get promoted to structure:
const url = `https://app.example.com/callback?next=${userSuppliedValue}`;If that value is https://evil.example/steal, the result declares a next parameter holding an absolute URL, and many backends will happily redirect there — an open redirect created by nothing more than a missing encodeURIComponent. The same concatenation breaks benignly when the value contains &page=2: the query ends early and a spurious page parameter appears.
The mechanism is percent-encoding: any byte outside the allowed set is replaced by % plus two uppercase hex digits. A space becomes %20, & becomes %26, / becomes %2F. A non-ASCII character is encoded per UTF-8 byte — 王 becomes %E7%8E%8B, three triplets for three bytes.
That last detail explains the most confusing symptom. Because one character maps to multiple bytes, a correctly encoded Chinese character is nine characters long. If a server decodes with a single-byte charset instead of UTF-8, those nine bytes become nine Latin characters or a wall of replacement characters. The percent-encoding did its job; the charset on the other end did not.
The solution: encodeURIComponent for values, encodeURI for structure
JavaScript gives you two functions, and choosing wrong is the most common encoding mistake in the language:
encodeURIComponent | encodeURI | |
|---|---|---|
| Intended for | A value — one component | A whole URI |
Escapes & = ? # + / | Yes | No |
| Typical use | Query values, form fields, path segments | A complete URL being assembled |
The rule: encode the values, not the URL. Build the URL with its delimiters intact and encode each piece that came from outside your code:
const base = 'https://app.example.com/callback';
const params = new URLSearchParams({ next: userSuppliedValue }).toString();
// → next=https%3A%2F%2Fevil.example%2Fsteal
const url = `${base}?${params}`;2
3
4
Using URLSearchParams rather than hand-built template strings is the durable fix — it handles escaping, correct form encoding and parameter ordering in one step. Reserve encodeURI for the rare case of encoding a complete URL that is itself a value inside another parameter.
The + trap
Two encodings are in common use and they disagree on how to write a space. Percent-encoding produces %20 — what encodeURIComponent and encodeURI emit, and what you want in a query string. Form encoding (application/x-www-form-urlencoded) produces +, which is what HTML form submission and URLSearchParams.toString() emit. Both are correct in their own context; the failure happens when data crosses between them.
A + in a form becomes a space. A user typing C++ into a search box makes the browser post q=C%2B%2B — correct. But a hand-built request sending q=C++ raw is decoded by most servers as C , because the plus signs read as spaces. This is why concatenated query strings must escape a literal + as %2B.
The same character also means different things per layer: in a path segment + is a literal plus and spaces must be %20. That is why "just encode the whole thing" is not a plan.
Encoding is not a security control
Percent-encoding is reversible and keyless: decodeURIComponent undoes it completely. Encoding a value before putting it in a URL does not make it safe — it only guarantees the URL's structure survives. Validation must happen after decoding, because %3Cscript%3E is just <script> once decoded, and any allowlist check applied to the encoded string can be bypassed by encoding the disallowed characters.
Tool walkthrough: finding which layer is broken
When a URL misbehaves, stop guessing and look at the two sides separately. The URL Encoder & Decoder encodes and decodes in one click, using the same encodeURIComponent your framework calls — so the output you get is the output your code produces, not an approximation.
The sequence that resolves most cases:
- Encode the raw value and compare. If it differs from what your code sent, the bug is in the encoding call — usually
encodeURIwhereencodeURIComponentwas needed. - Decode what the server received. Take the raw query string it logged and decode it. A space where the user typed
+, or replacement characters, points straight at the form-decoding or charset issue. - Check the length. Correct UTF-8 encoding of a 3-byte character is 9 characters. A value roughly three times the input is expected overhead; much larger expansion means double-encoding.
- Look for
%25. That is an encoded percent sign, which means the value was encoded twice — usually a "defensive" re-encoding in a proxy or framework.
For structural problems, the URL Parser is the complement: paste a full URL and it lists scheme, host, port, path, each query parameter and the fragment separately. It uses the browser's own parser, so the boundaries it shows are the ones the browser will enforce. That answers "what does this URL actually mean?" — and it is how you find a parameter that was never part of the query because an unescaped & ended it early.
Together they settle the classic case fast: parse the URL, read the parameter table, and if the value you expected is missing while an unexpected one appeared, the source value contained & or # and was never encoded.
FAQ
Should I use encodeURI or encodeURIComponent?encodeURIComponent for values — query parameters, form fields, individual path segments — and encodeURI only for a complete URI whose delimiters must survive. The difference: encodeURI leaves &, =, ?, # and / unescaped, so a value containing them breaks the URL's structure. In practice, prefer URLSearchParams.
Why does my + turn into a space? Form encoding represents a space as +, so a decoder applying those rules reads a literal + as a space. Percent-encoding avoids the ambiguity by writing a space as %20 and a literal plus as %2B. If your value can contain +, send %2B.
Is URL encoding enough to prevent injection? No. It is fully reversible and provides no security guarantee — it exists only to preserve URL structure. %3Cscript%3E decodes back to <script>, so validation performed on the encoded string can be bypassed. Decode first, then validate; keep encoding and sanitising as separate steps.
More developer tools on DigDevBox
- URL Encoder & Decoder — percent-encode and decode values with the same function your code uses
- URL Parser — break a full URL into scheme, host, path, parameters and fragment
- JSON to GET Params — build an escaped query string from a JSON object
- Base64 string encoder/decoder — for the cases where Base64, not percent-encoding, is the transport