Base64, URL Encoding, and HTML Entities: What Every Developer Should Know
If you have spent any time building web applications, you have run into encoding problems. A URL breaks because of a stray ampersand. An HTML page renders a < symbol as an arrow. An API returns a string that looks like SGVsbG8gV29ybGQ= and you have no idea what it contains. These are encoding problems, and three encoding schemes show up more often than any others: Base64, URL encoding, and HTML entities.
This guide explains what each one does, shows you concrete examples, and points you to tools that let you encode and decode instantly β no installation required.
What Is Base64?
Base64 is a way to represent binary data using only printable ASCII characters. It takes any sequence of bytes and maps them to a 64-character alphabet: AβZ, aβz, 0β9, +, and /. The output is always text, which makes it safe to embed in places that cannot handle raw binary β email attachments, JSON payloads, CSS data: URLs, and JWT tokens all rely on Base64.
Example β encoding "Hello World":
- Input:
Hello World - Base64 output:
SGVsbG8gV29ybGQ=
The trailing = is padding. Base64 always produces output whose length is a multiple of 4, so = signs fill the gap.
When to use it:
- Embedding images or fonts directly in CSS or HTML (
data:image/png;base64,...) - Encoding binary data for JSON APIs that only accept text
- Reading or writing JWT tokens (the header and payload sections are Base64url-encoded)
Base64 is not encryption. Anyone who sees the string can decode it in seconds. Use it for encoding, not for hiding information.
Try it now with the FileCrank Base64 Encode/Decode tool.
What Is URL Encoding (Percent Encoding)?
URLs can only contain a limited set of characters safely. Spaces, ampersands, equal signs, slashes, and non-ASCII characters all carry structural meaning in a URL β or simply are not valid. URL encoding (officially called percent encoding) replaces unsafe characters with a % sign followed by two hexadecimal digits representing the character's UTF-8 byte value.
Example β encoding "Hello World":
- Input:
Hello World - URL-encoded output:
Hello%20World
The space becomes %20. An ampersand & becomes %26. A forward slash / becomes %2F.
When to use it:
- Building query strings:
?q=hello%20world&lang=en - Passing special characters in form data
- Constructing API requests that include user-supplied input
Modern JavaScript gives you encodeURIComponent() for this, but when you need to inspect or debug an encoded URL, an online tool is faster than firing up a console.
Decode or encode any URL string with the FileCrank URL Encode/Decode tool.
What Are HTML Entities?
HTML uses certain characters as syntax. The angle brackets < and > delimit tags, the ampersand & starts an entity reference, and quotation marks " delimit attribute values. If your content contains any of these characters literally β say, you want to display source code on a page β you must escape them so the browser does not interpret them as markup.
HTML entities replace those characters with named or numeric references:
Example β encoding "Hello World" as an HTML heading:
If you want to display the literal string <h1>Hello World</h1> as text on a page, you write:
<h1>Hello World</h1>
When to use it:
- Displaying code samples in documentation or blog posts
- Preventing XSS vulnerabilities by escaping user input before inserting it into HTML
- Inserting typographic symbols (em dashes
—, copyright©, non-breaking spaces )
Use the FileCrank HTML Entities tool to escape or unescape HTML in seconds.
Comparing All Three: A Side-by-Side View
Here is the same input string processed by each encoding scheme:
They solve different problems and are not interchangeable:
- Base64 handles binary-to-text conversion. Use it when you need to embed binary data in a text-only context.
- URL encoding handles characters that are illegal or ambiguous in a URL. Use it when constructing URLs or query strings.
- HTML entities handle characters that are reserved by HTML syntax. Use them when outputting text inside HTML documents.
Common Mistakes
Mistake 1: Using Base64 for security. Base64 output looks like gibberish, but it is trivially reversible. Do not use it to "hide" passwords or tokens. Use proper encryption or hashing instead.
Mistake 2: Double-encoding URLs.
If you call encodeURIComponent() on a string that is already percent-encoded, you get %2520 instead of %20. Decode first, then encode.
Mistake 3: Forgetting to escape HTML in template strings. Inserting user-supplied data directly into HTML without escaping is the number-one cause of XSS vulnerabilities. Always escape before rendering.
Try the Tools
All three encoding tools are free, run entirely in your browser, and require no account:
- Base64 Encode/Decode β encode text or files to Base64 and back
- URL Encode/Decode β percent-encode and decode URL strings
- HTML Entities β escape and unescape HTML special characters
Bookmark them. You will reach for at least one of them every week.
Base64 Is Encoding, Not Encryption
This point appears in the Common Mistakes section above, but it deserves a full explanation because the misconception is widespread and genuinely dangerous.
Encryption takes plaintext and transforms it into ciphertext that cannot be read without a key. Base64 takes binary data and transforms it into ASCII text that anyone can reverse in milliseconds β no key required. The only reason Base64 output looks unreadable is that most people are not used to seeing it. A tool, a browser console, or a single line of Python will decode it instantly.
import base64
base64.b64decode("SGVsbG8gV29ybGQ=")
# b'Hello World'The practical risk: developers sometimes Base64-encode sensitive values β API keys, passwords, internal tokens β and then embed them in client-side code or public repositories, reasoning that the encoded form "hides" the value. It does not. Any attacker who finds that string will decode it in seconds.
The correct approach depends on what you are protecting:
- Passwords: use a proper hashing algorithm (bcrypt, Argon2). Never store plaintext or Base64-encoded passwords.
- Secrets in transit: use TLS (HTTPS). Base64 adds nothing on top of TLS, and nothing without it.
- Data you need to keep confidential: use AES-256 or a similar symmetric cipher, managed through a dedicated secrets service.
Base64 is a serialisation format, not a security mechanism. Treat it the same way you treat JSON β useful for moving data around, not for protecting it.
JWT Tokens and Base64URL Encoding
JSON Web Tokens (JWTs) are one of the most common places developers encounter Base64 in production. Understanding the structure demystifies a lot of authentication debugging.
A JWT looks like three chunks of text separated by dots:
eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiJ1c2VyXzEyMyIsInJvbGUiOiJhZG1pbiIsImlhdCI6MTcxMDAwMDAwMH0.SflKxwRJSMeKKF2QT4fwpMeJf36POk6yJV_adQssw5c
Each section is independently Base64URL-encoded (a variant that replaces + with - and / with _, and omits padding, so the string is safe inside a URL without further encoding):
- Header β algorithm and token type. Decode it and you get
{"alg":"HS256","typ":"JWT"}. - Payload β the claims: user ID, role, expiry, anything the server chose to include. Decode it and you get
{"sub":"user_123","role":"admin","iat":1710000000}. - Signature β a cryptographic signature over the header and payload, produced with a secret key. This part you cannot forge without the key.
The critical point: the header and payload are not encrypted. Anyone who intercepts a JWT can decode the first two sections and read their contents. This is intentional β JWTs are designed to be readable by any party that receives them. The signature only proves that the token was issued by someone who holds the signing key; it says nothing about confidentiality.
This means you should never put sensitive information β passwords, payment details, private personal data β inside a JWT payload. The payload is visible to the user, to any CDN, to any logging system, and to any attacker who intercepts the token in transit without TLS.
To inspect a JWT manually, paste it into the FileCrank Base64 Encode/Decode tool and decode each section. You will see the JSON structure immediately.
Why Base64 Increases File Size by Around 33 %
Base64 is not free. Every time you encode binary data as Base64, the output is roughly one third larger than the input. Understanding why helps you decide when the trade-off is worth making.
Binary data uses all 256 possible values for each byte (8 bits). Base64 only uses 64 characters (6 bits of information per character). To represent 3 bytes of binary data (24 bits), Base64 needs 4 characters (4 Γ 6 = 24 bits). So 3 bytes becomes 4 characters β a 33.3 % expansion.
Input: 3 bytes β Output: 4 Base64 characters
Input: 1 MB β Output: ~1.37 MB
Input: 100 KB β Output: ~137 KB
For small assets this is usually acceptable. For large files it adds up quickly:
- A 5 MB background image encoded as a Base64 data URI adds roughly 1.7 MB to your CSS file.
- That CSS file must be downloaded before the page renders, delaying your Largest Contentful Paint.
- It also bypasses the browser's separate image caching β the image is re-downloaded with the stylesheet on every page load.
The guidance most performance teams follow: Base64-embed assets smaller than roughly 4β8 KB, where the HTTP request overhead outweighs the size cost. For anything larger, serve the file separately and let the browser cache it.
Data URIs: Embedding Assets Directly in CSS and HTML
Data URIs are the main production use case for Base64 in front-end development. Instead of linking to an external file, you embed the file's contents directly in the attribute or stylesheet as a data: URL.
The syntax is:
data:[mediatype][;base64],<data>
In HTML β embedding a small PNG inline:
<img src="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAUA..." alt="Icon">In CSS β using a Base64-encoded SVG as a background:
.icon {
background-image: url("data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL...");
}Common use cases:
- Small icons in component libraries that you want to ship as a single CSS file with zero external dependencies.
- Email templates, where external image requests are often blocked by mail clients and must be embedded inline.
- Favicons inlined into the
<head>to avoid an extra request. - Placeholder images shown while the real image loads (a low-resolution, heavily compressed version encoded as Base64 and rendered immediately).
Limitations to keep in mind:
- Embedded assets cannot be cached independently. Every time the CSS file changes, the browser re-downloads the embedded image as well.
- Base64 data URIs cannot be lazy-loaded. The browser decodes them during initial CSS parsing.
- Some mail clients impose a size limit on embedded images and will strip anything above roughly 100 KB.
For a deeper look at how browsers handle text encoding underneath all of this β the UTF-8 layer that sits below Base64 and URL encoding β see our guide to Unicode and UTF-8 character encoding.
If you are deciding whether to encode files client-side before processing, our guide on why browser-based file tools are safer explains how in-browser encoding keeps your data off third-party servers.