XML URL Decoder Tool
Decodes XML entities and URL percent-encoding, alone or in either order, with %XX bytes read as UTF-8. It handles the five predefined XML entities plus decimal and hexadecimal character references, and each pass removes one layer, so %2520 becomes %20. Malformed percent signs, non-UTF-8 bytes, and references to invalid XML characters stay in place and are listed.
This tool sends nothing over the network. Everything you enter is processed on your device and never reaches our servers.
Documentation
An XML and URL decoder reverses two escaping schemes that protect special characters in transit. XML escaping replaces markup characters with entity references so a parser reads them as text, and URL percent-encoding (RFC 3986) replaces bytes that a URL cannot carry with a percent sign and two hex digits. The two often stack: a URL placed inside an XML or HTML attribute is percent-encoded first and then has its ampersands escaped as &.
XML 1.0 predefines exactly five named entities: & for &, < for <, > for >, " for a double quote, and ' for an apostrophe. Any other character can be written as a numeric character reference, decimal (é) or hexadecimal (é), and both of those mean é, code point 233. HTML names such as and © belong to HTML rather than XML and pass through unchanged. A numeric reference to something XML does not allow as a character, such as code point 0, a lone surrogate between D800 and DFFF, or anything above 10FFFF, is left as written and listed beside the result.
Percent-decoding turns each %XX into the byte XX and reads the bytes as UTF-8, so %C3%A9 is é and %F0%9F%9A%80 is a single emoji. A percent sign not followed by two hex digits, as in 100%, is left alone, and a byte that cannot start a valid UTF-8 sequence, such as %E9 from an older Latin-1 encoder, keeps its %XX form and is listed. The plus sign means a space only in HTML form encoding (application/x-www-form-urlencoded); in a URL path a + is a literal plus, which is why Settings offers form decoding as an option rather than applying it by default.
The four modes decode XML only, URL only, XML then URL, or URL then XML, and the order matters when layers are nested. URL then XML turns %26lt%3B into <, while XML then URL stops at <, because that entity only appears after the XML stage has already run. Each stage removes exactly one layer, so double-encoded input such as %2520 becomes %20 and &lt; becomes <; decoding the result again removes the next layer. The count beside the result is the number of entities and percent-encoded bytes decoded.
The string caf%C3%A9%20%26amp%3B%20th%C3%A9 in URL then XML mode first percent-decodes eight bytes: C3 A9 twice for é, 20 twice for the spaces, 26 for the ampersand, and 3B for the semicolon, giving café & thé. The XML stage then resolves the one entity, & to &, for a result of café & thé and a count of 9.
Encoded text appears across web development, data processing, API debugging, and content management. Decoding these sequences by hand is tedious and error-prone, especially when XML and URL encoding overlap in the same string.
- API Debugging: API responses that return XML-escaped query strings, such as
redirect=https%3A%2F%2Fexample.com%2Fpage%3Fid%3D5&token=abc, decode in XML then URL order to show the real parameter values and expose a malformed request. - Web Development: HTML fragments pulled from server logs or template engines, where ampersands appear as
&and angle brackets as<or>, come back as the original markup for review, shown as text rather than rendered. - RSS and Atom Feeds: A feed that stores an HTML description as escaped XML text escapes the links inside it twice, so an ampersand in a link appears as
&amp;. One XML pass per layer recovers the destination URL for validation or migration. - Database Cleanup: Content that was entity-encoded before insertion, such as user comments or product descriptions containing
"and', decodes to clean text for export or display in a new system. - Email Template Inspection: Tracking links in HTML email source stack percent-encoding inside entity encoding; decoding both layers reveals the actual redirect destination and its parameters.
- Log Analysis: Server and application logs that record request URIs in percent-encoded form, sometimes mixed with XML-escaped ampersands, turn into readable paths and parameter lists for troubleshooting. Form submissions in those logs use + for spaces, which the form encoding option handles.
- Content Migration: Content exported from CMS platforms that store numeric character references in database fields decodes to raw characters before import into a platform that expects them.
- Security Auditing: Encoded payloads in penetration testing reports or WAF logs, such as
<script>, decode to the original input that triggered an alert, and the decoded markup is shown as inert text.
Inputs, outputs, and what the XML URL Decoder Tool computes
What the XML URL Decoder Tool asks for and what it returns, as a plain list. Defaults, units, and ranges are the ones the form loads with.
Inputs
- Encoded Text
- Decode Mode · default: XML Entities Only
- Treat + as a space (HTML form encoding) · default: off
- Trim leading and trailing whitespace from result · default: off
- Decode a paste immediately (skip the half-second typing delay) · default: off
- Decoded Result (copyable)
Controls
Decode · Reset
Example
The string caf%C3%A9%20%26amp%3B%20th%C3%A9 in URL then XML mode first percent-decodes eight bytes: C3 A9 twice for é, 20 twice for the spaces, 26 for the ampersand, and 3B for the semicolon , giving café & thé.