Skip to main content
ToolsBay

DEVELOPER TOOLS

JSON vs XML: What a Round Trip Between Them Actually Loses

8 min read · ToolsBay editorial · Published · Updated

Just want to do it now?

Convert JSON structures into equivalent XML markup.

Open JSON to XML

Most comparisons of these two formats put a syntax sample of each side by side, note that XML has more characters in it, and stop. That is not the decision anyone faces. The real situation: something upstream speaks one format, something downstream speaks the other, and you need to know what breaks in the middle.

This site ships both directions — XML to JSON and JSON to XML — so rather than describe the gap, this post runs a document through it. Every output below is what those two pages actually produce.

A document that uses XML properly

<?xml version="1.0" encoding="UTF-8"?>
<invoice xmlns:acc="http://example.com/acc" id="A-1001">
  <!-- exported 2026-09-01 -->
  <acc:total currency="USD">42.50</acc:total>
  <note>Ship <em>before</em> Friday.</note>
  <memo></memo>
</invoice>

Six lines, and five things JSON has no equivalent for: a namespace binding, attributes, a comment, mixed content — text with an element sitting inside it — and an empty element. Converted, that becomes:

{
  "invoice": {
    "_attributes": { "xmlns:acc": "http://example.com/acc", "id": "A-1001" },
    "_comment": " exported 2026-09-01 ",
    "acc:total": { "_attributes": { "currency": "USD" }, "_text": "42.50" },
    "note": { "_text": ["Ship ", " Friday."], "em": { "_text": "before" } },
    "memo": {}
  }
}

The XML declaration comes through too, as a _declaration key trimmed above for width. Nothing was thrown away. Almost nothing survived intact either.

Attributes land in a reserved key

_attributes and _text exist because an element can carry both at once and JSON needs somewhere to put each. As the tool page says: a conversion keeping only one of them loses half the document.

What no page can promise is portability, because there is no standard here. Other libraries use @, $, -, or a flat merge that collides silently when an attribute and a child element share a name. Cross into a system built against a different convention and you are writing a translation layer — and that layer, not the format, is where the bugs live.

Mixed content loses its position

Look at note. The two text fragments sit in an array, and em is a sibling key. Nothing in that object records that <em> sat between them. Reassemble it naively and "Ship before Friday." comes back as Ship Friday.before.

For data-shaped XML this never comes up. For prose-bearing XML — DocBook, TEI, XHTML fragments in a CMS, anything with inline emphasis or links — it is the whole ballgame, and it fails quietly rather than throwing. If your XML has markup inside sentences, a JSON intermediate is the wrong pipeline, not a slower one.

Namespaces survive as characters, not as meaning

acc:total stays acc:total, and xmlns:acc stays an attribute. The prefix is preserved; the binding is not resolved. Two documents that mean exactly the same thing while binding different prefixes to the same URI produce two different JSON objects, and no consumer can tell they match.

Coming back is worse. XML element names cannot contain arbitrary characters, so the JSON to XML converter sanitises them, and a colon is not in the set it allows. xmlns:acc returns as <xmlns_acc>: the namespace stops being a namespace and becomes an ordinary element with an odd name.

Empty, missing and null all collapse

<memo></memo> converts to {}. Going the other way, JSON null becomes <note/>, which converts back to {}. An empty string arrives at the same place. Three distinguishable states in the source, one representation in the target. If your schema treats "field present but blank" differently from "field absent" — and billing and address systems usually do — that distinction does not survive the trip.

The array-of-one problem

The most common production failure in XML-to-JSON pipelines, seen plainly:

{"r":{"i":{"_text":"1"}}}
{"r":{"i":[{"_text":"1"},{"_text":"2"}]}}

Same schema, same tag. One <i> gives an object; two give an array. Your code works against every test fixture and then a customer with exactly one order line arrives. The tool's own FAQ is blunt about it: normalise after conversion rather than trusting the input to carry two or more.

Going the other way loses different things

Feed this JSON to the converter:

{ "order": { "total": 42.50, "active": true, "count": 3 }, "note": null }

XML has no types. Everything becomes text, so true and 3 come back as the strings "true" and "3", and rebuilding the original types is guesswork the converter declines to do.

42.50 is subtler. It is already 42.5 before XML is involved, because parsing it into a JavaScript number drops the trailing zero — XML, storing it as text, would have kept it. The same asymmetry bites harder on identifiers: 9007199254740993 parses to 9007199254740992, because IEEE-754 doubles run out of exact integers at 2^53. Snowflake IDs and 64-bit database keys are exactly that size. XML hands those digits back untouched; JSON rounds them silently unless the producer had the sense to quote them.

Order is the last one. RFC 8259 defines a JSON object as an unordered collection. JavaScript preserves insertion order in practice — except for integer-like keys, which are hoisted to the front in ascending order:

Object.keys(JSON.parse('{"b":1,"2":2,"a":3,"1":4}'))
// ["1", "2", "b", "a"]

XML guarantees document order among siblings. JSON does not, and the exception is not where anyone expects it.

The size argument is mostly wrong

XML's verbosity is the point everyone reaches for, so it is worth measuring. Five hundred inventory records, seven fields each, neither side indented:

| Format | Raw | gzip -9 | brotli |

| --- | --- | --- | --- |

| JSON | 67,081 B | 6,933 B | 4,091 B |

| XML | 94,127 B | 7,250 B | 4,313 B |

XML is 40% larger uncompressed. After compression it is 5% larger — 317 bytes on a 7 KB response. Repeated closing tags are the most compressible bytes in existence; removing that exact pattern is what gzip's back-references do.

If your transport compresses, and it almost certainly does, choosing JSON to save bandwidth optimises something already dealt with. The real costs of XML are parse time, parser surface area, and the code needed to get one value out. Those are good reasons. Payload size is not.

The reason a lot of teams actually left

XML's specification includes a document type definition, and a DTD can declare entities. One of those can point at a file:

<!DOCTYPE f [<!ENTITY xxe SYSTEM "file:///etc/passwd">]>
<r>&xxe;</r>

A parser that resolves external entities will read that file, or make that request, on behalf of whoever sent the document. XML External Entities had its own entry in the OWASP Top 10 in 2017; the 2021 revision folded it into Security Misconfiguration, since the fix is nearly always a parser flag someone forgot to set.

The second one needs no external reference at all. Define an entity as ten copies of a smaller entity, ten levels deep, and a few hundred bytes expand to a billion strings. That is the billion laughs attack, and it exhausts an eagerly-expanding parser using no bandwidth worth mentioning.

Both live in the DTD. JSON has no entity mechanism, no external references and no document type, which is why neither attack has a JSON analogue — and why "we moved to JSON" was so often a security decision wearing an ergonomics costume. JSON's own hazard sits downstream instead: parsing {"__proto__": {"x": 1}} creates an ordinary own property and pollutes nothing. The damage happens later, in a hand-rolled deep merge that assigns it.

Worth knowing what our converter does here, because it is a trade-off rather than a win. It rejects any entity it does not recognise: the five predefined entities and numeric character references work, while a custom &lol; returns "Invalid character entity" and the conversion fails outright. Nothing expands and no external DTD is fetched, since the parse happens in your tab and there is no server to fetch from. The cost is that a document legitimately relying on DTD entities will not convert at all.

Schemas are where XML still wins outright

XML has three validation languages, all finished: DTD, built into the XML specification itself; XSD, a W3C Recommendation since 2001 with version 1.1 following in 2012; and RELAX NG, standardised as ISO/IEC 19757-2. They express structure, datatypes, cardinality and co-occurrence constraints, and long-established tooling generates code from them.

JSON Schema is the counterpart, and the current release is draft 2020-12 — still an IETF draft series rather than a ratified standard, though implemented nearly everywhere. OpenAPI 3.1 aligns with it, which is why most people meet JSON Schema through an API description. It is sufficient for most contracts. It is also younger and less settled than XSD, and pretending otherwise helps nobody.

So which one

If you are producing data for a SOAP endpoint, an RSS or Atom feed, a sitemap, an office document format or a government integration, XML is not a choice you get to make. If you are shipping a new HTTP API, JSON is the default. Neither case is interesting.

The interesting case is the middle: a document that has to cross between them. Pick the direction that loses the least. Data-shaped XML — flat records, attributes carrying scalars, no inline markup — converts cleanly enough that a conversion plus a normalising step is a fine pipeline. Prose-shaped XML does not, and no converter will change that. Going the other way, decide how nulls, numbers and single-item arrays should look on the far side before you convert, because the converter will decide for you otherwise.

When a conversion goes wrong, checking that the JSON parses at all is the cheap first move. JSON Validator names the line and column the parser gave up on, which is usually enough to identify which stage mangled it.

Tools covered in this guide