JSON vs XML

Updated: August 8, 2026

JSON won the web API, XML kept the document. Understanding why explains what each is actually good at.

The Same Record, Twice

<?xml version="1.0" encoding="UTF-8"?>
<user>
  <id>101</id>
  <name>Alice Kumar</name>
  <active>true</active>
  <roles>
    <role>admin</role>
    <role>editor</role>
  </roles>
</user>
{
  "id": 101,
  "name": "Alice Kumar",
  "active": true,
  "roles": ["admin", "editor"]
}

The size difference is visible and is the most commonly cited advantage, but it is not the important one. Compression narrows it considerably, and XML's repeated closing tags compress well. The decisive difference is what happens after parsing.

The Difference That Actually Mattered

XML parses into a document tree — a generic structure of elements, attributes and text nodes. To get a value out, you traverse it and then convert whatever you find:

// XML — navigate, extract, convert
var name = xmlDoc.getElementsByTagName("name")[0].childNodes[0].nodeValue;
var id = parseInt(xmlDoc.getElementsByTagName("id")[0].textContent, 10);

// JSON — it is already the structure you wanted
var name = data.name;
var id = data.id;              // already a number

JSON's types map directly onto the types programming languages have: objects become dictionaries, arrays become lists, numbers become numbers. XML has one type — text — so every value needs converting, and the code doing the converting has to know what type to expect.

When AJAX made background data requests routine in the mid-2000s, that per-value overhead was being paid on every call in every application. That is what decided it, as covered on the history page.

Head to Head

JSONXML
Data typesSeven, explicitText only, unless a schema says otherwise
ArraysNativeImplied by repeated elements — ambiguous when there is one
AttributesNo equivalentYes — metadata without changing shape
CommentsNoYes
NamespacesNoYes
Mixed contentNoYes — markup inside running text
SchemaJSON Schema (separate standard)XSD, DTD, RELAX NG — long established
Queryingjq, JSONPath (informal)XPath, XQuery (standardised)
TransformationWrite codeXSLT

XML's Array Ambiguity

A genuine weakness worth knowing, because it bites anyone converting XML to JSON. XML has no array type — a list is just an element repeated. So a document with one item is structurally indistinguishable from a document with a single non-list value:

<roles>
  <role>admin</role>
  <role>editor</role>
</roles>
<!-- clearly a list -->

<roles>
  <role>admin</role>
</roles>
<!-- a list of one, or a single value? Nothing in the document says. -->

Converters must guess, and they guess inconsistently — sometimes producing "role": "admin" and sometimes "role": ["admin"] for the same input shape. Consumers then break whenever a collection happens to contain exactly one item, which is a memorably annoying bug to track down. JSON's arrays are explicit: ["admin"] and "admin" are different values and always were.

What XML Does That JSON Cannot

The comparison is usually framed as JSON winning, which undersells XML's genuine capabilities.

Attributes attach metadata to a value without altering the structure around it:

<price currency="INR" vatIncluded="true">1999.00</price>

In JSON this forces a choice: either the value becomes an object — changing the shape every consumer expects — or the metadata moves to sibling keys, separating it from what it describes.

{ "price": { "amount": "1999.00", "currency": "INR" } }
{ "price": "1999.00", "priceCurrency": "INR" }

Mixed content — markup interleaved with running text — has no JSON representation at all:

<p>The <em>quick</em> brown fox jumps over the <strong>lazy</strong> dog.</p>

This is why document formats stayed with XML. DOCX, SVG, EPUB and countless publishing pipelines are built on exactly this capability, and expressing it in JSON means inventing an ad-hoc node structure that reimplements XML badly.

Namespaces let vocabularies from different sources combine in one document without key collisions — essential when several organisations contribute to a shared schema. And XSLT transforms one document shape into another declaratively; the JSON equivalent is writing code.

Security

XML's larger feature surface includes some genuinely dangerous corners. XXE — XML External Entity injection — lets a document instruct the parser to read local files or make network requests:

<!DOCTYPE foo [ <!ENTITY xxe SYSTEM "file:///etc/passwd"> ]>
<user><name>&xxe;</name></user>

A related attack, the "billion laughs", uses nested entity expansion to consume unbounded memory from a few hundred bytes of input. Both are mitigated by disabling external entity resolution and DTD processing, which modern parsers usually default to — but the capability is in the format's design and has to be turned off.

JSON has no equivalent. There are no entities, no external references and no processing instructions, so a JSON document cannot instruct a parser to do anything. The residual risks, covered on the parser page, are resource exhaustion from very large inputs and prototype pollution during a subsequent merge — both milder, and neither inherent to parsing.

Choosing

Use JSON when

  • Building or consuming a web API
  • The data is records and collections
  • A browser or mobile client is involved
  • Parse speed and payload size matter
  • You want minimal tooling overhead

Use XML when

  • The data is a document, not a record
  • Mixed content or markup is required
  • Namespaces must combine vocabularies
  • XSD validation or XSLT is already in use
  • An industry standard mandates it

The honest summary: for anything resembling records crossing a network, JSON is the default and choosing XML needs a reason. For anything resembling a document — with structure inside its text, metadata on its values, or a standards body governing its shape — XML is still the better tool and is not going anywhere.

Frequently Asked Questions

Is JSON better than XML?

For web APIs, generally yes — it is lighter, parses faster, and maps directly onto the data structures programming languages already use. For documents with mixed content, metadata on values, and complex validation or transformation needs, XML remains the better tool. They solve overlapping but different problems.

Why did JSON replace XML for APIs?

Because extracting a value from XML required traversing a DOM and converting the result, while JSON arrived as the structure the code already wanted. When AJAX made frequent small data requests common, that per-call overhead became the deciding factor.

Does XML have anything JSON lacks?

Several things: attributes that attach metadata to a value without changing its shape, namespaces for combining vocabularies from different sources, mixed content for markup inside text, comments, and mature standards for validation (XSD), transformation (XSLT) and querying (XPath).

Is XML dead?

No. It is no longer the default for new web APIs, but it is deeply established in document formats such as DOCX and SVG, in publishing, in finance and healthcare messaging, and in enterprise systems like SOAP. Those uses play to strengths JSON does not have.

Related Resources

Related Resources