JSON in Python

Updated: August 8, 2026

Parsing, generating and troubleshooting JSON with Python's json module — including the type mismatches that catch people out.

The Four Functions You Need

Python's json module is in the standard library — no installation required. Its entire surface is essentially four functions, and the naming convention confuses newcomers exactly once: the trailing "s" means "string".

import json

json.loads(text)      # JSON string  → Python object
json.load(file_obj)   # JSON file    → Python object

json.dumps(obj)       # Python object → JSON string
json.dump(obj, file)  # Python object → written to a file

Mnemonic: loads loads from a string, load loads from a file. Almost every "TypeError: the JSON object must be str, bytes or bytearray, not TextIOWrapper" is someone reaching for loads when they had a file handle.

Reading JSON

import json

# From a string
raw = '{"name": "Alice", "age": 30, "active": true}'
data = json.loads(raw)
print(data["name"])        # Alice
print(type(data["active"]))  # <class 'bool'> — JSON true became Python True

# From a file — always specify the encoding
with open("data.json", encoding="utf-8") as f:
    data = json.load(f)

# From an HTTP response (requests handles the decode for you)
import requests
data = requests.get("https://api.example.com/users").json()

Parsing failures raise json.JSONDecodeError, a subclass of ValueError. It carries .lineno, .colno and .pos, which is far more useful than the message alone when you are logging the failure:

try:
    data = json.loads(raw)
except json.JSONDecodeError as e:
    print(f"Invalid JSON at line {e.lineno}, column {e.colno}: {e.msg}")

How Types Map Across

The conversion is not quite symmetrical, and the asymmetries are where bugs live:

JSONPython (decoded)Note
objectdictInsertion order preserved (3.7+)
arraylistTuples encode to arrays but decode back as lists
number (int)intArbitrary precision — Python does not lose large integers
number (real)floatStandard float precision limits apply
true / falseTrue / FalseCapitalisation differs — a frequent hand-editing error
nullNoneDifferent spelling again

One asymmetry deserves special attention. Python integers have unlimited precision, so json.loads will faithfully read a 30-digit ID that JavaScript would silently round. Your Python service can therefore be correct while a JavaScript client consuming the same endpoint is quietly corrupting the value — see the parser guide on numeric precision.

Writing JSON

import json

data = {"name": "Alice", "roles": ["admin"], "active": True}

# Compact — note the default still adds spaces after separators
json.dumps(data)
# '{"name": "Alice", "roles": ["admin"], "active": true}'

# Genuinely minimal output
json.dumps(data, separators=(",", ":"))
# '{"name":"Alice","roles":["admin"],"active":true}'

# Readable
json.dumps(data, indent=2)

# Deterministic — essential for files under version control
json.dumps(data, indent=2, sort_keys=True)

# Write to a file
with open("out.json", "w", encoding="utf-8") as f:
    json.dump(data, f, indent=2, ensure_ascii=False)

sort_keys=True is worth adopting as a habit for generated files. Without it, a dictionary rebuilt in a different order produces a completely different diff even when no values changed.

The Unicode Trap

By default json.dumps escapes every non-ASCII character, which is valid JSON but unreadable:

json.dumps({"city": "Bengaluru", "name": "José"})
# '{"city": "Bengaluru", "name": "Jos\u00e9"}'

json.dumps({"city": "Bengaluru", "name": "José"}, ensure_ascii=False)
# '{"city": "Bengaluru", "name": "José"}'

Both parse to the identical value, so this is cosmetic rather than a correctness issue. But for any file a human will open — translations, fixtures, seed data — ensure_ascii=False makes the difference between reviewable and opaque. Pair it with an explicit encoding="utf-8" when writing, or you risk a UnicodeEncodeError on systems whose default encoding is not UTF-8.

Serialising Types JSON Does Not Have

TypeError: Object of type X is not JSON serializable is the most common Python JSON error, and it is not a bug — JSON genuinely has no date, decimal, set or UUID type. You have to say what they should become.

The quick fix, for one-off scripts:

from datetime import datetime, date
from decimal import Decimal
from uuid import UUID
import json

def encode(obj):
    if isinstance(obj, (datetime, date)):
        return obj.isoformat()
    if isinstance(obj, Decimal):
        return str(obj)          # str, not float — preserves exactness
    if isinstance(obj, UUID):
        return str(obj)
    if isinstance(obj, set):
        return sorted(obj)
    raise TypeError(f"{type(obj).__name__} is not JSON serializable")

json.dumps(payload, default=encode)

The reusable version, for application code:

class AppEncoder(json.JSONEncoder):
    def default(self, obj):
        if isinstance(obj, (datetime, date)):
            return obj.isoformat()
        if isinstance(obj, Decimal):
            return str(obj)
        return super().default(obj)

json.dumps(payload, cls=AppEncoder, indent=2)

Note the Decimal handling. Converting to float is the obvious move and it is wrong for money — it reintroduces exactly the binary floating-point error you chose Decimal to avoid. Serialise as a string and parse back with Decimal(value) at the other end.

Large Files

json.load() builds the whole structure in memory, and the Python object graph typically occupies several times the file's size on disk. A 500 MB file can exhaust several gigabytes of RAM. Two approaches avoid this.

JSON Lines — one complete JSON document per line — is the simpler option, and the usual format for logs and data exports:

with open("events.jsonl", encoding="utf-8") as f:
    for line in f:
        event = json.loads(line)
        process(event)   # constant memory, regardless of file size

Streaming, when you are stuck with one enormous document, using ijson (pip install ijson):

import ijson

with open("huge.json", "rb") as f:
    for record in ijson.items(f, "data.item"):
        process(record)   # yields one record at a time

Frequently Asked Questions

What is the difference between json.load and json.loads?

The trailing 's' stands for 'string'. json.loads() takes a str or bytes containing JSON; json.load() takes a file-like object and reads from it. The same pattern applies to json.dump() and json.dumps() when writing.

Why does json.dumps fail with 'Object of type datetime is not JSON serializable'?

JSON has no date type, so Python's encoder does not know how to represent a datetime. You must convert it yourself — usually to an ISO 8601 string via .isoformat() — either by passing a default= function or by using a custom JSONEncoder subclass.

Does Python preserve the order of JSON object keys?

Yes, in practice. Since Python 3.7 dictionaries preserve insertion order, so a parsed object keeps the order it appeared in the document. The JSON specification does not guarantee this, so avoid relying on it when other systems are involved.

How do I stop json.dumps escaping non-ASCII characters?

Pass ensure_ascii=False. By default Python escapes every non-ASCII character to a \uXXXX sequence, which is valid but unreadable for names and text in most languages. With ensure_ascii=False you get the actual characters, and you should write the file as UTF-8.

How should I handle very large JSON files in Python?

json.load() builds the entire structure in memory, which can use several times the file size. For large files, either switch to JSON Lines and process one record per line, or use a streaming parser such as ijson that yields items incrementally.

Related Resources

Related Resources