Skip to main content
Dev Tools

JSON Validation Mistakes That Break Production Pipelines

Every developer has lost an hour to a JSON error that a validator would have caught in five seconds. The failure is rarely the syntax itself. It is the category of error that syntax checks do not touch: trailing commas that pass a lenient parser but break a strict one, numbers that silently lose precision, keys that collide after case normalization. Here are the mistakes that show up most often in real pipelines, and how to catch them before they reach production.

Trusting the parser instead of validating

JSON.parse() and its equivalents answer one question only: is this syntactically valid JSON? They say nothing about whether the structure matches what your code expects. A payload like {"items": []} parses perfectly and then crashes your renderer, which assumed at least one item. The fix is two-layer validation: parse first, then validate shape. JSON Schema remains the standard tool here — define required fields, types, and value ranges once, and run the schema check at every trust boundary: API ingress, queue consumers, file imports.

A practical rule: every byte of JSON entering your system from outside gets schema-validated, no exceptions. Webhooks from third parties are the worst offenders — vendors change payload shapes without notice, and your consumer discovers it at 3 a.m.

Duplicate keys and last-one-wins surprises

RFC 8259 says object keys SHOULD be unique, but parsers are free to accept duplicates. Most do, silently keeping the last occurrence. That means {"amount": 100, "amount": 999} parses as 999 in most runtimes — with no warning. If that JSON came from a merged config or a hand-edited file, the first value was probably the intended one. Validators like a JSON validator flag duplicate keys explicitly, which is why running untrusted JSON through one before processing is cheap insurance. Some strict parsers can be configured to reject duplicates outright; do it for anything machine-to-machine.

Number precision loss

JSON numbers are not integers and not floats — they are an abstract grammar. The moment a parser maps them onto a concrete type, precision rules kick in. The classic case: JavaScript reduces everything to IEEE 754 doubles, so 1234567890123456789 becomes 1234567890123456700. Database IDs, order numbers, and snowflake-style timestamps all exceed 2^53 and get corrupted in transit if anything in the chain re-serializes them. Twitter famously shipped IDs as both id and id_str for exactly this reason.

Practical defenses: keep large identifiers as strings in JSON, and never round-trip a payload through a runtime that re-parses numbers unless you control the precision. When debugging a value that changed between producer and consumer, compare the raw bytes of the payload, not the parsed values.

Assuming valid JSON means correct encoding

A parser accepts escaped Unicode happily. The problems start earlier and later: a BOM at the start of a "UTF-8" file breaks some parsers (RFC 8259 forbids it, real-world tools still add it), and unpaired surrogates like \uD834 parse on lenient engines and explode when the string hits a database with strict UTF-8. If you consume JSON from third-party sources, normalize encoding first, then parse. After parsing, watch for mojibake — é where é should be — which means double-encoded text somewhere upstream. The Unicode normalizer catches these cases before they land in your database and become permanently corrupted rows.

Manual quote-escaping in generated JSON

The most reliable way to produce broken JSON is to build it with string concatenation. The pattern looks harmless: '{"name": "' + user.name + '"}'. It works until a value contains a quote, a backslash, or a newline, and then you get a parse error — or worse, a valid JSON object with injected content. Always serialize with a real encoder: JSON.stringify(), Python's json.dumps(), Go's encoding/json. A telling failure mode from real projects: values pasted from Excel with embedded line breaks silently splitting rows when CSV output gets converted to JSON. The CSV to JSON converter handles the quoting rules of RFC 4180 correctly, which hand-rolled parsers routinely miss — especially multiline quoted fields.

Confusing format errors with content errors

When a parser rejects a document, the line number it reports is where syntax broke — not where the actual mistake is. A missing comma at line 4 manifests as an "unexpected token" at line 5. When debugging, start one line above where the parser points. Structured tools beat eyeballing here: a JSON formatter indents and pretty-prints the document, which exposes missing commas, mismatched brackets, and stray BOM characters within seconds. On the command line, python3 -m json.tool or jq empty file.json does the same job in CI.

Linting JSON like code

Formatted JSON costs bandwidth. A 2 MB pretty-printed payload with four-space indentation typically compresses to under 400 KB with minification — and over HTTP with gzip, the difference shrinks, but parse time still drops measurably. Minify before serving or storing; pretty-print only for humans. The mistake is doing it manually or inconsistently across environments — make it a build step.

The validation checklist

Before any JSON payload ships or gets consumed:

  • Syntax check with a strict parser that rejects duplicate keys and trailing commas.
  • Schema check against the expected structure at every trust boundary.
  • Encoding check: UTF-8 without BOM, no unpaired surrogates.
  • Serialization discipline: generated by encoders, never string concatenation.

None of this adds more than a few milliseconds of overhead. It removes an entire class of 3 a.m. incidents. The pattern behind every mistake above is the same: assuming the next layer will catch what this layer let through. Validating at the boundary is how you stop paying for that assumption.