5 JSON validation mistakes that break production apps
A JSON file can be syntactically perfect and still take your service down. Most teams discover this when a config file ships, an API response changes shape, or a third-party webhook sends something your parser never expected. Here are the five mistakes that show up over and over in real production incidents, and how to catch each one before deploy.
1. Checking syntax and calling it validation
"It parses, so the JSON is fine" is the most expensive assumption in config management. JSON.parse() and its equivalents only verify structure: matching braces, valid tokens, no trailing commas. They know nothing about your schema.
Real example: an upstream API renames a field from amount_cents to amount without telling you. Their JSON parses fine. Your code reads undefined * 100 and charges nothing, or NaN propagates until the reconciliation job fails at 3 a.m. Syntax checking would have passed this at every stage.
What actually works:
- Validate against a JSON Schema in CI, not just a parse check.
- Set
additionalProperties: falsewhere possible so renamed or extra fields fail loudly. - Run the schema check at every trust boundary: API ingress, queue consumers, file imports.
When you need a quick structural check on a file pulled from an API or a log dump, a JSON validator points you at the exact line and column of a syntax error instead of making you bisect an 8000-line file by hand. Treat it as step one, not the whole checklist.
2. Trusting numbers
JavaScript cannot represent every decimal. 0.1 + 0.2 is 0.30000000000000004, and JSON numbers follow IEEE 754 double precision. Fine for physics, catastrophic for money.
Take 19.99, multiply by 100 to get minor units: in JS you get 1998.9999999999998, and a naive Math.floor() turns it into 1998 cents. Off by one cent per transaction. At 50,000 transactions a month that is real money drifting in or out of the books, and it is invisible until someone audits.
The robust approach:
- Send money as integers in the smallest unit (cents) or as strings, and parse strings with a decimal library on both sides.
- Never round floating-point values with
Math.roundorfloorin a payment path. - Validate ranges too: a negative
discountis a schema failure, not just a math bug.
3. Assuming field presence and types
JSON has no required-field concept at the syntax level. {"user": null}, {} and {"user": ""} all parse. Your code should not treat them the same.
Three failure modes that hit constantly:
- Null vs missing vs empty string. API A returns
"phone": null; API B omits the key entirely. Code written against A crashes on B. - Numbers that arrive as strings.
"zip": "12345"from one source,"zip": 12345from another. Comparisons and formatting diverge silently. - Nested nulls.
order.customer.address.cityis not a chain, it is a minefield. One null in the middle and the whole expression throws.
Enforce nullability and types per field in your schema, and use optional chaining on every nested path you do not own. If an upstream contract is unclear, write the defensive version first and relax it later, never the reverse.
4. Letting encoding garbage through the door
JSON strings are Unicode, but the bytes around them are not. Three classic bugs:
- A UTF-8 BOM at the start of a file. Java and some Windows tools add it; JSON.parse rejects the whole document. Strip it or use a parser flag.
- Unescaped control characters from logs or pastes, which are invalid inside strings.
- Non-normalized Unicode.
caféstored as e plus combining accent vs precomposed é are different byte sequences, so equality checks, dedupe and API lookups fail. Normalize to NFC on ingest. Running text through a Unicode normalizer before it lands in a payload catches the invisible inconsistencies that later become "why does this key not match" tickets.
5. Skipping limits and depth
A schema that passes validation still may not survive the parser. Deeply nested JSON (a few thousand levels) can blow the recursion stack in many runtimes. Multi-megabyte documents make in-browser parsers choke or take seconds on mobile. Attackers know both: deeply nested payloads are a documented DoS vector against JSON APIs.
Practical guardrails:
- Enforce maximum string length and nesting depth in your schema where the tooling supports it.
- Cap request body size at the reverse proxy (nginx
client_max_body_size), not only in application code. - Never build a document by string concatenation; one unescaped quote and the JSON is invalid. Build objects and serialize with a real library.
The checklist that prevents most incidents
- Parse check in CI on every JSON file and fixture.
- Schema validation with required fields, types, enums and value ranges.
- Money as integers or strings, never floats.
- Explicit handling for null, missing and empty-string cases on third-party payloads.
- NFC normalization for any text from users or external sources.
- Body size and nesting depth limits at the proxy.
There is a pattern across all five mistakes: validation that checks only what the format allows, not what your system requires. The teams that rarely have JSON incidents are not smarter, they just validate one layer further. Add the schema step this week; it takes an afternoon and removes the most common category of 3 a.m. page.