Valid AI JSON Can Still Be Wrong
Validate AI output in three stages: JSON syntax, schema and real meaning, with explicit handling for missing facts and failed validation.
The essentials
- Valid JSON only means a parser can read the text.
- It does not guarantee the required fields exist, their types are correct or their values are supported by the input.
On this page
Valid JSON only means a parser can read the text. It does not guarantee the required fields exist, their types are correct or their values are supported by the input. Validate syntax, schema and meaning separately.
Define the output contract
Suppose your application extracts an invoice date and total. Decide whether an absent date should be null, omitted or an error. Define the currency separately from the numeric amount. Document whether a total includes tax.
Ollama's structured-output documentation supports supplying a schema and validating the result. Check feature support for the exact endpoint and model you use rather than assuming every compatible API behaves identically.
A schema should describe the result you actually accept, not merely make a demo response parse.
Use three validation layers
| Layer | Example failure | Response |
|---|---|---|
| Syntax | Truncated object | Reject incomplete response |
| Schema | Amount returned as an array | Report contract violation |
| Meaning | Amount from the wrong invoice | Compare with source evidence |
A response such as {"total": 100, "currency": "USD"} can pass the first two layers and still be false. Keep source identifiers or evidence spans where the task needs auditability.
The upstream schema error report also shows why generated schemas can expose version-specific compatibility issues. Preserve the actual schema in bug reports.
Test absence, ambiguity and truncation
Build a fixture with a missing field, another with two possible values, and a third that exceeds your normal response budget. The application should not substitute a plausible value when evidence is absent.
Treat incomplete streamed output as incomplete until the final result arrives. Do not execute an action using a partially parsed object simply because the opening fields look valid.
If you retry a failed extraction, bound the number of attempts and preserve the failure reason. Endless repair prompts can consume resources without resolving an ambiguous source.
Keep business rules outside generation
A model can propose an extracted value. Application code should decide whether the value is allowed: a date is within the accepted range, a referenced account exists, or a total reconciles with line items.
Do not let schema compliance authorize a payment, deletion or external message. Those actions need their own permission and consistency checks.
Establish a regression set
Keep the schema, fixture input, expected outcome and validation verdict together. Repeat it when changing model versions, schema libraries or endpoint adapters.
Use prompt regression testing for that record, and AI claim checking when the extracted fields depend on interpreting source text. The objective is an accurate usable object, not simply a green JSON parser.
Related troubleshooting
This guide draws on the linked documentation. Examples are illustrative unless explicitly identified as measured results.
Practical guides published by Lucivo, developed with AI assistance and references to official documentation. Examples are illustrative unless a guide explicitly documents a hands-on test. Check the linked sources for current product details.
Related articles
Stop Repeated AI Agent Tool Calls
AI Transcription Invents Words in Silence
Change Embedding Models Without Mixing Vectors
The Weekly Breakdown
High signal AI & software stories.
Direct to your inbox. No hype.
Independent analysis of AI models, developer tools, and computing architectures. Delivered every Sunday morning. 100% free.