Skip to content
TokIQ
Output

How to Get Reliable JSON From LLMs: Schemas, Validation, Retries

Getting valid JSON from an LLM takes more than asking nicely. Use schemas, structured output modes, validation and retries, and avoid the usual failure modes.

TokIQ Editorial5 min read
In this article
  1. Why do models break JSON in the first place?
  2. Step 1: Describe the output precisely in the prompt
  3. Step 2: Use the API's structured output feature
  4. Step 3: Validate everything, even with enforcement
  5. Step 4: Retry with the error, not blindly
  6. Where should reasoning go in a JSON response?
  7. Common design mistakes in output schemas
  8. Is prompting alone ever enough?

To get reliable JSON from a language model, define the structure as a JSON Schema, use the provider's structured output feature where available, validate every response in code, and retry with the validation error when it fails. Prompt wording alone gets you most of the way; the last few percent of reliability comes from the API and your own validation.

The distinction matters because "most of the way" is not good enough for code. A parser that fails on 2% of responses means a broken feature for someone every few minutes at scale.

Why do models break JSON in the first place?

Before fixing it, it helps to know what actually goes wrong. In roughly the order you will meet them:

  1. Markdown fences. The model returns ```json ... ``` because chat models are trained to present code nicely.
  2. Chatty preambles. "Sure! Here is the JSON you requested:" followed by the object.
  3. Wrong or invented keys. customerName instead of customer_name, or an extra notes field you never asked for.
  4. Wrong types. "price": "19.99" as a string, "count": "three", a single object where you wanted an array.
  5. Truncation. The response hits the max token limit and stops mid-object.
  6. Syntactically valid nonsense. Perfect JSON, wrong content: a date in the wrong format, a hallucinated phone number, null where the value was right there in the input.

Structured output features eliminate the first five, mostly. Nothing eliminates the sixth except good prompting, grounding and checks.

Step 1: Describe the output precisely in the prompt

Even with schema enforcement, the prompt is where you explain what each field means. A schema says due_date is a string; only the prompt can say it should be ISO 8601 and that relative dates should be resolved against the email's send date.

A weak prompt:

Extract the invoice details from this email as JSON.

A stronger one:

Extract invoice details from the email below.

Return a single JSON object with exactly these keys:
- "vendor": string, the company that issued the invoice
- "invoice_number": string, as written (keep leading zeros)
- "total": number, in the invoice currency, no currency symbol
- "currency": string, ISO 4217 code such as "USD" or "EUR"
- "due_date": string in YYYY-MM-DD format, or null if not stated

If a value is not in the email, use null. Do not guess.
Return only the JSON object, with no markdown and no explanation.

<email>
{{email_text}}
</email>

Two details carry a lot of weight here. "Keep leading zeros" prevents the model from "helpfully" turning 00417 into 417. And "use null, do not guess" gives the model a legitimate way out. Without an explicit null option, a model under a required-field constraint will invent a plausible value rather than leave it empty.

Step 2: Use the API's structured output feature

Most major providers now offer some way to constrain output to a schema. The names and details differ and change over time, so check current docs, but the main options look like this:

  • OpenAI has a basic JSON mode that guarantees syntactically valid JSON, and Structured Outputs, where you pass a JSON Schema with strict mode and the output is constrained to match it.
  • Google Gemini lets you set the response MIME type to JSON and supply a response schema.
  • Anthropic lets you define a tool with an input schema and have Claude call it, which yields schema-shaped arguments; its docs also describe native structured output support.

The important distinction is JSON mode vs schema enforcement. JSON mode only promises that the text parses. Your keys can still be wrong. Schema enforcement (often done with constrained decoding, which restricts which tokens the model is allowed to generate next) promises the shape.

A schema for the invoice example:

{
  "type": "object",
  "properties": {
    "vendor": { "type": ["string", "null"] },
    "invoice_number": { "type": ["string", "null"] },
    "total": { "type": ["number", "null"] },
    "currency": { "type": ["string", "null"], "enum": ["USD", "EUR", "GBP", null] },
    "due_date": { "type": ["string", "null"], "description": "YYYY-MM-DD" }
  },
  "required": ["vendor", "invoice_number", "total", "currency", "due_date"],
  "additionalProperties": false
}

Note the nullable types. In strict modes, every field is typically required, so "optional" has to be expressed as "required but may be null". Also note additionalProperties: false, which stops surprise fields. Strict modes usually support only a subset of JSON Schema, so check which keywords are allowed before building something elaborate.

If you use Python or TypeScript, define the schema as a Pydantic model or a Zod schema and generate the JSON Schema from it. Then the same definition validates the response, and the two cannot drift apart.

Step 3: Validate everything, even with enforcement

Schema enforcement guarantees structure. It does not guarantee:

  • that due_date is a real date (February 30 is a valid string),
  • that total matches the sum of line items,
  • that vendor appears anywhere in the source email,
  • that the model did not drop the third line item of five.

Write the checks that matter for your domain. Parse dates. Check that totals are non-negative. For extraction tasks, check that extracted strings actually occur in the input; this single test catches a surprising amount of hallucination.

Step 4: Retry with the error, not blindly

When validation fails, do not just resend the same request and hope. Send the error back:

Your previous response failed validation:
- "due_date": "March 3rd" does not match format YYYY-MM-DD
- "currency": "dollars" is not one of USD, EUR, GBP

Return the corrected JSON object only.

Models fix specific, named errors well. Cap retries at one or two, log every failure, and look at the logs weekly. Repeated failures on the same field usually mean the prompt or schema is unclear, not that the model is flaky. Libraries such as Instructor (Python) wrap this validate-and-retry loop for you, but the pattern is simple enough to write yourself.

For syntax-only breakage on providers without enforcement, a lenient parser that strips fences and trailing text before JSON.parse handles the most common cases. Treat it as a fallback, not a strategy.

Where should reasoning go in a JSON response?

If the task needs judgment (a classification with a justification, a risk score), put a reasoning field before the decision field:

{
  "evidence": "Customer mentions a duplicate charge on the March invoice.",
  "category": "billing",
  "confidence": "high"
}

Models generate fields in order. A decision written first gets justified afterward instead of informed by the reasoning. We go deeper on this in chain-of-thought prompting explained.

The flip side: reasoning fields cost tokens on every call. For simple extraction, leave them out.

Common design mistakes in output schemas

Deep nesting for no reason. Three levels of objects where a flat list would do. Every level is another place for the model to misplace a value.

Free-text fields that should be enums. If downstream code branches on status, make it an enum. Otherwise you will get "Pending", "pending", "In progress" and "awaiting review" for the same thing.

Ambiguous field names. date on an invoice could be the issue date, the due date or the payment date. Name it due_date and describe it.

Asking for too much in one call. A schema with forty fields extracted from a twenty-page contract invites omissions. Split it: one call per section, or one call to locate the relevant passages and another to extract from them.

Forgetting the empty case. What should the output be if the email is not an invoice at all? Add an is_invoice boolean or allow an empty result, or the model will extract an invoice from a lunch invitation.

Is prompting alone ever enough?

For prototypes, scripts you run by hand, or models without structured output support, yes. A precise format description, an example of the exact output, and a defensive parser get you to high reliability. For anything user-facing, use enforcement where available and validation always.

The underlying skill is the same in every case: describe the output so precisely that there is only one reasonable way to fill it in. The structured output topic drills those decisions, and if the JSON you produce is a tool call rather than a final answer, tool calling and AI agents picks up from here. For extraction grounded in source documents, see what RAG is.

Frequently asked questions

How do I force an LLM to return valid JSON?

Use the provider's structured output feature with a JSON Schema when it exists, because it constrains generation to the schema. Then still validate the result in code, and retry or repair when validation fails.

What is the difference between JSON mode and structured outputs?

JSON mode makes the model produce syntactically valid JSON but does not guarantee your keys or types. Structured outputs with a strict schema constrain the output to match a schema you supply.

Why does my LLM wrap JSON in markdown code fences?

Chat-tuned models are trained to format code nicely for humans. Tell the model to return only the raw JSON object, or use a structured output mode, and strip fences defensively in your parser.

Should I validate JSON from an LLM even with a schema?

Yes. A schema guarantees shape, not truth. The JSON can be perfectly valid and still contain a wrong date, an invented value or an empty field where data existed.

  • #structured output
  • #JSON
  • #JSON schema
  • #validation

Now practice it

TokIQ turns prompt engineering into short quizzes, with an explanation for every answer.

Coming soon onApp StoreComing soon onGoogle Play

Write better prompts, a few questions a day.

Short quizzes on real prompting decisions, with an explanation for every answer. Free to start on iPhone and Android.

Coming soon onApp StoreComing soon onGoogle Play