Skip to content

23. Structured Output

Structured output is the practice of constraining LLM generation to follow a specific format — JSON, XML, SQL, or code — ensuring the model’s output can be reliably parsed and used by applications.

LLMs generate text naturally. But applications need structured data. The model might write a beautiful paragraph when you need a clean JSON object. Structured output bridges this gap.

flowchart TD
LLM["🧠 LLM Output\n(raw text)"] --> PROMPT["Method 1: Prompt\n'Respond in JSON'"]
LLM --> SCHEMA["Method 2: JSON Schema\nDefine exact structure"]
LLM --> CONSTRAIN["Method 3: Constrained\nDecoding\n(token-level)"]
LLM --> FORMAT["Method 4: Post-Process\nParse + fix errors"]

ChallengeExample
Invalid JSONMissing comma, trailing comma, extra text
Wrong schemaExtra fields, missing fields, wrong types
Inconsistent formatDate is “Jan 15” in one call, “2024-01-15” in another

response = client.chat.completions.create(
model="gpt-4o",
messages=[{
"role": "system",
"content": "Respond with valid JSON only. No explanation."
}, {
"role": "user", "content": "Extract the name and age."
}],
response_format={"type": "json_object"}
)
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Extract person details."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "person",
"schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer", "minimum": 0},
},
"required": ["name", "age"]
}
}
}
)

Token-level constraints guarantee valid output by only allowing valid tokens at each step:

import guidance
program = guidance("""
Extract: {{input}}
{
"name": "{{gen 'name'}}",
"age": {{gen 'age' pattern='\\\d+'}}
}
""")
result = program(input=text)

  1. Use JSON Schema when available — Strongest format guarantees.
  2. Always validate server-side — Never assume the model followed instructions.
  3. Include examples — Few-shot examples improve compliance significantly.
  4. Handle failures — Have fallback logic when output doesn’t match the schema.

MethodReliabilityComplexitySpeed
Prompt engineeringLowNoneFastest
JSON SchemaHighLowFast
Constrained decoding100%HighSlowest
Post-processingMediumMediumFast

Previous: 22 — Function Calling

Next: 24 — Hallucinations

Related Topics: