Production Prompt Engineering
Production Prompt Engineering
Section titled “Production Prompt Engineering”The Problem
Section titled “The Problem”You’ve mastered prompt design. Your prompts work in the notebook. Now you need to:
- Serve 10,000 requests per minute
- Maintain 99.9% reliability
- Keep latency under 500ms
- Monitor costs in real-time
- Deploy updates without downtime
- Comply with SOC 2 and GDPR
This is production prompt engineering — where prompt craft meets software engineering.
Why Production Prompt Engineering Exists
Section titled “Why Production Prompt Engineering Exists”Production prompt engineering exists because:
- Scale changes everything — patterns that work for 10 users fail at 10,000
- Reliability is non-negotiable — downtime costs money
- Costs grow linearly — with volume, every token counts
- Monitoring is essential — you can’t fix what you can’t see
- Governance is required — compliance, audit, and security
“A prompt that works in a notebook is a prototype. A prompt that works at scale is an engineering achievement.”
Story: The Startup vs Enterprise
Section titled “Story: The Startup vs Enterprise”Startup: One prompt, one model, one developer. Works great for 100 users.
Enterprise:
- 47 prompts across 12 teams
- 8 different models
- A/B testing in production
- Latency SLO of 200ms
- Monthly token budget of $50K
- Audit requirements from 3 regulators
The same prompt engineering skills apply, but the infrastructure around them is completely different.
Enterprise Prompt Architecture
Section titled “Enterprise Prompt Architecture”Mermaid: Production Prompt Architecture
Section titled “Mermaid: Production Prompt Architecture”flowchart TD subgraph Client Layer A[Web App] B[Mobile App] C[API Clients] end
subgraph Gateway Layer D[Load Balancer] E[API Gateway] F[Rate Limiter] end
subgraph Prompt Layer G[Prompt Registry] H[Template Engine] I[Version Manager] end
subgraph LLM Layer J[Routing] K[Model Pool] L[Fallback] end
subgraph Observability M[Monitoring] N[Logging] O[Cost Tracking] end
A --> D B --> D C --> D D --> E E --> F F --> G G --> H H --> I I --> J J --> K K --> L L --> M M --> N N --> O
style Gateway Layer fill:#3b82f6,color:#fff style Prompt Layer fill:#8b5cf6,color:#fff style LLM Layer fill:#22c55e,color:#000 style Observability fill:#f59e0b,color:#000Production Components
Section titled “Production Components”1. Prompt Registry
Section titled “1. Prompt Registry”Centralized store for all prompts:
prompt_registry: storage: postgresql cache: redis access_patterns: - prompt_id + version - environment (dev/staging/prod)
schema: id: "customer-support-v3" version: "3.2.1" content: "You are a support agent..." model: "claude-3-opus" parameters: temperature: 0.3 max_tokens: 1024 metadata: owner: "support-team" created: "2025-01-15" tags: ["production", "critical"]2. Prompt Router
Section titled “2. Prompt Router”Route requests to the right prompt version:
class PromptRouter: def get_prompt(self, request: Request) -> PromptConfig: """Route to correct prompt based on context."""
# A/B test routing if self.ab_testing_enabled(request.user_id): variant = self.get_ab_variant(request.user_id) return self.registry.get("customer-support", variant)
# Feature flag routing if feature_flags.is_enabled("new_prompt_v2", request.user_id): return self.registry.get("customer-support", "2.0.0")
# Default routing return self.registry.get("customer-support", "production")3. Model Router
Section titled “3. Model Router”Route to the best model for each request:
class ModelRouter: def route_request(self, request: Request) -> str: """Route to appropriate model based on requirements."""
if request.requires_reasoning: return "claude-3-opus" # Best reasoning elif request.requires_speed: return "gpt-4o-mini" # Fastest elif request.is_simple: return "claude-3-haiku" # Cheapest else: return "gpt-4o" # BalancedMermaid: Request Lifecycle
Section titled “Mermaid: Request Lifecycle”sequenceDiagram participant Client participant Gateway participant Router participant Registry participant Model participant Monitor
Client->>Gateway: HTTP Request Gateway->>Gateway: Auth + Rate Limit Gateway->>Router: Route Request
Router->>Registry: Get Prompt Config Registry-->>Router: Prompt + Version
Router->>Router: Build Final Prompt
Router->>Model: LLM Call
par Monitoring Model->>Monitor: Log Latency Model->>Monitor: Log Tokens Model->>Monitor: Log Response end
Model-->>Router: Response Router->>Router: Output Validation Router-->>Gateway: Formatted Response Gateway-->>Client: HTTP Response 200
Monitor->>Monitor: Record MetricsMonitoring & Observability
Section titled “Monitoring & Observability”Key Metrics
Section titled “Key Metrics”| Category | Metrics | Alert Threshold |
|---|---|---|
| Latency | p50, p95, p99 | p95 > 2s |
| Throughput | RPM, RPS | Drop > 20% |
| Errors | 4xx, 5xx, timeout | Rate > 1% |
| Cost | Cost/request, daily spend | Budget > 80% |
| Quality | Accuracy, relevance | Score < 0.8 |
| Model | Token usage, context | > 90% of limit |
Logging Structure
Section titled “Logging Structure”{ "timestamp": "2025-06-15T10:30:00Z", "request_id": "req_abc123", "prompt_id": "customer-support-v3", "version": "3.2.1", "model": "claude-3-opus", "tokens": { "input": 450, "output": 120, "total": 570 }, "latency_ms": 340, "success": true, "user_id": "user_xyz", "environment": "production"}Cost Optimization
Section titled “Cost Optimization”Cost Breakdown
Section titled “Cost Breakdown”cost_model: token_costs: claude-3-opus: input: $15/M tokens output: $75/M tokens gpt-4o: input: $5/M tokens output: $15/M tokens claude-3-haiku: input: $0.25/M tokens output: $1.25/M tokens
daily_estimate: requests: 100,000 avg_input_tokens: 400 avg_output_tokens: 100 daily_cost: ~$200-800 monthly_cost: ~$6,000-24,000Optimization Strategies
Section titled “Optimization Strategies”| Strategy | Savings | Impact |
|---|---|---|
| Model tiering | 40-60% | Match complexity to model |
| Prompt compression | 20-30% | Shorter context = lower cost |
| Caching | 30-50% | Cache identical requests |
| Batching | 15-25% | Combine related requests |
| Response streaming | 0% | Better UX, same cost |
Caching Strategy
Section titled “Caching Strategy”class PromptCache: def __init__(self): self.cache = RedisCache(ttl=3600)
def get_or_compute(self, request, compute_fn): """Return cached response or compute new one."""
cache_key = self._build_key(request)
# Check cache cached = self.cache.get(cache_key) if cached: metrics.record("cache_hit") return cached
# Compute and cache response = compute_fn(request) self.cache.set(cache_key, response) metrics.record("cache_miss")
return responseScaling Strategies
Section titled “Scaling Strategies”Mermaid: Scaling Architecture
Section titled “Mermaid: Scaling Architecture”flowchart TD subgraph Load A[Traffic Spike] end
subgraph Auto Scaling B[Scale Up] C[Scale Out] end
subgraph Strategies D[Request Queueing] E[Rate Limiting] F[Load Shedding] G[Model Pooling] end
subgraph Fallbacks H[Cache Serve] I[Fallback Model] J[Graceful Degradation] end
A --> B A --> C
B --> D C --> E
D --> F E --> F F --> G
G --> H G --> I G --> J
style A fill:#ef4444,color:#fff style Strategies fill:#3b82f6,color:#fff style Fallbacks fill:#22c55e,color:#000Horizontal Scaling
Section titled “Horizontal Scaling”Requests → [Load Balancer] → [App Instance 1] → [App Instance 2] → [LLM API Pool] → [App Instance N]Fallback Chain
Section titled “Fallback Chain”fallback_chain: primary: "claude-3-opus" fallback_1: "gpt-4o" # Different provider fallback_2: "claude-3-sonnet" # Cheaper, faster fallback_3: "cache-only" # Degraded mode ultimate: "static_response" # Pre-written responseGovernance & Compliance
Section titled “Governance & Compliance”Required Controls
Section titled “Required Controls”| Requirement | Implementation |
|---|---|
| Audit Trail | Log all prompt versions + inferences |
| Approval Workflow | PR-based prompt updates |
| PII Protection | Automatic PII redaction |
| Retention Policy | Auto-delete logs after 90 days |
| Access Control | Role-based prompt access |
| Change Management | Versioned, reviewed, tested |
Approval Workflow
Section titled “Approval Workflow”flowchart LR A[Prompt Draft] --> B[Peer Review] B --> C{Approved?} C -->|Yes| D[Staging] C -->|No| A D --> E[QA Testing] E --> F{Passed?} F -->|Yes| G[Production Approval] F -->|No| A G --> H[Canary Deploy] H --> I{Monitor} I -->|OK| J[Full Rollout] I -->|Issue| K[Rollback]
style A fill:#3b82f6,color:#fff style J fill:#22c55e,color:#000 style K fill:#ef4444,color:#fffBad vs Good: Production
Section titled “Bad vs Good: Production”| Bad Practice | Good Practice |
|---|---|
| Prompts in code | Prompts in registry |
| Manual deploy | CI/CD pipeline |
| No monitoring | Full observability |
| Single model | Model pool + fallback |
| No caching | Multi-level caching |
| Deploy to all at once | Canary + staged rollout |
| No cost tracking | Real-time cost dashboard |
Production Checklist
Section titled “Production Checklist”production_readiness: reliability: - [ ] Load testing completed - [ ] Fallback chain configured - [ ] Rate limiting enabled - [ ] Timeout handling implemented
observability: - [ ] Latency monitoring - [ ] Error tracking - [ ] Cost dashboard - [ ] Quality metrics
security: - [ ] Input sanitization - [ ] Output validation - [ ] PII scanning - [ ] Rate limiting
operations: - [ ] Deployment pipeline - [ ] Rollback procedure - [ ] Runbook created - [ ] On-call rotation
compliance: - [ ] Audit logging - [ ] Data retention - [ ] Access control - [ ] Approval workflowCommon Mistakes
Section titled “Common Mistakes”| Mistake | Why It Hurts | Fix |
|---|---|---|
| No fallback strategy | Complete outage | Configure fallback chain |
| Missing cost monitoring | Budget shock | Real-time cost dashboard |
| Single model dependency | Vendor lock-in | Multi-model pool |
| No load testing | Fail under traffic | Regular load tests |
| Manual deployments | Human error | Automate pipeline |
| Ignoring latency | Poor UX | Set latency SLOs |
Best Practices
Section titled “Best Practices”| Practice | Description |
|---|---|
| Design for failure | Every component should have a fallback |
| Monitor everything | If it moves, measure it |
| Automate deploys | No manual prompt changes in production |
| Cost is a feature | Track and optimize aggressively |
| Test at scale | Load test with production traffic patterns |
| Document runbooks | Incident response procedures |
| Gradual rollouts | Canary → 10% → 50% → 100% |
Interview Questions
Section titled “Interview Questions”Beginner
Section titled “Beginner”- What changes when you move prompt engineering from development to production?
- Name three metrics you should monitor in production.
Intermediate
Section titled “Intermediate”- Design a fallback strategy for LLM API failures.
- How would you reduce LLM costs in production without sacrificing quality?
Senior
Section titled “Senior”- Design a production prompt architecture handling 10K requests per minute.
- How would you implement gradual rollouts and A/B testing for prompts?
Staff Engineer
Section titled “Staff Engineer”- Design a multi-region, multi-model prompt infrastructure with 99.99% availability.
- How do you balance cost, latency, and quality across different prompt use cases in a large organization?
Summary
Section titled “Summary”- Scale requires infrastructure — prompts need registries, routers, caches
- Monitor everything — latency, cost, errors, quality
- Design for failure — fallbacks, caching, graceful degradation
- Automate deploys — no manual changes in production
- Track costs — optimize aggressively as volume grows
- Governance is essential — audit, compliance, access control
Key Insight: Production prompt engineering is 20% prompt design and 80% software engineering around the prompt.
Next: Document 25 — Phase Summary