Skip to content

02. Build a Perplexity Clone

Build a production-grade AI search engine like Perplexity — combining real-time web search, RAG, source ranking, citations, and conversational follow-ups in a single interface.

Perplexity revolutionized search by combining LLM reasoning with real-time web data. This project teaches you to build every core feature: web crawling, content extraction, relevance ranking, citation generation, and conversational search.


Traditional search engines return links. Users want answers. An AI search engine should:

  • Understand the user’s question in natural language
  • Search the web in real-time
  • Extract relevant content from multiple sources
  • Generate a comprehensive answer with citations
  • Support follow-up questions that maintain context

A media company building a research assistant for journalists needs a Perplexity-like tool that:

  • Searches thousands of sources in real-time
  • Extracts and summarizes relevant content
  • Provides verifiable citations for every claim
  • Handles 100K+ queries/day
  • Supports multiple languages

#FeatureDescription
FR1Real-time web searchQuery multiple search engines (Bing, Google, SerpAPI)
FR2Content extractionParse and clean web pages
FR3Relevance rankingScore and rank search results
FR4Answer generationLLM generates answer from sources
FR5Citation generationLink every claim to its source
FR6Follow-up questionsContext-aware conversation
FR7Source managementAdd/remove/prioritize sources
FR8Search historyPersist search queries and results
FR9CollectionsOrganize searches into topics
FR10Pro searchDeep search with more sources
#RequirementTarget
NFR1Search latency< 3s end-to-end
NFR2Answer quality≥ 90% factually accurate
NFR3Availability99.9%
NFR4Citation accuracyEvery claim linked to real source
NFR5Cost efficiency≤ $0.05 per search query

LayerTechnologyPurpose
FrontendNext.js + Tailwind CSSSearch UI, streaming answers
BackendFastAPI (Python)Query processing, search orchestration
DatabasePostgreSQLUser data, search history
Vector DBQdrantDocument embeddings for relevance
CacheRedisSearch result caching
Search APISerpAPI / Bing SearchReal-time web search
AIOpenAI GPT-4o / ClaudeAnswer generation, summarization
QueueCelery + RedisAsync web crawling
File StorageS3Cached page snapshots
DeploymentDocker + K8s + AWSProduction infrastructure

flowchart TD
subgraph FRONTEND["Frontend"]
UI["Search UI\nNext.js"]
SSE["SSE Client\nStreaming"]
end
subgraph ORCH["Orchestration"]
GW["API Gateway"]
SEARCH["Search Service\nQuery + Crawl"]
RANK["Ranking Service\nRelevance scoring"]
EXTRACT["Content Extraction\nPage parsing"]
end
subgraph AI["AI Layer"]
SUMMARIZE["Summarization\nAnswer generation"]
CITE["Citation Engine\nSource linking"]
FOLLOWUP["Follow-up\nContext management"]
end
subgraph DATA["Data Layer"]
PG["PostgreSQL"]
QDRANT["Qdrant\nEmbeddings"]
REDIS["Redis\nCache"]
S3["Page Cache"]
end
subgraph EXTERNAL["External"]
BING["Bing API"]
GOOGLE["Google Search"]
SERP["SerpAPI"]
end
UI --> GW
GW --> SEARCH
SEARCH --> BING
SEARCH --> GOOGLE
SEARCH --> SERP
SEARCH --> EXTRACT
EXTRACT --> RANK
RANK --> QDRANT
RANK --> SUMMARIZE
SUMMARIZE --> CITE
CITE --> UI
SUMMARIZE --> REDIS
SEARCH --> REDIS
EXTRACT --> S3
style FRONTEND fill:#3b82f6,color:#fff
style ORCH fill:#8b5cf6,color:#fff
style AI fill:#22c55e,color:#fff
style DATA fill:#f59e0b,color:#fff
style EXTERNAL fill:#ef4444,color:#fff

sequenceDiagram
participant U as User
participant S as Search Service
participant Web as Web Search
participant Crawl as Crawler
participant Rank as Ranking
participant LLM as LLM
participant Cache as Cache
U->>S: "What is the latest AI news?"
S->>Cache: Check cache
Cache-->>S: Miss
S->>Web: Search multiple engines
Web-->>S: Raw results (URLs + snippets)
S->>Crawl: Fetch top 10 pages
Crawl->>Crawl: Extract content, strip HTML
Crawl-->>S: Cleaned content
S->>Rank: Score relevance
Rank->>Rank: Embed query + docs, cosine similarity
Rank-->>S: Ranked passages
S->>LLM: Generate answer from top passages
LLM->>LLM: Answer + citations
LLM-->>S: Generated answer
S->>Cache: Store result (TTL: 1 hour)
S-->>U: Answer with citations
U->>S: "What about advancements in robotics?"
S->>S: Use previous context
S->>LLM: Follow-up with context
LLM-->>S: Context-aware answer

flowchart LR
QUERY["User Query"] --> MULTI["Multi-Engine Search"]
MULTI --> BING["Bing Search\nAPI"]
MULTI --> GOOGLE["Google Custom\nSearch"]
MULTI --> SERP["SerpAPI\nGoogle results"]
BING --> MERGE["Merge + Deduplicate"]
GOOGLE --> MERGE
SERP --> MERGE
MERGE --> RANK["Rank by\nrelevance score"]
RANK --> TOP_K["Top-K results\n(5-10 pages)"]
style MULTI fill:#3b82f6,color:#fff
style MERGE fill:#f59e0b,color:#fff
style TOP_K fill:#22c55e,color:#fff
flowchart TD
URL["Web Page URL"] --> FETCH["HTTP Fetch\nUser-agent headers"]
FETCH --> PARSE["HTML Parse\nBeautifulSoup/Readability"]
PARSE --> CLEAN["Clean Content\nRemove ads, nav, scripts"]
CLEAN --> CHUNK["Chunk Content\n512 token segments"]
CHUNK --> EMBED["Generate Embeddings\nFor each chunk"]
EMBED --> INDEX["Store in Vector DB\nWith source metadata"]
style FETCH fill:#3b82f6,color:#fff
style CLEAN fill:#22c55e,color:#fff
style INDEX fill:#f59e0b,color:#fff
flowchart LR
PASSAGES["Retrieved Passages\n10-20 chunks"] --> CROSS["Cross-Encoder\nRe-ranking model"]
QUERY["User Query"] --> CROSS
CROSS --> SCORES["Relevance Scores\n0.0 - 1.0"]
SCORES --> TOP["Top 3-5 Passages\nFor answer generation"]
style CROSS fill:#8b5cf6,color:#fff
style TOP fill:#22c55e,color:#fff
flowchart TD
ANSWER["Generated Answer"] --> CLAIMS["Extract Claims\nSentence splitting"]
CLAIMS --> MATCH["Match to Sources\nSemantic similarity"]
MATCH --> VERIFY["Verify Claim\nIs source relevant?"]
VERIFY -->|"Yes"| CITE["Add Citation\n[1], [2], etc."]
VERIFY -->|"No"| DROP["Drop unsupported claim"]
CITE --> FINAL["Final Answer\nWith numbered sources"]
style CLAIMS fill:#3b82f6,color:#fff
style CITE fill:#22c55e,color:#fff
style DROP fill:#ef4444,color:#fff

MethodEndpointPurpose
POST/api/searchExecute a search query
GET/api/search/{id}Get search results
POST/api/search/{id}/followupFollow-up question
GET/api/historyGet search history
DELETE/api/history/{id}Delete search
POST/api/collectionsCreate collection
GET/api/collections/{id}Get collection searches
GET/api/sourcesAvailable search sources

erDiagram
USERS ||--o{ SEARCHES : creates
SEARCHES ||--o{ SEARCH_RESULTS : contains
SEARCHES ||--o{ FOLLOW_UPS : has
SEARCH_RESULTS ||--o{ SOURCES : cites
COLLECTIONS ||--o{ SEARCHES : groups
USERS {
uuid id PK
string email
string name
timestamp created_at
}
SEARCHES {
uuid id PK
uuid user_id FK
string query
text answer
json citations
int sources_count
timestamp created_at
}
SEARCH_RESULTS {
uuid id PK
uuid search_id FK
string url
string title
string snippet
float relevance_score
}
SOURCES {
uuid id PK
string url
string domain
string title
text cached_content
timestamp crawled_at
}
FOLLOW_UPS {
uuid id PK
uuid search_id FK
string query
text answer
int turn_number
}
COLLECTIONS {
uuid id PK
uuid user_id FK
string name
string description
timestamp created_at
}

flowchart TD
subgraph PROD["Production"]
CF["CloudFront\nCDN"]
ALB["Load Balancer"]
subgraph ECS["ECS Fargate"]
FE["Frontend\nNext.js"]
API["API\nFastAPI"]
WORKER["Crawler Workers\nCelery"]
RANK_SVC["Ranking\nService"]
end
subgraph DATA["Data"]
RDS["Aurora\nPostgreSQL"]
ELASTICACHE["Redis\nCache + Queue"]
S3["Page Cache"]
QDRANT_CLOUD["Qdrant\nVector DB"]
end
end
CF --> ALB
ALB --> FE
ALB --> API
API --> WORKER
API --> RANK_SVC
API --> RDS
API --> ELASTICACHE
WORKER --> S3
API --> QDRANT_CLOUD
style PROD fill:#1e293b,color:#fff
style ECS fill:#3b82f6,color:#fff
style DATA fill:#f59e0b,color:#fff

MetricMethodTarget
Search latencyP95 end-to-end< 3s
Citation accuracyHuman review sample> 95%
Answer relevanceLLM-as-a-Judge> 90%
Cache hit rate% of queries served from cache> 30%
Crawl success rate% of URLs successfully crawled> 95%
User satisfactionThumbs up/down> 85%

ConcernImplementation
Rate limitingPer-user: 10 searches/min, Per-IP: 100/min
Content filteringBlock malicious/pornographic sites from results
Data privacyNo storage of raw web content beyond TTL
API securityAll endpoints authenticated via JWT
Crawl ethicsRespect robots.txt, rate-limit crawling

FeaturePriorityComplexity
Image search integrationMediumMedium
PDF/document searchHighMedium
Custom knowledge basesHighHigh
Multi-language supportMediumMedium
Real-time news alertsLowHigh
Team workspacesMediumHigh

Q: Design the search pipeline for a Perplexity clone handling 100 queries/second.

Multi-stage pipeline: (1) Query understanding — Classify query type (factual, opinion, news), (2) Parallel search — Hit Bing, Google, and internal index simultaneously, (3) Content extraction — Async worker pool crawls top URLs, (4) Re-ranking — Cross-encoder scores passages, (5) Answer generation — GPT-4o synthesizes answer from top passages, (6) Caching — Store results with TTL based on query type (news: 5min, facts: 1hr). Use Redis for hot cache, S3 for warm cache.

Q: How do you ensure citation accuracy?

(1) Claim extraction — Parse answer into atomic claims, (2) Source matching — Embed each claim and find best matching source passage, (3) Verification — Check if source actually supports the claim (NLI model), (4) Confidence scoring — Only include citations above 0.85 threshold, (5) Human review — Sample 1% of answers for citation accuracy audit.

Q: Design the crawling infrastructure for a Perplexity clone.

Architecture: (1) Task queue — Celery with Redis broker, (2) Worker pool — 50-100 workers, respecting per-domain rate limits, (3) Content extraction — Readability algorithm + BeautifulSoup, (4) Cache — Store cleaned content in S3 with 7-day TTL, (5) Politeness — Track per-domain request timing, delay between requests, (6) robots.txt — Cache and respect per-domain rules.


FeatureImplementation
Web searchMulti-engine (Bing, Google, SerpAPI)
Content extractionReadability + HTML parsing
Re-rankingCross-encoder model
Answer generationGPT-4o with source-grounded responses
CitationsClaim extraction + source matching
CachingRedis (hot) + S3 (warm)
Async crawlingCelery worker pool
Follow-upsContext-aware conversation management

Previous: 01 — Build a ChatGPT Clone

Next: 03 — Build a NotebookLM Clone

Related Projects: