Skip to content

13 — File Storage Architecture

File storage architecture handles how files (images, videos, documents) are uploaded, stored, processed, and served. A well-designed storage system is scalable, cost-effective, and fast.

Analogy: File storage is like a library. Object storage (S3) is the main bookshelves where books live. Block storage (EBS) is a book on your desk that you’re actively editing. File storage (EFS) is a shared reading table where multiple people can access the same book.


Poor file storage design leads to:

  • Storage running out — no capacity planning for growing data
  • Slow uploads/downloads — single server bottleneck
  • Data loss — no redundancy or backup strategy
  • High costs — storing everything in expensive hot storage
  • CDN misses — origin server overloaded serving files

flowchart TB
Storage["File Storage Types"] --> Object["Object Storage<br/>S3, GCS, Azure Blob<br/>Unlimited scale<br/>HTTP API access"]
Storage --> Block["Block Storage<br/>EBS, SSD, HDD<br/>Fast, low-latency<br/>Attached to servers"]
Storage --> File["File Storage<br/>EFS, NFS, NAS<br/>Shared access<br/>POSIX compatible"]
Object --> ObjectUse["Use: Images, videos, backups, static assets"]
Block --> BlockUse["Use: Database storage, OS volumes, temp files"]
File --> FileUse["Use: Shared configs, home directories, logs"]
style Storage fill:#7c3aed,color:#fff
style Object fill:#3b82f6,color:#fff
style Block fill:#f59e0b,color:#fff
style File fill:#059669,color:#fff
FeatureObject Storage (S3)Block Storage (EBS)File Storage (EFS)
APIHTTP (REST)Block device (mounted)File system (NFS)
ScaleVirtually unlimitedLimited by volume sizeUp to petabytes
SpeedModerate (ms)Very fast (μs)Fast (ms)
AccessFrom anywhereSingle serverMultiple servers
CostLow per GBHigherMedium
Use caseStatic assets, backupsDB storage, app dataShared files, logs

flowchart TB
subgraph S3["Amazon S3 (Object Storage)"]
Bucket["📦 Bucket<br/>Globally unique name"]
Objects["📄 Objects (files)<br/>Key: path/to/file.jpg<br/>Value: binary data<br/>Metadata: content-type, size"]
Bucket --> Objects
end
subgraph Access["Access Patterns"]
HTTP["HTTP REST API<br/>PUT / GET / DELETE"]
SDK["AWS SDK / CLI<br/>s3 cp, s3 sync"]
CDN["CloudFront CDN<br/>Global distribution"]
end
Objects --> HTTP & SDK & CDN
subgraph Security["Security"]
BP["Bucket Policies"]
IAM["IAM Permissions"]
Enc["Encryption (SSE)"]
end
style S3 fill:#7c3aed,color:#fff
style Access fill:#3b82f6,color:#fff
style Security fill:#059669,color:#fff

sequenceDiagram
participant User as User Browser
participant App as Application
participant S3 as S3 / Object Store
participant CDN as CDN
participant Worker as Processing Worker
User->>App: Request upload URL
App->>App: Generate presigned URL
App-->>User: Presigned PUT URL (expires in 15 min)
User->>S3: PUT directly to S3 (via presigned URL)
S3-->>User: Upload complete ✅
Note over S3: S3 event notification triggers
S3->>Worker: File uploaded event
Worker->>S3: Download file
Worker->>Worker: Process (resize, compress, thumbnail)
Worker->>S3: Save processed versions
User->>App: Request file
App-->>User: CDN URL
User->>CDN: GET /images/photo.jpg
CDN-->>User: Cached file (or fetch from S3)

Data moves through storage tiers as it ages — hot → warm → cold → archive:

flowchart LR
Upload["📤 Upload"] --> Hot["Hot Storage<br/>S3 Standard<br/>💰 $$$<br/>Instant access"]
Hot -->|"30-90 days"| Warm["Warm Storage<br/>S3 Standard-IA<br/>💰 $$<br/>Instant access"]
Warm -->|"90-365 days"| Cold["Cold Storage<br/>S3 Glacier<br/>💰 $<br/>Min. retrieval"]
Cold -->|"> 365 days"| Archive["Archive<br/>Glacier Deep Archive<br/>💰 ¢<br/>12hr retrieval"]
style Hot fill:#3b82f6,color:#fff
style Warm fill:#059669,color:#fff
style Cold fill:#f59e0b,color:#fff
style Archive fill:#ef4444,color:#fff
TierS3 ClassCost/GB/moRetrievalRetention
HotStandard$0.023Instant0-30 days
WarmStandard-IA$0.0125Instant30-90 days
ColdGlacier$0.0041-5 min90-365 days
ArchiveDeep Archive$0.00112 hours1+ years

flowchart TB
subgraph Upload["Upload Path"]
User["User Device"] -->|Upload| UploadAPI["Upload API"]
UploadAPI -->|Store| S3["S3 Bucket<br/>Original files"]
end
subgraph Process["Processing Path"]
S3 -->|Event| Worker["Processing Worker<br/>Resize, compress, thumbnail"]
Worker -->|Store| Processed["S3 Bucket<br/>Processed files"]
end
subgraph Serve["Serving Path"]
Processed --> CDN["CloudFront CDN<br/>Edge caching"]
CDN -->|Serve| User
end
style Upload fill:#3b82f6,color:#fff
style Process fill:#7c3aed,color:#fff
style Serve fill:#059669,color:#fff

PatternDescriptionUse Case
Direct Upload (Presigned URL)Client uploads directly to S3Scalable uploads, no server bottleneck
Transcoding PipelineUpload triggers processing workflowVideo/image processing
CDN + Origin ShieldCDN caches files, shield protects S3Global media delivery
Multi-tier StorageAutomatic lifecycle transitionsCost optimization
Versioned StorageKeep all file versionsDocuments, collaborative editing
Geo-redundant StorageReplicate across regionsDisaster recovery, global access

DecisionProsCons
Object storageScalable, cheap, global accessNo POSIX, slower than block
Direct uploadNo server bottleneckMore client-side complexity
Server-side uploadFull control, validationServer bottleneck for large files
All hot storageFast, simpleExpensive
Lifecycle policiesCost-optimizedRetrieval delays for old data

StrategyDescription
Presigned URLsClients upload directly to S3 — servers don’t handle file data
S3 event notificationsTrigger processing asynchronously on upload
CDN with origin shieldReduce origin load, protect S3 from traffic storms
Multipart uploadLarge files uploaded in parallel chunks
Storage classes + lifecycleAuto-move old data to cheaper storage
Replication across regionsDisaster recovery, lower global latency

  1. What’s the difference between object, block, and file storage?
  2. How does a presigned URL work and why is it useful?
  3. How would you design a system that supports 1M file uploads per day?
  4. How do you optimize storage costs for user-generated content?
  5. What happens when you need to serve files to users worldwide?

SystemStorage Architecture
NetflixS3 for master files, CloudFront CDN for streaming
Google DriveGoogle File System (GFS) — distributed, replicated blocks
DropboxCustom block storage, deduplication, delta sync
InstagramS3 for photos, CDN for serving, lifecycle to Glacier for old content

  • Object storage (S3) = store any files, any size — scalable, HTTP API, great for static assets
  • Block storage (EBS) = fast drive attached to a server — for databases, app data
  • File storage (EFS) = shared network drive — multiple servers access the same files
  • Presigned URLs let clients upload directly to S3 — servers don’t handle file data
  • CDN + storage = files served globally from edge caches, origin storage protected
  • Use lifecycle policies to move old data to cheaper storage automatically
  • For video/images, use transcoding pipelines that process uploads asynchronously