TL;DR URL (Uniform Resource Locator) and URI (Uniform Resource Identifier) are often used interchangeably, but they're not identical. A URI identifies a resource; a URL locates it (specifies how to access it). Every URL is a URI, but not every URI is a URL. For example, urn:isbn:978-0-306-40615-7 is a URI (it identifies a book) but not a URL (no protocol, no location). In everyday web engineering, you'll never use URIs that aren't also URLs.
You type a URL and press Enter. A web page appears. Simple. But between those two events, your browser executes a precise choreography of protocol handshakes, cache lookups, and network negotiations โ and understanding every step is what separates engineers who guess at performance problems from those who solve them.
Read the Deep Dive โ Open Network Lab ๐ ๐ https:// example.com /products /shoes.html Table of ContentsYou've typed thousands of URLs. But do you actually know what every character means? A URL isn't just an address โ it's a structured instruction set that tells your browser how to connect, where to connect, and what to ask for. Missing any one of those four elements changes the entire request.
A URL has four parts. The scheme (also called protocol) comes first: http:// or https://. This isn't decoration โ it's a hard instruction to the browser specifying exactly which application-layer protocol to use for the connection. HTTP means plain text over TCP. HTTPS means TLS-encrypted TCP. The browser reads the scheme before it does anything else. The domain follows: example.com. This is a human-readable alias for an IP address โ your browser can't actually connect to "example.com" as-is, which is why DNS exists (and why we'll cover it next). The path and resource work together like a directory and filename: /products/shoes.html. The path tells the server which section of the site you want; the resource tells it which specific file or endpoint within that section.
Here's the thing most tutorials miss: the division between path and resource is a convention, not a requirement. On a REST API, /api/users/123 might be a "path" with no separate resource โ it's all just a string the server parses. On a static file server, /images/logo.png maps to a real filesystem path. The URL spec doesn't enforce the distinction โ the server interprets it however it wants. This is why some frameworks route on the full path and others strip the last segment as a filename.
URL (Uniform Resource Locator) and URI (Uniform Resource Identifier) are often used interchangeably, but they're not identical. A URI identifies a resource; a URL locates it (specifies how to access it). Every URL is a URI, but not every URI is a URL. For example, urn:isbn:978-0-306-40615-7 is a URI (it identifies a book) but not a URL (no protocol, no location). In everyday web engineering, you'll never use URIs that aren't also URLs, but the distinction matters in spec documents and when dealing with XML namespaces or RDF graphs.
from urllib.parse import urlparse
url = "https://example.com/products/shoes.html?size=10&color=red#reviews"
parsed = urlparse(url)
print(ff"Scheme: {parsed.scheme}") # https
print(ff"Domain: {parsed.netloc}") # example.com
print(ff"Path: {parsed.path}") # /products/shoes.html
print(ff"Query: {parsed.query}") # size=10&color=red
print(ff"Fragment: {parsed.fragment}") # reviews (browser-only, never sent to server!)
# The fragment (#reviews) is NEVER sent to the server.
# It's purely client-side โ the browser scrolls to the element with id="reviews".
# Servers never see URL fragments, which surprises many developers.
# This is why client-side routers (React Router) use fragments for routing.
Your browser has the URL example.com. But to open a TCP connection, it needs an IP address โ the numerical address computers actually use to route traffic. Domain names are for humans; IP addresses are for machines. The Domain Name System (DNS) is the translation layer between them, and it's one of the most elegant distributed systems on the internet.
The lookup process is a cascade of caches. First, your browser checks its own DNS cache โ it already looked up google.com ten minutes ago and stored the result with a TTL (Time to Live). No need to look it up again. If the browser doesn't have it, it asks the operating system, which has its own cache. Still nothing? The OS queries a DNS resolver โ usually provided by your ISP or configured as a custom resolver like 8.8.8.8 (Google) or 1.1.1.1 (Cloudflare). The resolver is the workhorse. It walks the DNS hierarchy: it asks the root name servers which authoritative server handles .com, then asks the .com name server who handles example.com, then asks example.com's authoritative server for the actual IP. The answer comes back and is cached at every level โ the resolver, the OS, and the browser โ each for the TTL specified in the DNS record.
The counterintuitive insight: DNS caching means that changing a domain's IP address doesn't propagate instantly. If you update your A record from IP X to IP Y, visitors with the old IP cached will keep going to IP X until their cache TTL expires. This is why DNS TTLs are a deployment tool โ lower them to 60 seconds before a migration, make the switch, then raise them back to 3600. Engineers who don't understand DNS TTLs have caused multi-hour outages from incorrect migration assumptions.
โ Reduce DNS Latency with Prefetch HintsFor web performance, DNS lookup latency (10-100ms per lookup) adds up quickly when your page loads resources from multiple domains. The fix: DNS prefetch hints tell browsers to resolve domains before they're needed: <link rel="dns-prefetch" href="//cdn.example.com">. For critical cross-origin resources, go further with <link rel="preconnect" href="https://cdn.example.com"> โ this performs DNS lookup, TCP handshake, and TLS handshake all in advance. The performance difference is measurable: 50-200ms per domain eliminated from your critical path.
# See full DNS resolution chain $ dig +trace example.com # Check cached DNS entry with TTL $ dig example.com # Answer section shows TTL in seconds โ how long result is cached # Look up using specific resolver (1.1.1.1 = Cloudflare) $ dig @1.1.1.1 example.com # Check what your OS has cached (macOS) $ sudo dscacheutil -statistics # Flush DNS cache (macOS) $ sudo dscacheutil -flushcache && sudo killall -HUP mDNSResponder # Check response time for DNS lookup $ time nslookup example.com # Real 0.025s โ from cache (25ms) # Real 0.098s โ from resolver (98ms) # Real 0.450s โ full recursive lookup (450ms)
Your browser now has the IP address. Before it can send a single byte of HTTP data, it must establish a TCP connection with the server. TCP (Transmission Control Protocol) is a connection-oriented protocol โ both sides must agree to communicate before data flows. This agreement is the three-way handshake: the client sends SYN (synchronize), the server responds SYN-ACK (synchronize-acknowledge), the client sends ACK (acknowledge). Three messages, at minimum one round-trip time (RTT) to complete. At 50ms network latency, that's 50ms before you've sent a single character of your HTTP request.
The three-way handshake isn't just ceremony โ it establishes sequence numbers on both sides (so both parties know which packets belong to which order), negotiates connection parameters, and verifies that the server is actually reachable and responsive before the client invests resources in sending data. TCP's reliability guarantees (ordered delivery, retransmission of lost packets, flow control) come from this connection-oriented design.
Keep-alive connections exist specifically to avoid paying the handshake cost repeatedly. Without keep-alive, every resource on a webpage (HTML, CSS, JS, images) would require a new TCP connection โ each with its own handshake. With keep-alive, the browser reuses an established connection for multiple requests. HTTP/1.1 has keep-alive on by default. HTTP/2 takes this further with multiplexing โ multiple requests fly over a single TCP connection simultaneously without waiting for each other. HTTP/3 (QUIC) eliminates TCP entirely in favor of UDP, removing head-of-line blocking at the transport layer.
๐ก TCP's Slow Start Affects Your First RequestTCP has a built-in congestion control mechanism called slow start: new connections begin with a small congestion window and increase it exponentially with each acknowledged packet. This means the first request on a new connection is deliberately throttled โ TCP doesn't trust the network yet. For large responses (> ~14KB initial window), slow start can significantly delay time-to-first-byte. This is one reason CDNs dramatically improve performance: they maintain warm, pre-established connections to end users, eliminating both the handshake overhead and slow start for every request. TCP Fast Open (TFO) is a newer mechanism that allows data to be sent in the SYN packet, eliminating one RTT for repeat connections.
If you're connecting to https://, after the TCP handshake completes, there's still another handshake before any application data can flow: the TLS (Transport Layer Security) handshake. This is the "expensive" part of HTTPS โ computationally expensive on the server, and latency-expensive for the user. Understanding why it costs what it costs, and how modern browsers minimize that cost, is essential knowledge for any web performance engineer.
A TLS 1.3 handshake works in roughly three steps: the client sends a ClientHello containing supported cipher suites and a key share. The server responds with its certificate (proves its identity) and a key share of its own. Both sides independently derive the same encryption keys from those key shares (using Diffie-Hellman key exchange โ mathematics that lets two parties agree on a shared secret without ever transmitting it directly). The total cost: 1 RTT in TLS 1.3 (down from 2 RTTs in TLS 1.2). For TLS 1.3 with 0-RTT resumption (for reconnecting sessions), it can be effectively zero additional RTTs โ data can be sent with the first message.
SSL session resumption is the browser's primary trick for reducing TLS cost on repeat visits. After a full handshake, the server issues a session ticket โ an encrypted token containing the session parameters. On the next connection, the client presents this ticket, and the server can resume the session without a full handshake. This is why HTTPS connections to sites you've visited recently feel as fast as HTTP โ the full handshake cost is amortized over many requests.
โ ๏ธ Certificate Validation Adds Hidden LatencyTLS certificate validation can add surprising latency that doesn't show up in basic timing measurements. When a browser validates a certificate, it may need to check if it's been revoked via OCSP (Online Certificate Status Protocol) โ this is a live network request to the CA's server. If the CA's OCSP server is slow or unreachable, certificate validation blocks the TLS handshake. Mitigation: OCSP Stapling, where the web server periodically fetches and caches the certificate's OCSP status and sends it to browsers as part of the TLS handshake, eliminating the browser's round trip to the CA. OCSP Stapling is a standard Nginx/Apache configuration option and can save 50-200ms on first-visit handshakes.
tls_timing.sh โ inspect TLS handshake cost# Time TLS handshake components with curl verbose
$ curl -w "DNS: %{time_namelookup}s\n
TCP: %{time_connect}s\n
TLS: %{time_appconnect}s\n
First Byte: %{time_starttransfer}s\n
Total: %{time_total}s\n" -o /dev/null -s https://example.com
# Typical output:
# DNS: 0.025s โ from cache
# TCP: 0.075s โ 50ms RTT for handshake
# TLS: 0.175s โ 100ms for TLS 1.3 (1 RTT)
# First Byte: 0.200s โ server processing time
# Total: 0.310s โ with response
# Check TLS version and cipher suite used
$ openssl s_client -connect example.com:443 -servername example.com 2>&1 | head -20
# Inspect certificate validity and OCSP stapling
$ echo Q | openssl s_client -connect example.com:443 -status 2>&1 | grep -E 'OCSP|Protocol|Cipher'
After all the infrastructure work โ DNS resolution, TCP handshake, TLS negotiation โ the actual HTTP exchange is almost anticlimactic in its simplicity. HTTP itself is a text-based protocol: the client sends a request message as plain text, the server reads it and sends back a response message as plain text (with a body that might be anything from HTML to JSON to binary image data). That's it.
An HTTP request has three parts: a request line (method, path, version: GET /products/shoes.html HTTP/1.1), headers (metadata: Host, Accept, Content-Type, Authorization, Cache-Control, etc.), and an optional body (for POST/PUT/PATCH requests that send data). An HTTP response mirrors this structure: a status line (version, status code, reason: HTTP/1.1 200 OK), headers (Content-Type, Content-Length, Cache-Control, Set-Cookie, etc.), and the response body (the actual HTML, JSON, image, etc.). The server processes the request, constructs the response, and sends it back over the same TCP connection.
Imagine you're ordering at a restaurant. The request line is your order: "One burger please" (method + resource). The request headers are extra context: "No pickles, allergic to gluten" (Accept, Authorization). The response status line is the kitchen's acknowledgement: "Order received, coming right up" (200 OK). The response headers are metadata: "ready in 10 minutes, serves 1" (Content-Type, Content-Length). The response body is the actual burger. HTTP is exactly this exchange, just with bytes instead of food.
โ Status Codes Are Contracts โ Respect ThemHTTP status codes are more than cosmetic. Caches use them: 200 means cache normally, 301 Moved Permanently means update bookmarks and cache forever, 304 Not Modified means use your cached copy, 503 Service Unavailable means retry later. Client applications react to them: 401 Unauthorized means missing credentials, 403 Forbidden means credentials present but insufficient, 404 Not Found means stop retrying this resource. The most common mistake: returning 200 OK with an error message in the body. This breaks caches, breaks retry logic, and breaks monitoring. If it's an error, use an error status code. 4xx for client mistakes. 5xx for server mistakes.
http_raw.txt โ what actually travels over the wire--- HTTP REQUEST (what browser sends) --- GET /products/shoes.html HTTP/1.1 Host: example.com User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) ... Accept: text/html,application/xhtml+xml,*/*;q=0.9 Accept-Encoding: gzip, deflate, br Accept-Language: en-US,en;q=0.9 Connection: keep-alive โ reuse this TCP connection Cache-Control: max-age=0 โ check for fresh version If-None-Match: "abc123" โ "I have this version cached" (blank line marks end of headers) --- HTTP RESPONSE (what server sends back) --- HTTP/1.1 200 OK Content-Type: text/html; charset=UTF-8 Content-Length: 4823 Content-Encoding: gzip โ compressed response body Cache-Control: max-age=3600 โ cache for 1 hour ETag: "abc123" โ fingerprint for conditional requests Connection: keep-alive (blank line, then compressed HTML body)
The browser receives the HTML response and starts parsing it. Almost immediately, it encounters references to more resources: a <link rel="stylesheet"> for CSS, <script src="..."> for JavaScript, <img src="..."> for images. Each of these is a separate resource that requires its own request. For a typical modern webpage, that can be 50 to 200 additional requests.
For each new resource, the browser goes through the same process: DNS lookup (usually cached), TCP connection (reused via keep-alive when possible), TLS handshake (resumed via session tickets), HTTP request and response. The difference from the initial page load is that many of these can happen in parallel. HTTP/1.1 browsers open up to 6 parallel TCP connections per domain (a practical limit hardcoded into browser implementations). HTTP/2 eliminates the per-domain connection limit by multiplexing all requests over a single connection. HTTP/3 further improves this by eliminating TCP's head-of-line blocking.
Here's the thing most web performance tutorials miss: the critical rendering path is not just about how many resources load, but which resources block rendering. CSS and synchronous JavaScript block the browser from rendering any content until they're downloaded and processed. Images and async scripts don't. A 50KB CSS file blocking render is far more damaging to user experience than a 500KB image that loads asynchronously. Understanding which HTTP requests are on the critical path โ and minimizing them โ is the foundation of web performance optimization.
๐ฌ Resource Hints Change When Requests FireModern browsers support resource hints that let you control when certain requests happen: rel="preload" fetches a resource immediately (before the parser would normally find it); rel="prefetch" fetches a resource for likely future navigation (low priority, runs in idle time); rel="preconnect" establishes the TCP+TLS connection to a domain before it's needed. These hints are powerful performance tools but require discipline โ over-prefetching wastes bandwidth, and over-preloading resources fight for bandwidth with critical-path resources. Always measure with WebPageTest or Chrome DevTools before and after adding resource hints.
The complete journey from URL entry to rendered page: browser parses the URL (extracts scheme, domain, path, resource) โ DNS lookup cascade (browser cache โ OS cache โ resolver โ authoritative DNS) โ TCP three-way handshake (SYN/SYN-ACK/ACK, 1 RTT) โ TLS handshake for HTTPS (1 RTT for TLS 1.3, or 0 with 0-RTT resumption) โ HTTP GET request (headers + no body for GET) โ server processing โ HTTP response (status, headers, body) โ browser parses HTML โ parallel requests for additional resources โ critical-path resources (CSS, sync JS) block rendering โ non-critical resources load asynchronously โ page renders.
Every step in this pipeline has a performance lever. DNS: cache aggressively, use prefetch hints for third-party domains. TCP: keep-alive and HTTP/2 multiplexing to minimize handshakes. TLS: TLS 1.3 + session resumption + OCSP stapling. HTTP: compression, caching headers, ETags for conditional requests, HTTP/2 server push for critical resources. Resource loading: preload critical assets, defer non-critical scripts, minimize render-blocking resources. Understanding the full pipeline is what lets you know which lever to reach for.
Four experiments: URL parser, DNS resolver, connection timing calculator, and HTTP request builder.
URL Anatomy Parser Enter any URL Parsed ComponentsDNS resolution cascade โ watch queries flow down and answers flow back up
DNS Lookup Simulator Domain to resolve Cache status Browser cache โ DNS hops โ Lookup time (est)Waterfall: time spent in each connection phase (ms)
Connection Timing Calculator Round-trip time (RTT ms) 50 Protocol HTTPS / HTTP/2 DNS cached? Yes TLS session resumption? Yes โ DNS โ TCP โ TLS โ Total before HTTP HTTP Request Builder Method URL Headers (include) Authorization: Bearer token Accept: application/json Cache-Control: no-cache Request body (POST/PUT) Generated HTTP Request Example Response