Architecture of a URL Shortener

Published on

Architecture!

I haven’t been following many YouTube channels lately, but one that always leaves me satisfied whenever I watch is Renato Augusto’s channel.

Recently, I watched the video he made about URL shorteners. Overall, it’s hard to argue that his solution is probably the best possible given the problem’s premises:

Functional Requirements:

  1. URL Shortening: given a long URL -> return a much shorter URL
  2. URL Redirection: given a short URL -> redirect to the original URL

Non-functional Requirements:

  1. The system must support 100 million URLs generated per day
  2. The shortened URL should be as short as possible
  3. Only numbers (0–9) and letters (az, AZ) are allowed in the URL
  4. For every 1 write operation, there will be 10 read operations
  5. The average length of a stored URL is 100 bytes
  6. URLs must be stored for a minimum of 10 years
  7. The system must operate with high availability (24/7)

Estimates

  • Write operations: 100 million URLs per day = 100,000,000 / 24 / 60 / 60 = 1160 requests per second
  • Read operations: 10:1 = 1160 * 10 = 11600 requests per second
  • URL storage period: 10 years = 100,000,000 * 365 * 10 = 365 billion records
  • Storage capacity: 100 bytes per URL = 365 billion * 100 bytes = 46.5 TB

Given these premises, he built an architecture using Cassandra to store URLs, Redis for auto-increment and caching, and a function that performs base62 obfuscation with hashid to generate the shortened URL.

My reaction!
Real images of my reaction to the video

Go watch the video. It’s good stuff.

Elegant. I spent some time thinking about it and I believe the base62 part is truly the best solution here. It’s simple, robust, self-contained, and (aside from Redis) has no “moving parts.”

In an ideal world, his solution is the best. But the problem is that, in practice, we usually have constraints around what technologies we can use: especially in established companies. With that in mind, I’ll expand the requirements by adding a hypothetical one:

  • the company uses MariaDB, and the DBA team has denied the use of Cassandra

This restriction introduces what, for me, is the major difficulty in the entire described system. Cassandra was solving the main architectural bottleneck:

1160 write requests per second

On the read side, we can optimize cache hits (even without Redis). For ID generation, we can keep using the base62 approach (the same logic works in any language). But the write problem: we were (correctly) delegating it to a database built for this scale. And once we remove Cassandra from the diagram, we now have to deal with that load ourselves.

In fact, in the “normal day-to-day” of web systems, this is the bottleneck we deal with the most: database overload. Unless your application has some atypical behavior, most mechanisms and patterns exist to avoid unnecessary hits to the database. For reads, some caching structure is always helpful: but what about when you need a lot of writes?

And to avoid “cheating” with some MariaDB-specific trick, I’ll attack the problem assuming the database is used without heavy custom optimizations, so we can solve the problem architecturally: even if the DB were any SQL-like system.

The Problem

It’s not just a matter of optimizing writes in this scenario: any turbulence in the database will directly impact the application, whether backup routines, another process stealing resources, or maintenance tasks. The objective is to increase application availability by keeping the main bottleneck out of the hot path.

Here’s an example of the standard synchronous approach. Run a load test using the synchronous endpoint, increase the volume, and then cause some instability in the database. This is exactly what would happen in production and the one who suffers is the user.

Synchronous
No surprises

There are many ways to deal with this. Let’s cover a few:

Asynchronous Write with Eventual Consistency

I briefly talked about this in another article. The idea is to reply quickly to the client, not with the final answer, but with a promise of execution. The client sends a POST, we return a 200-range status, and the client can poll to verify when the processing is finished.

Here’s an example using RabbitMQ as a processing queue.

Asynchronous

Notice that we now have two different mechanisms: one that receives the client request, and one that processes the queue. This means we can optimize both parts independently. We can even automatically scale the number of consumers if the queue grows, and do so respecting database limits. This allows your application to remain minimally responsive even if the DB becomes unstable.

In this URL shortener scenario, there’s an interesting architectural detail: generating URLs isn’t computationally heavy (thanks to the elegant base62 mechanism mapping the incremental ID directly to the shortened URL). This means that, in the client’s initial request, it is theoretically possible to already return the shortened URL.

The problem is that, with the queue in the middle, we would return the shortened URL in the POST response, but the data would not yet exist in the database, meaning the GET would fail. And this brings us to the next architectural pattern:

Asynchronous Write with Write-Behind Cache

Since we know the “final URL” just from the incremental ID, we can store in cache the “things to be written,” so that even before the queue is consumed, the application can correctly respond to GET. This gives you all the benefits of the two independent mechanisms mentioned earlier while increasing self-sufficiency.

Your cache becomes more cluttered, since it must handle two types of data, and you need extra care with entries that can or cannot be automatically evicted. Additionally, the queue consumer needs a mechanism to “promote” the write-pending cache entry into a regular cache entry after persistence.

Asynchronous with write-behind

Notice that the operational complexity increased significantly. This is never great, but availability and robustness increase. Here’s an example implementation.

In this scenario, with a well-configured read cache, your application stays up even if the database goes down. Of course, it won’t respond correctly to everything, since not all URLs will be in the cache. But it lets you escape the dreaded “100% outage.” It’s also important to note that this is only fully achievable without the database because of the external incremental ID and the simplicity of the processing.

There’s a reasonable question: Why do we need the queue if the write cache already has all the information? Wouldn’t it be easier to periodically flush the cache into MariaDB?

The simple answer is no. The queue solves a crucial problem: durability at high write throughput. Unless your scenario accepts data loss (most applications don’t), relying purely on in-memory cache (even clustered) still means volatile data. You must write to something durable to ensure final storage consistency.

Now, if your domain can tolerate some data loss…

Memory as Primary Storage with Periodic Consistency

Since the required throughput volume is extremely high, and we want to simplify architecture by minimizing components, we can store data in memory and treat memory as the primary source of truth.

The idea is to have a large cache with hot data (both writes and reads) and a periodic mechanism that persists this data to durable storage. The durable storage can also hold historical data—so if something isn’t in the cache, we can fetch it from the DB. But writes go straight to memory, and only occasionally do we flush them to the DB.

Memory as Primary Storage

In this scenario, there is a window where data may be lost if the cache goes down. For a URL shortener, this is unacceptable, but there are scenarios where this volatility gives you the speed you need with lower complexity and moderate risk.

As a general recommendation, it’s a “avoid”, but use your judgment.

There are other approaches, such as event sourcing, but that’s content for another post.

A complete example of all scenarios can be found here. It’s not production code, but it’s enough to compare the approaches.

Comments

I feel that comments on specific blogs have been dying down as the times goes. If you have any questions or want to talk about the post, contact me through the below links.