AI system design·rate limitingintermediate9 min
AI system design interview: what changes when a model is the backend
AI system design is classic system design with a model as the slowest and most expensive component. The interview still asks about latency, capacity, caching, rate limits and failure, and each of those gets a new number.
Say you're asked to design a customer support assistant. A user types a question, the app adds the account's recent orders, and a model writes the answer. The classic questions still apply: how many users, how fast must it answer, what happens when it fails. The answers look different, because the backend is a 70-billion-parameter model that needs about 140 GB of memory to load and takes seconds to write one reply.
Weights, 70B
140 GB
Per token
320 KB
First token
~400 ms
Stream
~40 tok/s
This guide opens the AI track. It walks through the six places where the model changes the design, with the numbers interviewers expect you to know. Later guides take each one on its own: the KV cache, time to first token, semantic caching and the LLM gateway. After those come the design problems: ChatGPT, RAG and an AI agent.
Scroll →
First token
400 ms
Tokens/s per chat
40
Chats in memory
10
429 per second
0
The gateway counts the prompt's tokens, finds no cached answer, and streams the reply from the model servers. The first token arrives in about 400 ms, then about 40 a second.
Step through the figure. Each step matches one of the sections below.
The reply takes seconds, so the latency budget has two numbers#
A classic API answers in tens of milliseconds and the whole response arrives at once. A model writes its answer one token at a time. A token is a piece of a word, about three quarters of a word on average, so a 300-word reply is about 400 tokens.
Two numbers describe the experience. Time to first token is how long the user waits before anything appears. Tokens per second is how fast the text then flows. For a chat, the first token should arrive in under a second, and 40 tokens a second is faster than anyone reads.
The first number comes from prefill. The model reads the whole prompt before it can write a word. A 2,000-token prompt on one modern GPU takes a few hundred milliseconds of pure compute. The second number comes from decode. For each new token the model reads all of its weights once, all 140 GB. Spread over two GPUs with about 3 TB/s of memory bandwidth each, that is roughly 25 ms per token, around 40 tokens a second.
So the design streams. The app sends tokens to the phone as they come, and the user starts reading while the model is still writing. A reply that takes 10 seconds in total feels fine when the first words show up in half a second.
In an interview, give both numbers and say which one the product cares about. A chat cares about time to first token. A batch job that summarises 10,000 tickets overnight cares about total tokens per second across the fleet. The p99, the time the slowest one in a hundred requests takes, is measured on the first token.
Capacity is GPU memory, and the KV cache uses it up#
A classic service runs out of CPU. An inference server runs out of memory, and the thing that fills it is the KV cache.
While the model works on a conversation, it keeps a record of every token it has read so far. That record is the KV cache. Without it, each new token would mean re-reading the entire conversation. For a 70B model, each token costs about 320 KB. That is 80 layers, times 8 key-value heads, times 128 numbers per head, times 2 for key and value, times 2 bytes each.
Now do the arithmetic for one conversation. A 4,000-token context, the prompt plus the reply so far, is 4,000 times 320 KB, about 1.3 GB. Put the model on two 80 GB GPUs. The weights take 140 GB of the 160, and the server needs a few GB for its own work, so call it 15 GB for conversations. That is about a dozen conversations in memory at the same time.
A dozen is the number interviewers want to hear you work out, because it is the real concurrency of that pair of GPUs. Request number thirteen waits, however many app servers sit in front of the model.
The GPUs can serve those twelve at nearly the price of one. Decode is limited by reading the weights, and one read serves every conversation in the batch. This is called continuous batching. The server adds and removes conversations from the batch as they arrive and finish, and each token step is one pass over the weights for all of them.
Three levers raise the ceiling. Shorter contexts, so each conversation uses less memory. Quantised weights, 8-bit instead of 16-bit, which halves the 140 GB. More GPUs, which is the lever that costs money. The capacity estimate in this interview is memory per token, times context length, times concurrent conversations, against the GPU memory left after the weights.
Rate limits count tokens, because one request can cost a thousand#
An API rate limiter counts requests. A model's cost grows with the size of each request. One call that pastes in a 90,000-token document costs as much GPU time as a thousand short chats. A limit of 60 requests a minute lets that one key take the whole cluster.
So the gateway in front of the model keeps a budget in tokens per minute per key, for example 100,000. The usual tool is a token bucket, a counter that refills at a steady rate, and each request spends one token. Here the word token is doing two jobs. The bucket's tokens are units of budget, and each request spends as many of them as its prompt and reply contain.
There is one catch. The prompt's size is known when the request arrives. The reply's size is only known when the model stops. So the gateway reserves the request's maximum reply length up front, then gives back what the reply didn't use.
When a key is over budget the gateway returns 429, the status that tells one client it has sent too many requests. The other keys never notice. When the model servers themselves are full, the right status is 503, the status that says the service itself is unavailable to everyone, and every client should back off.
The gateway is also where you put the things every AI request needs. That means a key check, logging of prompt and reply sizes, a per-team cost meter, and the fallback described below.
Caching an answer means deciding how close is close enough#
Two requests to a classic cache either have the same key or they don't. Two questions to a model are almost never spelled the same, so a cache that needs an exact match hits almost nothing.
A semantic cache stores each question as an embedding, a list of numbers that places similar sentences near each other. It answers a new question from an old reply when the two are near enough. Near enough is a threshold you choose, for example a similarity of 0.95. "How do I reset my password" and "I forgot my password, how do I change it" land on the same reply. The hit costs 20 ms and no GPU time.
The threshold is the whole design. Set it low and "What is my balance" matches "What is my limit", and the user gets a wrong answer stated with confidence. Set it high and the hit rate drops to a few percent. Personal and time-sensitive questions should never be served from the cache at all, whatever the threshold.
The cheaper cache lives inside the model server. Most requests share the same system prompt, the instructions the app puts in front of every question, often 2,000 tokens or more. The server can keep that prefix's KV cache entries and skip prefill for it. Providers charge a fraction of the price for cached prefix tokens, a tenth with some, and the first token arrives sooner.
Both caches have a TTL, how long a cached value is allowed to live. When a popular cached answer expires, every request for it misses at once and all go to the model together. That is the same cache stampede a classic cache has, a thundering herd, many callers missing at once and all hitting the database. Single-flight, where one request does the work while identical ones wait for its result, fixes it the same way.
Agents make calls that must be safe to repeat#
An agent is a model in a loop. It reads the task, picks a tool, reads the result, and picks again until it decides it's done. Each tool call is an ordinary request to one of your services, and the loop can run twenty of them for one user question.
The model is unreliable in a specific way. It may call the same tool twice, or the loop may be retried after a crash halfway through. If the tool is "refund this order", twice means a real refund paid twice. So every tool call that changes something carries an idempotency key. The tool is idempotent, safe to repeat, because doing it twice has the same result as doing it once.
The loop also takes time, often a minute or more. A request that holds an HTTP connection for a minute will time out somewhere between the phone and the server. So the app accepts the task, returns 202, accepted, meaning saved and promised, not yet done, and puts the job on a queue. A worker runs the loop and streams progress back over a connection the phone opens for it.
Queues bring at-least-once delivery, where every message is delivered, and some are delivered more than once. The same idempotency keys cover that case too.
A step that spends money or sends a message to a customer gets an approval gate. The loop pauses, a person confirms, and the loop continues. Interviewers ask where that gate goes. It goes in the loop, before the tool call, so the tool itself stays simple.
When the model fails, degrade before you refuse#
Classic dependencies fail by going down. A model also fails by being slow, by returning nonsense, and by running out of budget at the end of the month. The design needs an answer for each.
The gateway wraps the model in a circuit breaker, a switch that stops calling a failing dependency for a while so it can recover. When the breaker opens, requests go to a fallback: a second provider, or a smaller model on separate hardware. The fallback may answer slower or less well, and nobody sees an error.
Under load, shed in order. Shorten the maximum reply length first, so each request costs less. Then route free-tier traffic to the smaller model. Then return 503 to the lowest-priority callers. Each step keeps paying customers on the full model a little longer.
Watch the numbers that show trouble early: time to first token at the p99, GPU memory in use, tokens per minute per key, cache hit rate, and cost per day. A model that is slowing down shows up in the first two before anyone complains.
What to say in an interview#
Ask five questions before you draw anything.
- Is the model ours or an API? Self-hosted means you own the capacity math. An API means you own the budget and the fallback.
- What is the latency budget, in time to first token and tokens per second?
- How long are the prompts and the replies? That sets memory per conversation.
- Which answers can be cached, and how old may they be?
- What happens when the model is down or the budget is spent?
Then map each classic topic to its new number.
| Classic topic | With a model behind it | The number to say |
|---|---|---|
| Latency | First token, then a stream | Under 1 s to first token, 40 tokens a second |
| Capacity | GPU memory per conversation | 320 KB per token, about a dozen 4k chats on two 80 GB GPUs |
| Rate limiting | Tokens per minute per key | One 90k-token request costs as much as 1,000 chats |
| Caching | Similarity threshold, shared prefix | 0.95 similarity, cached prefix at a tenth of the price |
| Queues | Agent loop with idempotent tool calls | 202 and a job, an idempotency key on every write |
| Failure | Breaker, fallback model, shed by tier | Shorter replies first, 503 last |
Practice this in the app: graded drills, with the misses brought back.
Start freeRelated guides
Updated