Find your breaking point before your users do: designing realistic load scenarios, driving them with Locust or k6, reading the latency percentiles that matter, locating the true bottleneck, and turning results into capacity plans and autoscaling signals.
Every app has a point where added load turns healthy latency into a cliff — requests queue, timeouts cascade, and the site falls over. The only question is whether you find that cliff in a controlled test or during your busiest hour in production. This tutorial covers load-testing Django properly: designing realistic scenarios, driving them with Locust or k6, reading the percentiles that matter, finding the true bottleneck, and turning results into capacity plans and autoscaling signals rather than a reassuring but meaningless "it handled a thousand requests".
Functional tests prove the app is correct for one user; they say nothing about what happens under a thousand concurrent ones. Systems do not degrade gracefully — they are fine, fine, fine, then they hit a resource limit (connections, CPU, a lock) and collapse non-linearly. Load-testing finds that limit deliberately, so you know your real capacity, can plan for growth and traffic spikes, and can prove a change made things faster rather than hoping. Skipping it means your first true capacity test is a real incident with real users.
Settle the scope before any script runs. To the target, a load test looks exactly like a denial-of-service attack: thousands of requests per second from a handful of addresses, deliberately pushed until something breaks. Only ever point load at systems you own or have explicit, written authorisation to test. That rule covers more than the obvious case of someone else's website. It also covers the infrastructure and third parties your own app depends on:
Write the scope down: the target hostnames (for example staging.example.com), the time window, the peak rate, who is on hand to stop the test, and the condition that aborts it. Pass the target host explicitly on each run and never bake a default into the script. That way nobody fires the stress profile at production, or at a host they have no right to test, just by forgetting a flag.
Three distinct tests answer three questions. A load test holds an expected level of traffic and checks the system meets its latency targets there. A stress test pushes past expected load to find the breaking point and see how it fails — cleanly with shed load, or catastrophically. A soak test runs moderate load for hours to surface slow problems: memory leaks, connection exhaustion, disks filling, caches degrading. You need all three eventually; they catch different classes of failure and a system can pass one while failing another.
The most common load-testing mistake is testing something users never do — hammering one cheap endpoint and declaring victory. Real traffic is a mix: some users browse, some search, a few check out, each with think-time pauses between actions. Model that. Weight the scenarios by real proportions from your analytics, include the expensive paths (search, checkout, report generation) that actually strain the system, and add pauses so you simulate users rather than an unrealistic tight loop. A test that mirrors real behavior predicts real capacity; one that does not produces confident nonsense.
Two tools dominate. Locust is Python, so scenarios are Django-developer-friendly and easy to make dynamic, with a live web UI. k6 is a Go-based tool scripted in JavaScript, excellent for CI and high load per machine. Either works; pick Locust if you want to stay in Python and reuse app knowledge, k6 if you want a lean binary and first-class automation. Avoid ad-hoc tools like a bare ab loop for anything beyond a smoke check — they cannot express realistic multi-step user behavior.
A Locust scenario is a class of tasks with weights and wait times, expressed in plain Python.
from locust import HttpUser, task, between
class ShopUser(HttpUser):
wait_time = between(1, 5) # think-time between actions
@task(10)
def browse(self):
self.client.get("/products/")
@task(3)
def search(self):
self.client.get("/search/?q=django")
@task(1)
def checkout(self):
self.client.post("/cart/add/", json={"sku": "ABC"})
self.client.post("/checkout/")
The weights encode the real mix — ten browses per checkout — and between(1, 5) gives each simulated user human pauses. Run it ramping users up gradually and watch where latency starts to climb.
k6 describes load with scenarios, and each scenario has an executor that decides how iterations are scheduled. The key choice is between VU-based executors, which keep a fixed number of virtual users looping like Locust does, and arrival-rate executors, which start a fixed number of iterations per second regardless of how slowly the server responds. For capacity work, arrival-rate is usually what you want, for reasons the next section explains.
import http from 'k6/http';
import { check } from 'k6';
const BASE = __ENV.BASE_URL; // required: -e BASE_URL=https://staging.example.com
if (!BASE) { throw new Error('BASE_URL is required'); }
export const options = {
scenarios: {
browse: {
executor: 'ramping-arrival-rate',
exec: 'browse',
startRate: 20,
timeUnit: '1s',
preAllocatedVUs: 100,
maxVUs: 1000,
stages: [
{ target: 50, duration: '3m' }, // warm-up
{ target: 50, duration: '5m' }, // steady baseline
{ target: 400, duration: '30m' }, // slow climb through the knee
],
},
},
thresholds: {
http_req_failed: ['rate<0.01'],
'http_req_duration{name:product_detail}': ['p(95)<300', 'p(99)<800'],
http_req_duration: [
{ threshold: 'p(99)<3000', abortOnFail: true, delayAbortEval: '1m' },
],
},
};
export function browse() {
const id = Math.floor(Math.random() * 50000) + 1;
const res = http.get(`${BASE}/products/${id}/`, { tags: { name: 'product_detail' } });
check(res, { 'status is 200': (r) => r.status === 200 });
}
Unlike the Locust version, there is no sleep(). With an arrival-rate executor, pacing comes from the configured rate, and think time inside the iteration would only tie up VUs. Thresholds can target tagged sub-metrics, so each endpoint gets its own target instead of sharing one blended number; add a scenario per user journey, each with its own rate, to reproduce the traffic mix. The abortOnFail threshold is a safety valve that stops a breakpoint test automatically once it has clearly gone past the cliff. Watch the dropped_iterations metric as well. It counts iterations k6 wanted to start but could not, because every VU up to maxVUs was busy waiting on slow responses. A rising count means either the server is past its knee or maxVUs is too low.
This is the subtlest way a load test can lie. In a closed model, a fixed population of users each waits for a response and then pauses before sending the next request. Locust users and k6 VU-based executors work this way. The consequence: when the server slows down, the generator automatically sends less. The offered load is throughput = users / (think time + response time). With 100 users, 1 s of think time and 100 ms responses, that is about 91 requests per second. If responses slow to 2 s, the same 100 users send only about 33 per second. The requests it did not send are never measured, and the recorded tail looks far better than real users would experience. That effect is called coordinated omission.
Real public traffic is mostly an open model. New visitors arrive at their own rate whether or not your server is keeping up. When service time exceeds the arrival rate's budget, the queue grows without bound, and that is the cliff. An open-model test, such as k6's constant-arrival-rate or ramping-arrival-rate, reproduces this faithfully.
You can still find the cliff with Locust if you read its results with this in mind. Watch for achieved requests per second going flat or falling while the user count keeps rising. That is the closed model backing off, and it means you are already past the knee.
Three numbers matter together, and no single one is enough. Throughput (requests per second) is capacity. Latency is the user experience, but only as a distribution, not an average. And error rate is the honesty check — throughput means nothing if a third of responses are 500s or timeouts. A system "handling" high throughput while erroring or timing out is not handling it; always read all three at once.
Average latency is the most misleading number in performance work. A mean of 100ms can hide that one request in a hundred takes three seconds — and on a page that makes twenty backend calls, roughly one page in five hits that slow tail (1 − 0.9920 ≈ 18%), and over a session of a hundred requests most users will. Read percentiles: p50 (median) is the typical experience, p95 and p99 are the tail that real users feel and that averages bury. Set targets on the tail — "p99 under 500ms" — because the worst-case latency, multiplied across a page's many requests, is what determines whether the site feels fast or flaky.
A load test tells you the system slowed down; it does not tell you why. To find the cause, watch resource metrics on every tier during the test — application CPU, database CPU and connection count, cache hit rate, external-API latency. The bottleneck is whatever saturates first: usually the database (connection exhaustion or a missing index turning slow under concurrency), sometimes CPU, sometimes a rate-limited upstream. Correlate the latency climb with which resource maxed out, because optimizing anything other than the actual bottleneck moves no needle.
A load test is only as trustworthy as its environment. Testing against an empty database hides the index and query problems that appear at real data volume; testing on a laptop tells you nothing about production's network and infrastructure. Run against an environment that mirrors production — comparable instance sizes, a database loaded with production-scale (anonymized) data, the real cache and proxy in front. The most common way load tests lie is by being too easy: a fast result on a tiny, empty, local system that has nothing to do with production reality.
Turn numbers into decisions. If one instance sustains 200 requests per second at acceptable latency and you expect peaks of 800, you need four instances plus headroom for failures and spikes — never plan to run at 100% of tested capacity, because there is no margin for a bad deploy or a traffic surge. Little's Law (concurrency = throughput × latency) helps size worker and connection pools from your measured numbers. The output of load-testing is a defensible statement: "we can serve X, we provision for Y, we scale at Z".
Suppose a breakpoint test on one application instance (4 vCPU, Gunicorn with 9 sync workers) shows the knee at about 180 requests per second, with mean in-app latency of 45 ms and p99 of 350 ms at that point. Expected peak is 800 requests per second. Here is how that becomes a plan:
| Step | Calculation | Result |
|---|---|---|
| Safe operating point per instance | 60% of the 180 rps knee | ~108 rps |
| Instances for peak | 800 / 108, rounded up | 8 |
| Survive losing one instance (N+1) | 8 + 1 | 9 |
| Average in-flight requests per instance (Little's Law) | 108 rps × 0.045 s | ~5 of 9 workers busy |
| Persistent DB connections from web tier | 9 instances × 9 workers | 81 |
| Plus Celery workers, cron, admin shells | 81 + 20 + 5 | 106 |
The last row is the trap. PostgreSQL's default max_connections is 100. The ninth instance, the one added for resilience, pushes the connection count past the server's limit. The resulting failures (FATAL: sorry, too many clients already) look like a load problem even though no query is slow. Little's Law gives you the other key insight: average concurrency is only about 5 per instance, so the 81 persistent connections spend most of their time idle. That is the argument for a connection pooler such as PgBouncer in transaction mode between the app and the database, rather than simply raising max_connections.
Load-testing also tells you what to autoscale on. The right signal is the one that predicts saturation for your workload — often request concurrency or latency rather than CPU, since an I/O-bound Django app can be at its limit while CPU looks calm. Use the test to find which metric moves first as you approach the cliff, and scale on that, with thresholds set below the breaking point so new capacity comes online before users feel pain rather than after.
Performance regresses silently — an N+1 sneaks in, a cache is disabled, a dependency slows down — and you want to catch it in a pull request, not a postmortem. Run a scaled-down load test in CI against a representative environment and fail the build if p99 or throughput regresses beyond a threshold. Treating performance as a tested, gated property rather than something checked once before a big launch is what keeps an app fast as it changes over months and years.
Beyond empty databases and localhost testing, the recurring traps are: warm-vs-cold caches, where a test that runs long enough to warm every cache overstates capacity for real cold-start traffic; the load generator itself becoming the bottleneck, so you measure its limits not the server's; unrealistic uniform data that makes every cache hit and every query cheap; and reading only averages. Each one makes the numbers prettier and the conclusion wronger. A load test's job is to be honest, and most bad ones are simply too easy on the system.
How you apply load matters as much as how much. Slamming a system with full load instantly measures cold-start behavior — empty caches, unconnected pools — which is a real but different scenario from sustained peak. Use a ramp profile that adds users gradually so you can watch latency climb and pinpoint the exact concurrency where it turns the corner, and include a warm-up period so steady-state numbers are not skewed by initialization. The shape of the load curve is a parameter of the experiment; choose it to answer the question you actually have, whether that is "how do we handle a slow build-up" or "how do we survive a sudden spike".
One machine can only generate so much traffic before it becomes the bottleneck, and then you are measuring your load generator, not your server. For high target loads, run the generator distributed across several workers — both Locust and k6 support this — and confirm the generators themselves are not saturated on CPU or network. A telltale sign you have hit this trap is latency that plateaus no matter how many virtual users you add: often the server has room and the load box is maxed out.
Realistic scenarios are not static URLs — users log in, receive session and CSRF tokens, get back ids they use in the next request. Your load script must correlate this: capture the token from a login response and send it onward, use returned ids in follow-up calls, and vary inputs so every virtual user is not hitting the identical cached row. Static, uncorrelated scripts both fail against real auth and paint an unrealistically cache-friendly picture. Handling dynamic data is what separates a genuine end-to-end load test from a trivial URL hammer.
Plot latency against concurrency and you get a characteristic shape: flat and healthy, then a knee where it bends sharply upward as a resource saturates and requests start queuing. That knee is your practical capacity — not the point where errors begin, which is already past the cliff. Your target operating load should sit comfortably left of the knee, with the gap as your headroom. Learning to spot the knee, rather than chasing the maximum requests-per-second before total collapse, is the core skill of reading load-test results.
Streaming responses, server-sent events, and WebSockets need different tests than request-response endpoints, because their cost is in held open connections over time, not requests per second. Measure how many concurrent long-lived connections a worker sustains and what happens as that number climbs, using a tool that can hold connections open (k6 supports WebSockets; Locust can with an appropriate client). An HTTP-only load test badly understates the resource pressure of an app built around live connections, so test the connection model you actually run.
Here is how a real breakpoint campaign runs, with illustrative numbers from a stepped k6 ramp against a single staging instance. The instance is a production-sized copy with the anonymized production dataset, and the team runs it with written approval for a two-hour window. Each row is one three-minute step:
| Offered rps | Achieved rps | p50 | p95 | p99 | Errors | DB active conns | App CPU |
|---|---|---|---|---|---|---|---|
| 50 | 50 | 38 ms | 95 ms | 160 ms | 0% | 3 | 18% |
| 100 | 100 | 40 ms | 110 ms | 190 ms | 0% | 5 | 34% |
| 150 | 150 | 44 ms | 140 ms | 260 ms | 0% | 8 | 49% |
| 175 | 174 | 52 ms | 210 ms | 480 ms | 0% | 9 | 55% |
| 200 | 188 | 140 ms | 900 ms | 2,400 ms | 0.4% | 9 | 58% |
| 225 | 181 | 1,100 ms | 4,800 ms | 9,000 ms | 6% | 9 | 57% |
Reading it: the knee sits between 150 and 175 rps, where p99 nearly doubles while p50 barely moves. At 200 the gap between offered and achieved load opens, the queue starts growing, and the cliff follows at 225, where throughput actually falls as load increases. CPU never passes 60%, so this is not a CPU problem, and scaling on CPU would never have fired. The telling column is database connections, pinned at 9: one per sync worker, all busy. pg_stat_statements from the run shows the top query by total time is a per-product stock lookup called 24 times per product page. It is an N+1 hidden inside a template loop.
The team fixes it with prefetch_related, adds an index on the stock table's foreign key and product-status filter, and re-runs the identical profile. The knee moves to about 300 rps and p99 at 150 rps drops to 140 ms. Only then do they redo the sizing arithmetic. They also record the scaling signal the run revealed. Worker saturation, meaning busy workers as a fraction of the total, crossed 80% one full step before p99 broke, while CPU stayed flat. That makes worker saturation the autoscaling metric, with a threshold at 70%.
Most confusing results fall into a few recognisable patterns:
| Symptom | Likely cause | How to confirm |
|---|---|---|
| Throughput plateaus while server CPU and DB look idle | Load generator saturated, or a closed model backing off | Check generator CPU and network; compare offered vs achieved rate; watch k6 dropped_iterations |
| Client latency rises, Django timings flat | Requests queueing before reaching workers | Compare Nginx $request_time with $upstream_response_time and your in-app timings |
| Sudden wall of 502/504 errors at a precise load | Worker or proxy timeouts, pool exhaustion, connection limit | Proxy error log, pg_stat_activity count vs max_connections |
| Everything is suspiciously fast | Cache hits on identical inputs, empty tables, error pages returned as 200 | Randomize inputs; check row counts; assert on page content |
Load-test before any major launch or expected traffic event, after significant architectural changes, and continuously in a lightweight form as a regression guard. You do not need a full stress campaign every week, but you should never discover your capacity for the first time during a real spike. The goal is to always know, with evidence, roughly where your cliff is and to be provisioned comfortably below it — so growth is a planning decision, not an emergency.