More machines, not bigger ones
Instead of upgrading the machine, buy more of them. Run several copies of your service on ordinary instances, put a load balancer in front, and capacity grows roughly in line with the number of copies.
That changes the ceiling. The biggest machine in the catalogue is a hard limit. A fleet is not, and nobody has hit a meaningful limit on servers behind a balancer.
Availability is the bigger half
Capacity is only half the sale, and honestly the smaller half. The bigger win is having more than one of something. One machine at 99.9 percent is down almost nine hours a year, all at once, at the worst possible time. Three machines behind a balancer that checks their health keep serving while one dies, deploys, or reboots for patches.
Everything else falls out of that. Rolling deploys, canary releases and maintenance without downtime all need more than one of everything. It is why plenty of teams go horizontal at traffic levels one machine would handle comfortably: they are buying availability, not throughput.
Elasticity arrives with it. Your provider watches a metric, usually processor load or requests per instance, and adds or removes machines against it. Traffic that swings tenfold between 3am and the evening peak stops being a provisioning problem. It becomes a line on your bill, because you pay for peak capacity only during the peak.
The fine print matters, though. Horizontal scaling is something you design for, not a box you tick. It works only when any instance can serve any request, which means no state living on individual machines.
That constraint sounds mild and is not, and the next lesson is a tour of everything it breaks. Note the asymmetry too: stateless application servers scale out trivially and databases do not, which is why replication and sharding get chapters of their own.
Worked example
Hana runs ticket sales for a regional events company on 3 c5.large instances behind an ALB. A boy band announces a reunion show and the on-sale hits 9x normal traffic at exactly 10 am. Her autoscaling group is set to hold average CPU at 55 percent, so between 9:58 and 10:07 it launches 11 more instances, reaching 14. Latency at p95, the figure 95 requests in every 100 come in under, holds around 140 ms through the spike. By 1 pm the group has scaled back to 4. The extra machines cost her about 4 dollars total for the three hours they ran. The previous year, on a single large box, the same on-sale produced 40 minutes of timeouts and a public apology.