Horizontal Scaling

AI Cost Management

Adding more compute resources (servers/GPUs) in parallel to handle increased load. Horizontal scaling is cost-effective for distributed inference but requires proper load balancing and can increase latency slightly.

← Back to glossary