Scalable architecture is a way of building a platform in which growing load is met by adding resources rather than by rewriting the code. For most projects that means a modular monolith with clear boundaries between its parts, a database designed for growing volume, state kept outside the application, and monitoring. Microservices are not for everyone: they solve the problem of independent teams and releases, not load in itself.
Key takeaways
Scalability is the ability of a system to absorb growing load without being rebuilt: there are more users, the platform answers just as quickly, it simply uses more resources. There are two routes, and they do not replace each other.
That caveat about state is not a formality. While user sessions, uploaded files and counters sit in the memory of one server, a second server will not help: some requests will land in the wrong place. Beyond that, every system keeps work that cannot be split across machines — writing to a single database table, for instance. That share sets the real limit of horizontal growth.
A stateless application keeps nothing between requests that cannot be lost along with the server. The Twelve-Factor App methodology puts it plainly: application processes are stateless and share nothing, and any data that must persist lives in an external backing service, usually a database. Tying a user to one particular server, known as sticky sessions, is called a violation of the principle.
In practice that means sessions in a shared store such as Redis, uploaded files in object storage rather than on the application server's disk, and long operations in a task queue. Then any server can handle any request, and a failed one is replaced by a new one without data loss.
Important You can test this before any load arrives: run two copies of the application behind a balancer and work with both. If a user is logged out when switching between copies, and an uploaded file is visible on only one of them, state still lives inside the application.
A monolith is a platform built as a single application: catalogue, orders, payments and notifications live in shared code and a shared database. For a first release that is normal and usually the right call. The difference between the three approaches is not fashion but the price you pay for independence between the parts.
| Approach | When it fits | What it gives | What you risk |
|---|---|---|---|
| Monolith | First version of a product, one team, load still unknown | Fast start, simple debugging, one database and one deployment | Over time the parts grow together and any change touches its neighbours |
| Modular monolith | Product is growing, several groups in the team, load is predictable | Boundaries between modules are set in advance, any module can be split out later | Boundaries need discipline, otherwise it turns into an ordinary monolith |
| Microservices | Several teams with different release cycles, parts with very different load | Independent updates, targeted scaling of the overloaded parts | A network between services, distributed failures, separate infrastructure and monitoring |
The line that a failure in one service does not bring down the whole system only holds with failure isolation: timeouts on every external call, a limit on retries, and a fallback for when a neighbour is unavailable. Without that, microservices produce cascading failures — a slow service holds the connections of the others and stops the whole chain.
The sensible route for most projects is a modular monolith with clean boundaries: modules talk through an API, the programming interface for exchanging data, and any of them can later be split into a separate service without rewriting the platform. How the server side works and how such interfaces are designed is covered in our article on the backend and API for a mobile app.
The usual order is: indexes and query tuning, caching, serving static files through a CDN, queues, and read replicas of the database. Start with the cheapest measure — each of the following is harder to run than the one before, and switching them all on at once means paying for what you are not using yet.
Autoscaling in the cloud adds and removes servers according to current load. It helps with sharp peaks but does not replace these steps: if the bottleneck is the database, extra application servers will only lengthen the queue to it.
From the system side, four measures matter. The Google SRE book calls them the golden signals: latency (how long a request takes to service), traffic (how much demand arrives), errors (the rate of failed requests) and saturation (how full the most constrained resources are). Latency rising while traffic stays flat is the first sign that headroom is running out.
From the user side, speed is measured by Core Web Vitals. Google sets the thresholds at which a metric counts as good and evaluates the result at the 75th percentile of page loads, separately for phones and desktops.
Load testing is done before the peak, not after it. Scenarios are taken from real work — registration, search, checkout — load is raised gradually, and you record at what number of concurrent users latency climbs and errors appear. That number becomes a known limit rather than a surprise on the day of a campaign.
You do not need your own server to comply with the law: a cloud in a data center inside the country is enough. Since 27 March 2026 only biometric and genetic data and data of telecom operators' users must be stored in Uzbekistan — that is set by Article 27-1 of the law on personal data as amended by Law No. ZRU-1125 of 26 March 2026.
Important Other personal data may be stored abroad under one of the conditions of the same article. The conditions, the list of countries and latency to foreign data centers are covered in our article on cloud or server in Uzbekistan.
Your own servers make sense with constant high load, where renting the same capacity over several years costs more than buying it, or where hardware has to be physically isolated. A platform with changing load is easier to run in the cloud: the configuration changes without a purchase and a delivery wait.
The price depends on how much capacity the platform actually needs and on the services around it: backups, monitoring, DDoS protection and administration. On our servers and cloud page the scope of work is set out for three tasks.
The bill grows not with the number of registered users but with what they do: the size of the database and backups, outbound traffic, heavy reports and background file processing. Decisions made at the start add less to the budget than reworking the architecture under load, when the platform is already live and cannot be stopped.
Readiness for growth is measured not by the size of the server but by what can be done without rewriting code. Go through the list before the load arrives.
Outside the list there is operations: who is on call during an outage, how quickly they respond and who updates the system. Without that even a sound architecture will not save you — how that work is organised is covered in our article on technical support for IT infrastructure.
Syntra Systems sizes the load before launch: how many concurrent users are expected, how much data accumulates in a year, when the peaks fall and which operations are the heaviest. We choose the configuration from those numbers, decide what to move into background processes, and after launch review it against the actual figures.
Example At the UzFranchise Expo trade show the load was at the entrance: 3,539 QR tickets were scanned over two days, and checking one guest took 0.2 seconds. Load like that does not arrive evenly, so headroom is calculated from the peak minutes rather than from the daily average. The details are in the trade show case study.
We build platforms as modules with clean boundaries, keep state outside the application and move reporting and file processing into separate processes. We host them in Syntra Cloud, in a data center in Uzbekistan, with daily backups and availability monitoring.
Let’s discuss your project
Tell us what you need, and we will estimate the timeline and cost and suggest a solution.
A system can be built this way if it's split from the start into independent blocks, such as orders, warehouse and payments. Each block works with its own data and talks to its neighbours only through interfaces defined in advance, so a change in one block doesn't break the others, and an overloaded block can be moved to its own server without rewriting the whole system. This approach is called a modular monolith.
When the module gets its own team with its own release schedule, or when its load differs sharply from the rest of the platform and it has to be scaled on its own. Splitting for architectural tidiness alone, without one of those reasons, usually costs more than leaving things as they are.
Measure first, rewrite later. A slow query log and latency metrics show exactly where time is lost: in the database, in an external service or in serving files. In most cases the first improvements come from indexes, caching and moving heavy operations into the background.
Containers are useful for almost everyone: they make a launch identical on any server. Kubernetes adds orchestration of many services along with its own operational complexity, so it is adopted when there are many services and people to run them.
Before launch, before a seasonal peak or an advertising campaign, and after major changes to the database or integrations. Between those points monitoring is enough: if latency rises while traffic stays flat, it is worth repeating the test earlier than planned.
The difference at the start is small: it is a matter of order in the code and the data rather than extra technology. The opposite is expensive — reworking the architecture when the platform is live, cannot be stopped, and the data has to be migrated without loss.