Syntra Systems
Cases Services Products About Blog IT Caravan
+998 70 010 68 44 +7 999 900 22 12
RusEngUzb
Rows of identical lilac and blue cubes stacked in steps, an image of a platform that grows by adding identical blocks
Websites and platforms

Scalable web platform architecture: monolith or microservices

By Nikita Zhulin · · 9 min read · updated

Scalable architecture is a way of building a platform in which growing load is met by adding resources rather than by rewriting the code. For most projects that means a modular monolith with clear boundaries between its parts, a database designed for growing volume, state kept outside the application, and monitoring. Microservices are not for everyone: they solve the problem of independent teams and releases, not load in itself.

Key takeaways

  • Scalable architecture meets growing load by adding resources rather than by rewriting the code.
  • Horizontal growth only works if the application is stateless: sessions and files are moved to external storage.
  • Most projects are served by a modular monolith; microservices solve the problem of independent teams and releases, not load.
  • The order to add things: indexes and queries, cache, CDN, queues, read replicas, and only then splitting data across nodes.
  • A cloud in a data center in Uzbekistan is enough for compliance: localisation is mandatory only for certain categories of data.

What scalability is and how vertical growth differs from horizontal

Scalability is the ability of a system to absorb growing load without being rebuilt: there are more users, the platform answers just as quickly, it simply uses more resources. There are two routes, and they do not replace each other.

That caveat about state is not a formality. While user sessions, uploaded files and counters sit in the memory of one server, a second server will not help: some requests will land in the wrong place. Beyond that, every system keeps work that cannot be split across machines — writing to a single database table, for instance. That share sets the real limit of horizontal growth.

Why the application has to be stateless

A stateless application keeps nothing between requests that cannot be lost along with the server. The Twelve-Factor App methodology puts it plainly: application processes are stateless and share nothing, and any data that must persist lives in an external backing service, usually a database. Tying a user to one particular server, known as sticky sessions, is called a violation of the principle.

In practice that means sessions in a shared store such as Redis, uploaded files in object storage rather than on the application server's disk, and long operations in a task queue. Then any server can handle any request, and a failed one is replaced by a new one without data loss.

Important You can test this before any load arrives: run two copies of the application behind a balancer and work with both. If a user is logged out when switching between copies, and an uploaded file is visible on only one of them, state still lives inside the application.

Monolith, modular monolith or microservices

A monolith is a platform built as a single application: catalogue, orders, payments and notifications live in shared code and a shared database. For a first release that is normal and usually the right call. The difference between the three approaches is not fashion but the price you pay for independence between the parts.

ApproachWhen it fitsWhat it givesWhat you risk
MonolithFirst version of a product, one team, load still unknownFast start, simple debugging, one database and one deploymentOver time the parts grow together and any change touches its neighbours
Modular monolithProduct is growing, several groups in the team, load is predictableBoundaries between modules are set in advance, any module can be split out laterBoundaries need discipline, otherwise it turns into an ordinary monolith
MicroservicesSeveral teams with different release cycles, parts with very different loadIndependent updates, targeted scaling of the overloaded partsA network between services, distributed failures, separate infrastructure and monitoring

The line that a failure in one service does not bring down the whole system only holds with failure isolation: timeouts on every external call, a limit on retries, and a fallback for when a neighbour is unavailable. Without that, microservices produce cascading failures — a slow service holds the connections of the others and stops the whole chain.

The sensible route for most projects is a modular monolith with clean boundaries: modules talk through an API, the programming interface for exchanging data, and any of them can later be split into a separate service without rewriting the platform. How the server side works and how such interfaces are designed is covered in our article on the backend and API for a mobile app.

What to add as load grows, and in what order

The usual order is: indexes and query tuning, caching, serving static files through a CDN, queues, and read replicas of the database. Start with the cheapest measure — each of the following is harder to run than the one before, and switching them all on at once means paying for what you are not using yet.

  1. Slow queries and indexes. The first bottleneck is almost always in the database. A slow query log shows which ones eat the time, and missing indexes or oversized result sets are fixed without touching the architecture.
  2. Cache. Results of frequent, rarely changing queries are kept in memory, for example in Redis. The hard question is not how to store them but when to drop them: the expiry rule is designed together with the business logic, otherwise users see stale data.
  3. Serving static files through a CDN. Images, scripts and stylesheets are delivered by a network of nodes closer to the user. That takes most of the requests off the application server and speeds up page loads.
  4. Queues. Sending email, generating documents, processing files and building exports move out of the user's request into background workers. The user gets an answer immediately while the heavy work runs separately and survives spikes.
  5. Read replicas of the database. The PostgreSQL documentation describes a hot standby — a server that accepts connections and serves read-only queries. Reports and analytics are moved to such servers so that they do not interfere with order processing.
  6. Splitting data across nodes. Sharding is left as a last resort: it complicates queries, backups and upgrades. You move to it once the earlier steps are done and a single database can no longer cope.

Autoscaling in the cloud adds and removes servers according to current load. It helps with sharp peaks but does not replace these steps: if the bottleneck is the database, extra application servers will only lengthen the queue to it.

How to tell that the platform is running out of headroom

From the system side, four measures matter. The Google SRE book calls them the golden signals: latency (how long a request takes to service), traffic (how much demand arrives), errors (the rate of failed requests) and saturation (how full the most constrained resources are). Latency rising while traffic stays flat is the first sign that headroom is running out.

From the user side, speed is measured by Core Web Vitals. Google sets the thresholds at which a metric counts as good and evaluates the result at the 75th percentile of page loads, separately for phones and desktops.

Load testing is done before the peak, not after it. Scenarios are taken from real work — registration, search, checkout — load is raised gradually, and you record at what number of concurrent users latency climbs and errors appear. That number becomes a known limit rather than a surprise on the day of a campaign.

Where to host the platform in Uzbekistan and what the law requires

You do not need your own server to comply with the law: a cloud in a data center inside the country is enough. Since 27 March 2026 only biometric and genetic data and data of telecom operators' users must be stored in Uzbekistan — that is set by Article 27-1 of the law on personal data as amended by Law No. ZRU-1125 of 26 March 2026.

Important Other personal data may be stored abroad under one of the conditions of the same article. The conditions, the list of countries and latency to foreign data centers are covered in our article on cloud or server in Uzbekistan.

Your own servers make sense with constant high load, where renting the same capacity over several years costs more than buying it, or where hardware has to be physically isolated. A platform with changing load is easier to run in the cloud: the configuration changes without a purchase and a delivery wait.

What infrastructure costs and what makes the bill grow

The price depends on how much capacity the platform actually needs and on the services around it: backups, monitoring, DDoS protection and administration. On our servers and cloud page the scope of work is set out for three tasks.

The bill grows not with the number of registered users but with what they do: the size of the database and backups, outbound traffic, heavy reports and background file processing. Decisions made at the start add less to the budget than reworking the architecture under load, when the platform is already live and cannot be stopped.

Signs of an architecture ready to grow

Readiness for growth is measured not by the size of the server but by what can be done without rewriting code. Go through the list before the load arrives.

Outside the list there is operations: who is on call during an outage, how quickly they respond and who updates the system. Without that even a sound architecture will not save you — how that work is organised is covered in our article on technical support for IT infrastructure.

How we design platforms

Syntra Systems sizes the load before launch: how many concurrent users are expected, how much data accumulates in a year, when the peaks fall and which operations are the heaviest. We choose the configuration from those numbers, decide what to move into background processes, and after launch review it against the actual figures.

Example At the UzFranchise Expo trade show the load was at the entrance: 3,539 QR tickets were scanned over two days, and checking one guest took 0.2 seconds. Load like that does not arrive evenly, so headroom is calculated from the peak minutes rather than from the daily average. The details are in the trade show case study.

We build platforms as modules with clean boundaries, keep state outside the application and move reporting and file processing into separate processes. We host them in Syntra Cloud, in a data center in Uzbekistan, with daily backups and availability monitoring.

Let’s discuss your project

Tell us what you need, and we will estimate the timeline and cost and suggest a solution.

Discuss platform architecture

Frequently asked questions

Can a system be built from the start so growth doesn't mean rewriting it from scratch?

A system can be built this way if it's split from the start into independent blocks, such as orders, warehouse and payments. Each block works with its own data and talks to its neighbours only through interfaces defined in advance, so a change in one block doesn't break the others, and an overloaded block can be moved to its own server without rewriting the whole system. This approach is called a modular monolith.

When is it time to split a module into a separate service?

When the module gets its own team with its own release schedule, or when its load differs sharply from the rest of the platform and it has to be scaled on its own. Splitting for architectural tidiness alone, without one of those reasons, usually costs more than leaving things as they are.

What should we do if the platform is already slow under load?

Measure first, rewrite later. A slow query log and latency metrics show exactly where time is lost: in the database, in an external service or in serving files. In most cases the first improvements come from indexes, caching and moving heavy operations into the background.

Should we plan for Kubernetes and containers from the start?

Containers are useful for almost everyone: they make a launch identical on any server. Kubernetes adds orchestration of many services along with its own operational complexity, so it is adopted when there are many services and people to run them.

How often should load testing be run?

Before launch, before a seasonal peak or an advertising campaign, and after major changes to the database or integrations. Between those points monitoring is enough: if latency rises while traffic stays flat, it is worth repeating the test earlier than planned.

How much more expensive is it to design a platform for growth?

The difference at the start is small: it is a matter of order in the code and the data rather than extra technology. The opposite is expensive — reworking the architecture when the platform is live, cannot be stopped, and the data has to be migrated without loss.

Read also