Microservices Tradeoffs and Design Patterns

Letting aside the reasons why we should or should not jump into Microservices from previous post, here we will talk more about the trade-offs of Microservices and the design patterns that are created to address them.

Building microservices is not as easy as installing some packages into your current system. In fact, you will need to install a lot of components. The beauty of microservices lies in the separation of services that enables each module to be developed independently and keeps each module simple. However, that separation also causes new challenges.

1. More I/O operations

The first issue that we can easily recognize is the emergence of I/O calls between separate services. It resembles the integration of our system with third-party services; however, this time, all those third-party services are actually our internal ones. To ensure correct API calls, there will be efforts to document and synchronize knowledge between the teams handling different services.

But here is the bigger problem: if every service has to keep a list of other services’ addresses (to call their APIs), they become tightly coupled, meaning there is a strong dependence between them, which destroys the promised scalability of microservices. This is when the event-driven style comes to the rescue.

1.1. Event Driven Design Pattern

Example tools: RabbitMQ, ActiveMQ, Apache Kafka, Apache Pulsar, and others.

The main idea of this pattern is to allow services to operate without knowing each other’s addresses. Each service just needs to know of an event pipe or a message broker and entrust it with distributing its messages and feeding back data from other services. There will be no direct API calls between services; each service only fires events to the pipe and listens for events coming from the pipe.

Along with this design pattern, the mindset on how to store data requires some escalation as well. We will not only store the state of entities but also the stream of events that construct that state. This storage strategy is also very effective when dealing with concurrent modifications of the same entity that can cause inconsistencies in the data. There are two approaches to storing and consuming events: using a queue and using a log, which we will discover in later topics.

2. More Complex Query Mechanism

It is obvious that there will be moments when we need to query some data that requires cooperation between multiple services. In the past, with a monolithic style, when all data from all services was located in the same database, writing an SQL query was simple. However, in a microservices architecture, this is not the case. Each service secures its own database as a recommended practice. We suddenly can’t JOIN tables, we lose the out-of-the-box rollback mechanism from the database’s transaction feature in case something goes wrong with storing data, and we may experience longer delays while each service waits for data from other services. These obstacles make an event-driven approach a “must-have” design for microservices systems, as this architecture forms the foundation to support patterns that solve this querying issue, with the most common being event sourcing, CQRS, and Saga.

2.1. Event Sourcing

It can be a bit confusing to distinguish between the terms Event Driven and Event Sourcing. Event Driven refers to the communication mechanism between services, while Event Sourcing focuses on the coding solution within each service to retrieve the state of an entity; instead of fetching the entity from the database, we reconstruct it from an event stream. The event stream can be stored in various ways: it can be stored in a database table, or it can be read from Event-Driven supported components such as Apache Kafka or RabbitMQ, or by using dedicated event stream databases like EventStore, etc. This method brings new responsibilities for developers, as they must create and maintain the reconstruction algorithms for each type of entity.

2.2 CQRS (Command and Query Responsibility Segregation)

As mentioned in the previous section, this strategy is helpful when dealing with concurrent data modification scenarios, such as collaboration features found in Google Docs or Google Sheets, or simply to address situations where two users hit “Save” on the same form in very close timeframes. However, this reconstruction method is not as friendly to more complex queries, which are typical in traditional databases like Oracle or PostgreSQL, such as the SELECT * WHERE queries. To address these drawbacks, each service usually also maintains a traditional database to store the states of entities and uses it for querying. This combination forms a new pattern called CQRS (Command and Query Responsibility Segregation), where the read and write operations on an entity occur in different databases.

As mentioned above, this pattern separates read and update operations for a data store. A service can use the Event Sourcing technique to update an entity or construct an in-memory database, such as H2, to quickly store updates on entities while persisting the calculated states of those entities back to a SQL database, for example, as quickly as possible. This pattern prevents data conflicts when multiple updates occur on a single entity simultaneously while also maintaining a flexible interface for querying data.

This pattern is effective for scaling purposes since we can scale the read database and the write database independently, and it is suitable for high-load scenarios when write requests can complete more quickly because it reduces calls to the database and potential delays from locking mechanisms. Quicker responses mean there will be more room for other requests, especially in thread-based server technologies such as Servlets or Spring.

A drawback of this pattern is its coding complexity. As more components join the process, there will be more problems to handle. Therefore, it is not recommended to use this approach in cases where the domain or business logic is simple. Simple features fit nicely with the traditional CRUD method. Overusing anything is not advisable. I also want to remind you that if the whole system does not have special needs regarding load or write-heavy features, it is not advisable to switch to microservices as well. (The reason is here)

2.3. Saga

Saga means a long heroic story. The story about transactions inside microservices is truly heroic and lengthy. A transaction is an important feature for a database that aims to maintain data consistency, as it prevents partial failures when updating entities. With distributed services, we are dealing with distributed transactions. Now, the mission is to coordinate those separate transactions to regain the attributes of a single transaction: ACID (atomicity, consistency, isolation, durability) across distributed services. Simply put, a saga is a design pattern aimed at forming transactions for microservices.

Saga patterns address what a system must do in the event of a failure within a service. It should somehow reverse some previously successful operations to maintain data consistency. The simplest way to achieve this is by sending messages to request that certain services roll back specific updates. To create a Saga, developers may need to anticipate many scenarios in which an operation can fail. A higher-level solution for the rollback mechanism is to implement techniques such as semantic locking or versioning entities. We can discuss this in other topics. However, the point here is that it also introduces significant complexities to the source code. The recommendation is to structure services carefully to avoid writing too many Sagas. If certain services are tightly coupled, we should consider merging them back into a single monolithic service. Sagas are less suitable for tightly coupled transactions.

3. More Deployment Effort

Back in the realm of Monoliths, deployment means running a few command lines to build API instances and a client-side application. With Microservices, we obviously have more than one instance, and we need to deploy each instance one by one.

To reduce this effort, we can use CI/CD tools such as Jenkins or various available cloud-based CI/CD options. We can also write our own tools; it won’t be difficult. However, there are still more issues than just running command lines.

3.1. Log Aggregation

Logging is a vital practice when building any kind of application, as it provides insight into how the system is performing and helps troubleshoot issues. Checking logs across separate services can be inconvenient in microservices, so it is recommended to stream logs to a single center. Many tools are dedicated to this purpose nowadays, such as Graylog and Logstash. The most famous stack for collecting, parsing, and visualizing logs is currently the ELK stack, which combines Elasticsearch, Logstash, and Kibana. The drawback of these available logging technologies is that they require a considerable amount of RAM and CPU, primarily to support log searching. For small projects, preparing a machine strong enough to run the ELK stack may not be very affordable. Logstash requires about 1-2 GB, which is generally sufficient. Graylog requires Elasticsearch, so it also needs about 8 GB of RAM and 4 CPU cores. ELK demands much more than that.

3.2. Health Check & Auto restart

Beside logging, we must also have a way to keep track of the availability of services. Each service may have its own API /healthcheck that we can use a tool to periodically call to check whether it’s alive. Alternatively, we can employ proactive monitoring tools such as Monit or Supervisord to monitor ports and processes, and configure their behavior when certain errors occur, such as sending emails or notifications to the Slack channel.

Beside the health check, each service should have an auto-restarting ability when something takes it down. We can configure a process to start up whenever the machine is up by adding scripts to /etc/init.d or /etc/systemd for most Linux servers. For processes, we can make use of Docker to automatically bring services back up right after they go down. For the machine itself, if we use a physical machine, we should enter the BIOS and set it up to auto-restart when power is restored. If we use cloud machines, there is no worry.

Those techniques are not only recommended for microservices but also for any monolithic system to ensure availability.

3.3 Circuit Breaker

This is for when bad things happen and we have no way to deal with them except to accept it. There is always such a situation in life. For some reason, one or more services are down or become so slow due to network issues that it makes users wait a long time just for a button click. Most users are impatient, and they will likely retry the pending action many times, which can worsen the system’s performance. This is when a Circuit Breaker takes action. Its role is similar to that of an electric circuit breaker; it prevents catastrophic cascading failures across the system. The circuit breaker pattern allows you to build a fault-tolerant and resilient system that can survive gracefully when key services are either unavailable or have high latency.

The Circuit Breaker must be placed between the client and the actual servers containing services. The Circuit Breaker has two main states: Closed and Open. The rules governing these states are:

  • In the Closed state, the Circuit Breaker only forwards requests from clients to the backend services.
  • Once Circuit Breaker discovers a failed request or high latency, it changes status to Open.
  • In the Open state, the Circuit Breaker will return errors to client requests immediately, so the user acknowledges the failure, which is better than letting users wait, and it also reduces the load on the system.
  • Periodically, the Circuit Breaker makes a retry call to the backend services to check their availability. If the backend services are functioning again, it changes to the Closed state; if not, it remains in the Open state.

Luckily, we may not have to implement this pattern ourselves. There are tools available out there, such as Hystrix—a part of Netflix OSS—or Istio—the community one.

3.4. Service Discovery

As we mentioned in the Event Driven section, services within a Microservices architecture do not need to know each other’s addresses when using an event channel. But what if the team is not familiar with the event style and decides not to use it, or if the services are simple enough to just expose REST APIs? Using event-driven architecture is not a requirement, and in this case, how do we solve the addressing problem between services?

When a system needs to be scaled, more instances of one or many services need to be added, removed, or simply moved around. To let every service know the addresses (IP, port) of others, we need a middleman that maintains the records of service addresses and keeps them up to date. This module is called Service Discovery and is usually used alongside Load Balancing modules. We may discuss this further in other topics.

We also do not need to create this component from scratch; there are some tools available, such as etcd, Consul, and Apache ZooKeeper. Let’s give them a try.

4. Conclusion

Above is an overview of what we need to know when moving to microservices. Make sure you Google them all before really starting. Each of the patterns will have its pros and cons, along with solutions for overcoming challenges that other topics will cover.

Build – Secure – Evolve with the-tech-lead.com


Leave a Reply