Oskar Dudycz

Pragmatycznie o programowaniu

Long-polling, how to make our async API synchronous

2021-11-17 oskar dudyczAPI

cover

I’ll continue today a topic of handling eventual consistency that I started in the previous article. This time let’s learn the trick called “long-polling”. It helps in cheating on the API surface that our operations are synchronous.

Let’s imagine that we’re either have an asynchronous process making changes, or just our database (e.g. MongoDB, ElasticSearch) has built-in eventual consistency. Having that, we cannot be sure if changes were already applied or not. The best was if we had a push notifications mechanism (e.g. WebSockets-based), informing the client application about the end of processing. Then the client could take the needed operation.

However, for various reasons, this might not be a viable solution. Sometimes we don’t have enough experience with such solutions, and our coworkers get defensive. WebSockets are nice but tricky to configure, especially if we’re using load balancers. We may also be making an API-first approach where API is treated as our product. Some of our clients might not want to use push notifications. It’s also a mind shift experience, and UI/UX needs to reflect the asynchronous nature of backend logic.

Too often, we decide to perform retries on our clients until we get satisfying data. I had the case once when client applications were doing DDoS attach on our backend, right in the middle of the important demo for the client. Blind client-side retries are no-go for the production systems.

What to do then? We can consider using long-polling. We’re shifting retries from the clients to the backend side. Instead of being banged by the API retries, we’re waiting until the operation is finished or retrying the calls internally. Thanks to that, we’re getting more control of them. We can easier include logic like backpressure, circuit breakers etc.

How to implement it? I’ll use NodeJS and TypeScript endpoint as an example. In other environments, it can be done accordingly. Let’s say that we’re opening the shopping cart. We’re processing the change and storing the data. We’re using MongoDB, which doesn’t guarantee reading-own-writes by default.

Our handler looks as follows:

export const route = (router: Router) =>
  router.get(
    '/clients/:clientId/shopping-carts/:shoppingCartId',
    async function (request: Request, response: Response, next: NextFunction) {
      try {
        const query = mapRequestToQuery(request);

        const result = await getShoppingCartDetails(query);

        if(result === null) {
          response.sendStatus(404);
          return;
        }

        response.set('ETag', toWeakETag(result.revision));
        response.send(result);
      } catch (error) {
        next(error);
      }
    }
  );

function mapRequestToQuery(
  request: Request
): GetShoppingCartDetails {
  if (!isNotEmptyString(request.params.shoppingCartId)) {
    throw 'Invalid request';
  }

  return {
    shoppingCartId: request.params.shoppingCartId,
  };
}

We’re getting shopping cart id from the URL, getting the result together with ETag to enable optimistic concurrency handling (read more in How to use ETag header for optimistic concurrency). The query handling is using MongoDB api:

export type GetShoppingCartDetails = Query<
  'get-shopping-cart-details',
  {
    shoppingCartId: string;
  }
>;

export async function getShoppingCartDetails(
  query: GetShoppingCartDetails
): Promise<ShoppingCartDetails | null> {
  const collection = await getMongoCollection<ShoppingCartDetails>(
    'shoppingCartDetails'
  );

  return collection.findOne({
    shoppingCartId: query.data.shoppingCartId,
  });
}

So far, so good. But what will happen if the newly opened shopping cart does not exist yet? We’ll get an unexpected 404 status. It’s getting more probable depending on how unlucky we are or how fast we’re making requests.

To do long-polling, we need to retry our query until the result is available. We’ll use for that retryPromise and retryIfNotFound methods introduced in the previous article.

export type RetryOptions = Readonly<{
  maxRetries?: number;
  delay?: number;
  shouldRetry?: (error: any) => boolean;
}>;

export const DEFAULT_RETRY_OPTIONS: Required<RetryOptions> = {
  maxRetries: 5,
  delay: 100,
  shouldRetry: () => true,
};

export async function retryPromise<T = never>(
  callback: () => Promise<T>,
  options: RetryOptions = DEFAULT_RETRY_OPTIONS
): Promise<T> {
  let retryCount = 0;
  const { maxRetries, delay, shouldRetry} = {
    ...DEFAULT_RETRY_OPTIONS,
    ...options,
  };

  do {
    try {
      return await callback();
    } catch (error) {
      if (!shouldRetry(error) || retryCount == maxRetries) {
        console.error(`[retry] Exceeded max retry count, throwing: ${error}`);
        throw error;
      }

      const sleepTime = Math.pow(2, retryCount) * delay + Math.random() * delay;

      console.warn(
        `[retry] Retrying (number: ${
          retryCount + 1
        }, delay: ${sleepTime}): ${error}`
      );

      await sleep(sleepTime);
      retryCount++;
    }
  } while (true);
}

export async function assertFound<T>(() => find: Promise<T | null>): Promise<T> {
  const result = await find();

  if (result === null) {
    throw 'DOCUMENT_NOT_FOUND';
  }

  return result;
}

export function retryIfNotFound<T>(
  find: () => Promise<T | null>,
  options: RetryOptions = DEFAULT_RETRY_OPTIONS
): Promise<T> {
  return retryPromise(() => assertFound(find), options);
}

Method assertFound throws an exception, when the record was not found, to trigger promise rejection and retry made by retryPromise. It uses the recommended by AWS retry policy. It has configurable exponential backoff to increase the delay between retries, a random factor to not spam our database.

We can wrap our query handler with them:

export const route = (router: Router) =>
  router.get(
    '/clients/:clientId/shopping-carts/:shoppingCartId',
    async function (request: Request, response: Response, next: NextFunction) {
      try {
        const query = mapRequestToQuery(request);

        const result = await retryIfNotFound(() =>
          await getShoppingCartDetails(query);
        );

        response.set('ETag', toWeakETag(result.revision));
        response.send(result);
      } catch (error) {
        if(error ===  'DOCUMENT_NOT_FOUND') {
          response.sendStatus(404);
        }
        next(error);
      }
    }
  );

Thanks to that, we can tune the retry options to get the expected result. However, as I explained in Tell, don’t ask! Or, how to keep an eye on boiling milk article, relying on the timing is never okay. There may be the case when we don’t calculate timeout properly, and the document will still be unavailable. We need to be prepared for such a case. We cannot just increase the number of maximum retries. This will keep a hanging connection to our service and may cause connection pool exhaustion. The best is to fail fast and recover. It’s always good to provide a maximum deadline to cut it if it takes longer.

We can do that by adding an additional method that will cut the async call. In JS/TS we can extend Promise as follows:

declare global {
  interface Promise<T> {
    withTimeout(timeout: number): Promise<T>;
  }
}

export type TIMEOUT_ERROR = 'TIMEOUT_ERROR';

Promise.prototype.withTimeout = function <T>(timeout: number) {
  return Promise.race<T>([
    this,
    new Promise(function (_resolve, reject) {
      setTimeout(function () {
        reject('TIMEOUT_ERROR');
      }, timeout);
    }),
  ]);
};

We’re using Promise.race method, which takes an array of promises and finishes immediately when one of the promises succeeds or fail. We’re passing two promises:

  • first one with our async call,
  • second one is with a timer that will just reject promise after the defined timeout.

If our call is fast enough, it’ll just return the result, and the timer promise will be finished. Otherwise, the timer will end our async call.

We’re also extending the general _Promise type by extending its prototype. Thanks to that, we can use it as:

const result = await retryIfNotFound(() =>
    await getShoppingCartDetails(query);
).withTimeout(1000);

Which will kill our retries if they take longer than expected.

Long-polling is a simple technique that shouldn’t be used by default. We should at first consider changing our UI/UX strategy, consider using push notifications. However, sometimes we just have to pragmatically get the job done, and that’s one of the tools that used wisely can help us on that.

Cheers!

Oskar

👋 If you found this article helpful and want to get notification about the next one, subscribe to Architecture Weekly.

✉️ Join over 11500 subscribers, get the best resources to boost your skills, and stay updated with Software Architecture trends!

Loading...
Event-Driven by Oskar Dudycz

cover

Through my window, I see the result of good plans but poor execution. Opposite my flat, there is a partially completed construction place. Buildings were supposed to be eye-catching Mediterranean style apartments. Delivery date? Two years ago. Actual? More and more unknown.

Some time ago, I heard that using Event Sourcing makes creating Event-Driven Architecture easier. The arguments were correct, that if we’re already publishing events to trigger business workflows, then at some point, we may want to also store events to not lose information. Agreed. However, I also heard that keeping the state as events will simplify things. We’ll have a source of truth with a record of the system behaviour. This will allow, e.g. to confront the results of the operations with the recorded state. I’d agree with that, with one distinction. It’s easier as long as you already know Event Sourcing.

Many people in the DDD community claim that the essential is to properly break down the system into autonomous parts called bounded contexts. Once we have it, the rest is secondary and will sort itself out. For sure.

Many seasoned programmers speak similarly about new technologies. They claim that they can translate past experience into new technologies. That’s true that by analogy, they can catch the big picture quicker. But isn’t it a bold assumption to say that Win.Forms specialist will learn Angular quickly?

The end result may differ a lot from the initial ideas. I saw the plan of those buildings next to me. Now I can see the effects of the execution. Or actually, the lack.

I believe that we should carefully acknowledge not only the point of view of our authorities but also their seating point. If we want to find out how to form a wall, do we ask an architect or a foreman? An architect may know the theory, but the practice is what we’re looking for. On the other hand, if you want to know where to put the wall, you prefer the architect to do measurements. At least if you don’t want to have the roof falling to your head.

After I had torn a ligament in my knee, I went to two qualified orthopedists. One said I should have surgery and do a reconstruction. The second stated that there is no need for that; rehabilitation should be enough. Guess which one had a specialization in surgery and which in rehabilitation?

People usually give us advice from the point where they’re currently standing. They are entitled to a biased view. An architect who rarely does programming will tend to downplay the value of implementation and tactical patterns. Midlevel developers will focus on technicalities instead of the global system impact. The team manager or consultant will emphasize the importance of soft skills (or esoteric techniques known only to them).

The truth is that we need all of them. The excellent plan will fall on the bad execution. The best execution for the wrong case will be just a waste of time. We should carefully evaluate the advice considering what we need and what an expert can give us.

Therefore, when we’re reading an article, watching a talk, let’s also pay attention to the place where the person is standing. The perspective from there may be much different from where we are right now. That can be good, as it may push us in the right direction. But it may also be misleading, as we accidentally take biases of this person without understanding the tradeoffs. Personally, I prefer to follow not only people from pedestal but also those that are closer to my position. A bit further in the journey, but not too far. That helps me to calibrate my view as those people are more relative to my daily struggles.

Polish historical leader Józef Piłsudzki reportedly used to say: “Right is like an ass, everyone has its own”.

Cheers!

Oskar