Skip to content
NOSTREL
Reliability

There is no uptime
figure on this page.
Here is why.

We are pre-launch and have processed no live volume. Any percentage we printed would be either a number from a development environment dressed up as production, or an aspiration formatted as a fact, and both are worse than the empty space where the number usually goes.

What we can show is the machinery: which jobs run and how often, which states they resolve, and which alarms exist. A percentage tells you how last quarter went. This tells you what happens on the day something breaks, which is the thing you were actually asking.

Status page

01When it breaks

Four ways this goes wrong, and what each one does

Including what we need from you in each case, which most pages of this kind leave out.

Safaricom is unavailable

Your request is accepted or refused, never silently half-done. A collection we cannot start comes back as an error you can retry with the same idempotency key, which makes retrying safe rather than frightening. Anything already in flight is resolved by the sweeper as soon as the provider can answer again.

Your part

Retry with the same key. Do not generate a new one, because a new key is a new payment.

Your server is down

Deliveries retry on a backoff and nothing is lost while they do, because the event was written with the payment rather than handed to a queue. What exhausts its retries is dead-lettered where you can see it and replay it.

Your part

When you come back, reconcile from the collection records rather than assuming nothing happened while you were away.

A callback never arrives

This is the one that quietly costs money elsewhere, because the obvious reaction, retrying a payout, pays somebody twice. Nothing here is retried blindly. We ask the provider what actually happened and take their answer.

Your part

Trust the record, not the absence of a message. A payment you heard nothing about is not a payment that did not happen.

Two systems disagree

Which they will, occasionally, because that is what distributed systems do. The daily job finds it and raises an exception with a timestamp and a reason, in a queue with a state, rather than leaving a difference for somebody to notice at month end.

Your part

Nothing, usually. If it concerns your account you will hear from us rather than the other way round.

02The machinery

Six jobs, each one chasing a specific way a payment gets stuck

They run on exactly one instance at a time, by lease, so running three copies of the service does not mean reconciling the day three times.

  1. 01

    Drain the outbox

    every 5 seconds

    Events are written in the same transaction as the thing that happened, then delivered from there. So an event cannot be lost by a queue being down, because it was never in a queue when it mattered.

  2. 02

    Retry failed deliveries

    every 15 seconds

    Your endpoint was restarting, or briefly unhappy. Deliveries whose backoff has elapsed go out again.

  3. 03

    Confirm collections

    every minute

    A payment parked in awaiting_confirmation gets asked about directly. Yes credits it, a definite no fails it and returns the fee, and no answer at all leaves it parked for the next pass.

  4. 04

    Reap stale collections

    every 5 minutes

    A prompt nobody answered becomes timed_out rather than sitting open forever. Deliberately does not touch awaiting_confirmation: those got a callback and are waiting on us, not on the payer.

  5. 05

    Sweep stuck payouts

    every 5 minutes

    A payout processing with no result callback gets a transaction status query against the identifier we kept from the original request, which is the one thing we still hold when a callback is lost.

  6. 06

    Reconcile the day

    at 02:00

    Our ledger against the provider record for the same window. Anything that does not line up becomes an exception with a state on it, so a person closes it rather than it ageing quietly.

03Alarms

Eight things wake somebody up

Naming them is the claim. A dashboard screenshot is decoration, because anyone can draw a graph; the question is what is wired to a phone.

  1. The API is down

    Nothing is answering. The one everybody has.

  2. HTTP errors are elevated

    Something is wrong that has not taken the whole service with it, which is usually worse to debug and better to catch early.

  3. The scheduler has stalled

    The jobs above are not running. Nothing looks broken from outside while this is true, which is exactly why it is alarmed.

  4. Reconciliation exceptions are open

    Money that has not been explained is ageing. This is the alarm a finance person would write.

  5. Payout failures are elevated

    Money leaving is failing more than it should, which is either a provider problem or ours, and both need somebody now.

  6. Webhook deliveries are dead-lettered

    Somebody is not receiving their events and may not know it yet.

  7. The outbox is backing up

    Events are being written faster than they are delivered. A leading indicator rather than an outage.

  8. The audit chain is broken

    The tamper-evident log does not verify. This one pages, immediately, whatever the hour.

04Limits

What this does not protect you from

A reliability page that implies total coverage is describing a system that does not exist.

Safaricom being down

If the rail is unavailable, payments do not happen. Nothing architectural changes that, and anyone who implies otherwise is describing a queue, not a payment. What we can do is make sure nothing is lost or double-counted while it lasts, and that it resolves itself when the rail returns.

The payer not answering

The slowest part of every M-Pesa payment is a person finding their phone. That is not latency we can engineer away, and designing a checkout as though it were is the most common cause of a payer paying twice.

A distributed flood

Our edge limits stop one source being rude, cheaply and well. Volumetric defence is a different layer and a different purchase, and we would rather name that than imply a protection we have not bought.

Your own retry logic

We make retrying safe by requiring an idempotency key. We cannot make you reuse it. A fresh key per attempt turns one payment into several, and that is the one failure mode on this page that lives entirely in your code.

When we launch, this page gets a number

A real one, measured, with the method next to it. Until then the machinery is the honest answer and the empty space is deliberate.