Status

All systems operational.

All five components are meeting their targets right now: the API, provisioning, payments, coverage data and support. If that changes, this page moves within five minutes of the incident being declared, and if we can tell who is affected we will have contacted you before you read it here.

A status page is only worth anything if it is willing to say something bad. This one is updated during an incident rather than after it, which means you will sometimes see the word investigating next to a component while we still do not know why.

Components

ComponentStateTargetRight now
APIOperational99.9 percent availabilityNo errors above baseline in the last 24 hours
ProvisioningOperational99.5 percent of paid orders reach a working profileThe rolling hour is above target, median install 41 seconds
PaymentsOperationalAuthorisation and automatic refund inside 60 secondsThe processor is reporting normal and refunds are firing on schedule
Coverage dataOperationalPublished figures refreshed daily, the 25 sample floor heldThe last refresh completed and no country is suppressed for staleness
SupportOperationalFirst response under 60 seconds, 24 hours a dayThe queue is inside target on every staffed language

Operational means the component is meeting its target over a rolling hour. It is not a promise about the next hour.

99.5%

provisioning success target

60 s

automatic refund when provisioning fails

30 min

incident length that triggers a public post mortem

72 h

deadline to publish that post mortem

What each state means

StateWhat it meansWhat we do
OperationalThe component is meeting its target over a rolling hourNothing. This is the normal state and it is not a promise about the future
DegradedWorking, but slower or less reliable than the targetNamed on this page, affected customers contacted, credits applied without a request
Partial outageFailing for an identifiable group, such as one carrier or one regionThe group is named, purchases are blocked where they would fail, refunds fire automatically
Major outageFailing broadlyAn incident commander is appointed and updates land at least every 30 minutes until it is resolved
MaintenancePlanned work with a known windowAnnounced at least 72 hours ahead and scheduled against the lowest traffic hour

There is no state that means fine, probably. A component is either meeting its target or it is named on this page.

The incident policy

Two commitments carry the whole thing, and both are uncomfortable on purpose. The first is that we tell affected customers before they notice. The second is that anything over 30 minutes gets a written post mortem published within 72 hours.

The post mortem names systems and decisions rather than individual employees, and it is written by the person who held the pager rather than by a manager describing somebody else.

If we are going to miss the 72 hour deadline, we publish the delay and the reason inside the 72 hours, because a missed deadline announced late is two failures.

  • Detection is automated from provisioning success and payment authorisation rates, so an incident starts when the numbers move rather than when somebody complains
  • The component state on this page changes within five minutes of the declaration, before the cause is known
  • An incident commander runs the response and a separate person handles communications, so neither job starves the other
  • Affected customers are identified and contacted, with a credit already applied where service was degraded, rather than offered on request
  • Purchases are blocked in any path where we know they would fail, because taking money we will have to refund is worse than losing the sale
  • Updates land at least every 30 minutes during a major outage, even when the update is that we still do not know
  • Within 72 hours of resolution, a post mortem is published with a timeline, the customer impact in numbers, the cause, and every fix with an owner and a date
Your phone
Carrier Acongested, 4 Mbps
Carrier Bclear, 61 Mbps
Moved, without asking you
A single carrier provider cannot make this move. It is the whole reason we buy from more than one.
Where a territory has more than one partner network, congestion moves you rather than becoming an incident.

Why the target is not one hundred percent

Because a network is involved and one hundred percent would be a lie. A 99.5 percent provisioning target says out loud that roughly five orders in a thousand will fail somewhere between the payment and the working profile.

What matters is that the path for those five is designed. The refund fires inside 60 seconds without a ticket, the message explains what happened and what to try instead, and nobody has to decide whether the customer deserves it.

The same logic applies to coverage. Where a partner network is congested in a territory we serve with more than one carrier, we move you rather than logging an incident, and the note says which move was made.

Where there is only one network and it is having a bad day, we say so on the country page. Our coverage pages publish the median, the slowest tenth and the sample count for exactly this reason.

Incident history

No incident over 30 minutes has been declared in the current reporting period.

That sentence is worth exactly as much as our willingness to change it, which is why the policy above is written down in this much detail. Every past post mortem stays published permanently rather than ageing out after ninety days.

Questions people actually ask

What does operational actually mean here?
That the component is meeting its target right now, not that nothing has ever gone wrong. Provisioning is operational when paid orders are reaching working profiles at or above the 99.5 percent target over a rolling hour.
How quickly does this page update during an incident?
The component state moves within five minutes of an incident being declared, before we know the cause. That means you will sometimes read investigating rather than an explanation. We would rather be early and vague than late and tidy.
Will you contact me, or do I have to watch this page?
We contact you. If we can identify who is affected, you hear from us before you notice, with any credit already applied. This page exists for everyone else and for the people who want the detail.
When do you publish a post mortem?
Within 72 hours for any incident lasting more than 30 minutes. It carries a timeline, the customer impact in numbers, the cause, and the fixes with owners and dates. If we are going to be late publishing it, we publish the delay and the reason on time.
What if I cannot load this page because I have no data?
That is the case this whole product is built around. Every account carries 100 MB a month of free rescue data in 145 countries, so you can reach this page, reach support and buy a plan even when your balance is empty. Support answers in under 60 seconds, at any hour.
Do carrier problems in one country show up here?
Yes, as a degraded coverage data component with the country named. Where our partners have more than one network in a territory we move you to the clearer carrier without asking, and the incident note says which move was made.