Status
All systems operational.
A status page is only worth anything if it is willing to say something bad. This one is updated during an incident rather than after it, which means you will sometimes see the word investigating next to a component while we still do not know why.
Components
| Component | State | Target | Right now |
|---|---|---|---|
| API | Operational | 99.9 percent availability | No errors above baseline in the last 24 hours |
| Provisioning | Operational | 99.5 percent of paid orders reach a working profile | The rolling hour is above target, median install 41 seconds |
| Payments | Operational | Authorisation and automatic refund inside 60 seconds | The processor is reporting normal and refunds are firing on schedule |
| Coverage data | Operational | Published figures refreshed daily, the 25 sample floor held | The last refresh completed and no country is suppressed for staleness |
| Support | Operational | First response under 60 seconds, 24 hours a day | The queue is inside target on every staffed language |
Operational means the component is meeting its target over a rolling hour. It is not a promise about the next hour.
99.5%
provisioning success target
60 s
automatic refund when provisioning fails
30 min
incident length that triggers a public post mortem
72 h
deadline to publish that post mortem
What each state means
| State | What it means | What we do |
|---|---|---|
| Operational | The component is meeting its target over a rolling hour | Nothing. This is the normal state and it is not a promise about the future |
| Degraded | Working, but slower or less reliable than the target | Named on this page, affected customers contacted, credits applied without a request |
| Partial outage | Failing for an identifiable group, such as one carrier or one region | The group is named, purchases are blocked where they would fail, refunds fire automatically |
| Major outage | Failing broadly | An incident commander is appointed and updates land at least every 30 minutes until it is resolved |
| Maintenance | Planned work with a known window | Announced at least 72 hours ahead and scheduled against the lowest traffic hour |
There is no state that means fine, probably. A component is either meeting its target or it is named on this page.
The incident policy
Two commitments carry the whole thing, and both are uncomfortable on purpose. The first is that we tell affected customers before they notice. The second is that anything over 30 minutes gets a written post mortem published within 72 hours.
The post mortem names systems and decisions rather than individual employees, and it is written by the person who held the pager rather than by a manager describing somebody else.
If we are going to miss the 72 hour deadline, we publish the delay and the reason inside the 72 hours, because a missed deadline announced late is two failures.
- Detection is automated from provisioning success and payment authorisation rates, so an incident starts when the numbers move rather than when somebody complains
- The component state on this page changes within five minutes of the declaration, before the cause is known
- An incident commander runs the response and a separate person handles communications, so neither job starves the other
- Affected customers are identified and contacted, with a credit already applied where service was degraded, rather than offered on request
- Purchases are blocked in any path where we know they would fail, because taking money we will have to refund is worse than losing the sale
- Updates land at least every 30 minutes during a major outage, even when the update is that we still do not know
- Within 72 hours of resolution, a post mortem is published with a timeline, the customer impact in numbers, the cause, and every fix with an owner and a date
Why the target is not one hundred percent
Because a network is involved and one hundred percent would be a lie. A 99.5 percent provisioning target says out loud that roughly five orders in a thousand will fail somewhere between the payment and the working profile.
What matters is that the path for those five is designed. The refund fires inside 60 seconds without a ticket, the message explains what happened and what to try instead, and nobody has to decide whether the customer deserves it.
The same logic applies to coverage. Where a partner network is congested in a territory we serve with more than one carrier, we move you rather than logging an incident, and the note says which move was made.
Where there is only one network and it is having a bad day, we say so on the country page. Our coverage pages publish the median, the slowest tenth and the sample count for exactly this reason.
Incident history
No incident over 30 minutes has been declared in the current reporting period.
That sentence is worth exactly as much as our willingness to change it, which is why the policy above is written down in this much detail. Every past post mortem stays published permanently rather than ageing out after ninety days.