Australian owned · Operating since 2014 · Sydney, NSW Support & SLAs 24×7 incident line

Home/Services/Managed AWS

Service 01 — Managed AWS Operations

We hold the pager.

A complete operations function for your AWS estate, delivered against a written SLA by engineers based in Australia. Not a ticket queue — a team that knows your architecture.

What's included

The whole run function, not the easy half

Monitoring is configured by us, alerts route to our on-call roster first, and no alert is allowed to page a human until there is a written, tested runbook behind it.

  • Monitoring & alerting — CloudWatch, synthetics and log-based detection tuned to your SLOs, not vendor defaults
  • Incident response — 24×7×365 first-line ownership, incident command for P1s, post-incident report within five business days
  • Patch & lifecycle — Systems Manager patch baselines by environment, with maintenance windows you approve
  • Backup & recovery — AWS Backup policy per workload, restore testing on a schedule, documented RTO/RPO
  • Change management — peer-reviewed Terraform, policy checks in CI, approval gates on production
  • Capacity & performance — trend review against agreed SLOs before it becomes an incident
  • Cost visibility — spend reporting bundled in, with active FinOps on Managed and Enterprise plans
pagerduty · incident #4821
02:14:07 P1 rds-prod-au writer unreachable 02:14:19 ACK seacow-oncall (12s) 02:14:31 runbook RB-018 rds-failover opened 02:15:02 customer channel notified #acme-seacow 02:16:44 failover to ap-southeast-2b complete 02:18:10 writes recovered, p95 204 ms 02:31:00 RESOLVED — 17 min total, 0 data loss # post-incident report due 2026-08-14

Alert hygiene comes first

Before we accept the pager we spend two weeks deleting or fixing alerts that fire without a useful action. On a typical estate that removes 60–70% of the volume — and it is the single biggest reason on-call stops hurting.

Service levels

Contractual, measured, and credited when we miss

Response is measured from alert receipt or ticket creation to a human acknowledging and starting work — not to an autoresponder.

SeverityDefinitionResponseUpdate cadenceTarget resolution
P1Production unavailable, data loss risk, or active security incident15 min, 24×7Every 30 min4 hours
P2Major degradation with a workaround, or single-AZ redundancy lost30 min, 24×7Every 2 hours1 business day
P3Minor fault, limited user impact, non-urgent change4 business hoursDaily5 business days
P4Service request, access change or question1 business dayOn changeBy agreement
15 minP1 response target, 24×7, contractual
30 minP2 response target, 24×7, contractual
5 dayspost-incident report after every P1
Creditsapplied automatically when we miss

These are commitments, not averages. We deliberately do not publish measured performance figures we cannot show you the workings for — ask in a first meeting and we will walk through our actual SLA reporting for a comparable customer.

Onboarding

Three weeks to pager transfer

We do not accept the pager until the runbooks exist. Taking on-call for a system you do not understand is how MSPs end up escalating everything back to you at 2am.

Our full engagement model
  • Week 0

    Discovery

    Read-only access, automated inventory, architecture walkthrough with your team, and agreement on what is in scope and what is not.

  • Week 1

    Alert hygiene & monitoring

    Existing alerts audited: each is fixed, deleted or given a runbook. New coverage deployed for the gaps.

  • Week 2

    Runbooks & access

    Runbooks written jointly with your engineers and tested in non-production. Cross-account roles, MFA and session logging established.

  • Week 3

    Shadow then transfer

    We run alongside your roster for a week, then take first-line paging. Your team stays as escalation, not front line.

Operating rhythm

What working with us actually feels like

DAILY

Shared channel

A Slack or Teams channel with our engineers in it. Questions get answered by the person who built the thing.

WEEKLY

Operations check-in

Thirty minutes: open incidents, upcoming changes, capacity concerns, and anything about to consume your included change capacity.

MONTHLY

Service review

SLA performance, every incident and its cause, spend trend, risk register and the next three optimisations we recommend.

ANNUAL

Review & DR test

A full Well-Architected review and a live disaster recovery exercise with a written result. Semi-annual on Enterprise.

“We went from three unplanned outages a quarter to none. The monthly review is the first vendor meeting I don't dread.”
Head of EngineeringNational logistics operator

Related case study

The customer moved onto Managed at go-live after an eleven-wave data centre exit — and has not had a P1 in the eight months since.

Read the case study
Access & security

Least privilege, fully auditable, revocable in one action

We never ask for IAM users or long-lived keys. Access is by cross-account role assumed through our identity provider, and every action we take appears in your own CloudTrail under a named engineer.

  • Cross-account IAM roles with an external ID, scoped by permission boundary
  • MFA enforced at our IdP; no shared credentials, ever
  • Just-in-time elevation for privileged actions, time-boxed and logged
  • Background-checked delivery staff under confidentiality agreements
  • You can revoke the trust relationship at any time without contacting us

What we will not do

  • Hold your infrastructure code in our own repositories
  • Deploy an agent we will not let your security team review
  • Make an unreviewed production change outside an incident
  • Mark up your AWS consumption

All four are written into the services agreement.

FAQ

Managed AWS questions

Can you work alongside our internal platform team?

That is the most common arrangement. We usually take 24×7 first-line operations and the undifferentiated heavy lifting; your team keeps the product-specific work. The boundary gets written down during onboarding so nobody assumes the other side has it.

What if an incident is in our application code?

We diagnose, contain and escalate to you with the evidence attached — logs, traces, timeline and what we ruled out. We do not sit on an incident waiting for business hours because the cause turned out to be above the platform.

Do you require specific tooling?

No. We work with CloudWatch natively, and equally with Datadog, New Relic, Grafana or Splunk if you already have them. We will tell you honestly when a tool is costing you more than it returns.

Is on-call handled offshore?

No. The after-hours roster is staffed from Australia and New Zealand. Every engineer on it has worked on your account during business hours.

What happens if we want to leave?

Thirty days' notice after the initial term, plus a documented handover: infrastructure code, runbooks, dashboards, a credential rotation plan and a walkthrough session. All of it lives in your accounts throughout the engagement anyway.

See what your estate looks like to us

The free Well-Architected review covers exactly the ground we would be taking over. No cost, no obligation.