DevOps / SRE Engineer

You are not buying Terraform.
You are buying quiet nights.

Every DevOps job description lists the same eight tools. The thing worth paying for is what happens to your on-call rota ninety days later. Our engineers work full-time in your team and stay on the Virtus payroll.

$38–68
hourly band
7–12 days
to day one
1 month
minimum term
24/7
cover available
The on-call week

The same week, ninety days apart

A real client's paging volume, one representative week each side of a single senior SRE engagement. No re-platforming, no new vendor — just the five things in the list below.

Week 0 — before83 pages
Mon
Tue
Wed
Thu
Fri
Sat
Sun
00
03
06
09
12
15
18
21
Week 13 — after7 pages
Mon
Tue
Wed
Thu
Fri
Sat
Sun
00
03
06
09
12
15
18
21
01

Alerts that mean a human is needed

Two-thirds of the pages on the left were informational. An alert nobody acts on is a training exercise in ignoring alerts.

02

Deduplicate and correlate

One database failover was paging four services independently. One incident should produce one page.

03

Fix the top three causes, not all thirty

Pages follow a power law. Three root causes were sixty per cent of the volume, and two of them were a missing retry.

04

Auto-remediate the boring ones

Disk cleanup, pod restarts and certificate renewals do not need a person at 4am. They need a runbook that runs itself.

05

Error budgets, agreed out loud

Once the SLO is written down, “should we ship this Friday” has an answer instead of an argument.

The rest of the role

Six things they own between incidents

Infrastructure as code

Terraform or Bicep, modularised, with a plan reviewed like any other code — and no click-ops drift left unaccounted for.

Pipelines that are boring

Build once, promote the artefact, gate on tests, deploy progressively. Rollback is a button, not an archaeology project.

Kubernetes without the folklore

Resource requests based on measurement, sane probes, PodDisruptionBudgets, and an autoscaler that has actually been tested.

Observability that answers questions

Metrics, logs and traces correlated by one id. Dashboards built around the four golden signals rather than around what was easy to graph.

Cost and capacity

Right-sizing, spot where it is safe, and an owner attached to every line of the bill.

Security in the pipeline

Image scanning, signed artefacts, secrets in a vault and least-privilege service accounts — enforced in CI rather than in a policy document.

Seniority

Three bands

A mid-level engineer keeps the platform running well. Turning the left-hand grid into the right-hand grid is senior work — it is judgement about what deserves a human, not tooling.

Mid-level

3–5 yrs$38–46/hr

Maintains pipelines and infrastructure modules, handles routine incidents, follows the runbook well.

Senior

5–8 yrs$46–58/hr

Owns the platform and the on-call quality. Designs the SLOs and drives the top causes to zero.

Lead / SRE Manager

8+ yrs$58–78/hr

Sets the reliability strategy across teams, runs incident review, and negotiates the error budget with product.

Export last month's alerts and send them over

We will group them, name the top three causes, and tell you how much of your grid is genuinely necessary. That analysis is free.