Skip to content
Glass 3D puzzle pieces on a lime green gradient

What to Hand Off in CI/CD Operations, and What to Keep

OpsWerks
OpsWerks
 
Three glass puzzle pieces fitted togetherA fourth glass puzzle piece lifted out of the set
Operating Models

What to Hand Off in CI/CD Operations, and What to Keep

Three tests to sort your CI/CD queue into what your team keeps, what you agree first, and what can move to someone else.

A build fails at 2 am and the release stops.
A team can’t deploy because a runner is stuck.
A new hire wants to know why their pipeline won’t pick up a secret.

All of it lands on the engineers who own the platform. Your platform team has become the help desk, and the roadmap they were hired to build slips a sprint at a time.

So which CI/CD work should leave your platform team, and which has to stay? Here’s how we sort it. We’ve been supporting enterprise engineering teams since 2015.

Why CI/CD support swallows platform teams

90%
of organizations have adopted platform engineering
38%
say staffing is stable on paper while their scope keeps growing
22%
name platform and developer support bottlenecks their biggest operational challenge; only alert fatigue and incident overload ranks higher

When 90% of organizations have adopted platform engineering, the platform is how most teams ship, and their questions come with it. The most common staffing picture in our survey is flat headcount with growing scope. Platform and developer support bottlenecks are the second most commonly named top challenge, ahead of migrations and legacy systems.

Sources: DevOps Research and Assessment (DORA), 2025 State of AI-assisted Software Development; the most recent OpsWerks State of SRE Survey

Two glass figures Every new team adds to the queue

Every team that ships through your pipelines brings its own questions. The team that owns the platform doesn’t grow when they arrive.

Burst of glass plates Interrupts cost more than the ticket

A quick fix rarely costs only the time it takes. It also costs the focused work it broke into, and that was roadmap work.

Three tilted glass plates Two platforms, one team

The legacy pipeline can’t retire until the last service moves off it, so for months you run two. The engineers building the new platform are the ones paged for the old one.

Glass document Toil never shows up on the roadmap

Support work arrives as tickets and chat messages. It never gets an epic or an estimate, so it stays invisible until it has eaten a deadline.

“Can’t build the new thing because I’m supporting production.”
Andrew, Director of Infrastructure Software

Due to strict enterprise confidentiality requirements, customer names and organizations are anonymized.

Three tests for sorting the work

Start with your own queue, not a list of services. Apply three tests to every recurring task, in this order. A Keep stops there, and an Agree on it first goes on to the third test once the rule is written down.

Test the task itself, not the area it touches. Setting the access policy is a decision. Granting access under that policy is a task.

Before the tests, ask whether the task should exist at all. The same how-to question every week means the docs are missing something, and a runner that sticks every Monday has a bug. Fix those first, or make fixing them part of the handoff.
Keep it
if it sets policy or needs your engineers’ depth: what the platform is for, who gets access, how much risk the business accepts, and fixes that need deep knowledge of your platform.
  • Architecture and roadmap
  • Access policy
  • Release risk tolerance
  • Shared pipeline libraries and platform code
  • Level 3 (L3) engineering and major incident command

These set the platform’s direction or need your engineers’ depth. They stay with you whoever runs the queue.

Agree on it first
if it changes how a policy is enforced in the pipeline, or what the platform costs to run.
  • Adding a scan stage
  • Changing branch protection
  • Resizing the runner fleet
  • Changing release strategy, such as canary or blue/green

Each of these puts a Keep decision into practice. Agree on the rule and write it down, then run the task through the third test.

Hand it off
if it’s repeatable, and someone who has never met your engineers could do it from a written runbook.
  • Flaky-test reruns and stuck deploys
  • Runner capacity alerts, within agreed scaling limits
  • Plugin and agent upgrades, through your change process
  • Access grants under an approved policy
  • Legacy pipeline break-fix until retirement
  • First response and triage, up to a declared major incident

Triage counts, because the steps to diagnose and route a fault repeat even when the fault is new. Work that can’t wait for business hours costs your team the most sleep, so move it early, once the runbooks have held up in daylight. A task that passes none of the three tests stays with your team.

Eight tickets, sorted

Here’s how the tests sort a typical week of tickets.

Hand it off Grant a new engineer access to the deploy pipeline. The access policy is already set, so granting access under it is routine work.
Keep it Decide who can approve production deploys. That’s access policy, a decision only your team can make.
Agree on it first Add an image scan stage to every pipeline. It changes how security policy is enforced. Once the rule is agreed, maintaining the stage and triaging its failures can be handed off.
Hand it off Rerun a flaky integration test and tag the owning team. The rerun follows a runbook. The fix belongs to the team that owns the test, so tagging them is the root-cause step.
Fix it first A build runner hangs every Monday morning. That’s a bug. Fix it, or make the fix part of the handoff, so nobody inherits a weekly restart.
Fix it first The same “how do I get secrets into my pipeline?” question, every week. The docs are missing something, so write the page and hand off what’s left.
Hand it off A production deploy fails at 2 am with an error nobody has seen. First response, including rollback under the agreed rule, follows a runbook, so it can be handed off. The fix may need L3, which stays with you.
Keep it Decide how much disruption a release may cause before it rolls back. That’s release risk tolerance. Agree on the rollback rule that puts it into practice, and running releases under that rule can be handed off.

Try it on last month’s queue. Export a month of platform tickets and chat requests, tag each one “keep,” “agree on it first,” “hand it off,” or “fix it first,” and note roughly how long it took.

Then total the hours in each pile. The handoff pile, plus any Agree on it first work that can follow it, is the roadmap time you could win back and the starting scope for any handoff.

Questions to settle before you hand off

Whoever takes the work, get these four answers in writing.

01
What happens to a ticket the runbook doesn’t cover? The right answer: it reaches a named person within an agreed time, and the fix becomes a new runbook.
02
Where do the runbooks live, and who can change them? Ask for them in your own repository or wiki, where your engineers can read and edit them.
03
Which issues escalate to your engineers, and how fast do they reach them? Agree who can declare a major incident, and at what severity.
04
Could you take the work back on three months’ notice? If not, find out what would stop you, and fix it before the handoff closes.

How to hand it off without losing control

Handing off operations goes wrong when it happens all at once, or when the knowledge leaves with it. The steps below apply whether the work goes to a partner or to another team inside your company.

If nobody can take the work yet, rotate one engineer a week onto the queue, and have them write or update a runbook for each type of ticket they close. The rest of the team works the roadmap, and you build the handoff as you go.
01 Name one owner. One person on your side decides what goes, in what order, and when it’s done.
02 Hand off once, into writing. Whoever takes the work writes the runbooks from your engineers’ knowledge, and your engineers review each one before it’s used. A usable runbook names the symptom, the access it needs, the exact steps, how to check the fix, how to roll back, who to escalate to, and who keeps it current.
03 Move one queue at a time. Start with one team, platform, or queue, and agree in advance when it counts as done. The receiving team has the access its runbooks need, no more, with every action logged. It has worked the queue for an agreed period with your engineers watching, and closed real tickets using only the runbooks, with escalations below an agreed rate.
04 Measure it every month. Set a baseline before anything moves: time to first response, time to resolution, how often work comes back to your engineers, hours spent on interrupts, and ticket volume per team, which should fall as root causes get fixed. Add DORA’s delivery metrics (change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate) to see whether your teams ship faster too.

Build the handoff so you can take it back

From 2018 to 2024, we ran platform support for a central developer platform serving 9 business units at a Fortune 100 technology company, on GitHub, Artifactory, internal CI, and Spinnaker. We stabilized it and handed it back to the internal team.

Whoever runs the work, keep the runbooks and escalation paths in your own systems, so you can take the work back.

CI/CD Platform Operations: the scope we operate, and the decisions your team keeps.

See what we run
Lime glass wrench illustration

Planning a handoff?

Bring a month of your platform queue. We’ll go through it with you and show which work could move first.

You define the outcomes. We own the delivery.

Share this post