We are looking for an experienced IT Service Recovery & Runbook Consultant to support a consultancy engagement focused on improving the ability of a large organisation to restore and recover the IT components underpinning its critical services.
This is a technical discovery and consultancy assignment, rather than a traditional Business Continuity or Operational Resilience role.
The initial engagement will assess the organisation's existing recovery documentation, establish a standardised approach to technical runbooks, identify the IT components requiring runbooks and determine the work required for a subsequent implementation programme.
The assignment
Working closely with platform engineering teams, technical SMEs and IT component owners, you will:
-
Define a standard service restoration runbook template, including the minimum technical processes and information required to recover IT components successfully.
-
Develop a runbook standard covering runbook creation, content, ownership, maintenance and ongoing management.
-
Review existing CMDB, service, application and architecture information to identify the IT components requiring recovery runbooks.
-
Determine the appropriate level of component decomposition and when infrastructure/components require standalone runbooks.
-
Build out an initial runbook catalogue, mapping critical services to the applications, middleware, databases, servers and other supporting components on which they depend.
-
Review existing recovery procedures, build documentation and technical artefacts to establish what can be reused.
-
Conduct a gap analysis across identified components, assessing existing documentation against the new runbook standard.
-
Assess remediation effort for individual runbooks using High/Medium/Low complexity and effort estimates.
-
Work directly with technical SMEs to create runbooks for two representative IT components.
-
Test and validate these runbooks, including technical recovery steps, commands, prerequisites, dependencies and validation checks.
-
Use the discovery and validation findings to develop a high-level delivery plan for the subsequent implementation phase.
-
Estimate the resources, technical SME effort, dependencies, timescales and costs required to create and test the remaining runbooks.
-
Develop an initial prioritisation approach and programme RAID log.
What we're looking for
We are particularly interested in people with a strong technical background across IT infrastructure, platforms and service recovery.
You should have experience of:
-
IT service restoration, recovery or disaster recovery within large and complex technology environments.
-
Creating or improving detailed technical runbooks, recovery procedures, operating procedures or operational documentation.
-
Understanding how applications and services depend upon underlying servers, databases, middleware, networks and infrastructure/platform components.
-
Working with infrastructure and platform engineering teams to understand how systems are built, operated and recovered.
-
Translating knowledge held by technical SMEs into clear, repeatable and testable recovery procedures.
-
Reviewing existing technical documentation and identifying gaps against an agreed standard.
-
Understanding service/component dependencies and the sequencing required for successful end-to-end recovery.
-
Working with CMDB and service/application information to understand technology estates and dependencies.
-
Facilitating technical discovery workshops and engaging effectively with engineers, architects, service owners and senior stakeholders.
-
Estimating the effort required for technical remediation or documentation programmes.
Knowledge of areas such as backup and restore, server recovery, databases, middleware, application recovery, infrastructure platforms, monitoring, patching, certificates, access requirements and disaster recovery would be highly beneficial.
The person
This role requires someone who can operate at both a consultancy and technical level.
You will need to be comfortable going into an environment where documentation may be incomplete, speaking directly with engineering teams to establish how services and components actually work, identifying what is missing and turning that information into a structured and practical approach.
We are therefore looking for someone who is technically credible with infrastructure and platform teams, but who can also step back and define the standards, methodology, estimates and programme of work required to address the wider estate.
This would suit someone from an IT Service Recovery, Infrastructure Recovery, Disaster Recovery, Infrastructure Resilience, Platform Engineering or Technical Service Management background who has subsequently moved into technical consultancy or programme delivery.
#LI_DNI