“A hands-on systems quality engineering, test automation, quality leadership, coaching and delivery accountability.“
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
The Systems Quality Engineer is responsible for improving the quality, reliability, and release confidence of Cloud Direct’s internal and customer facing business systems through test automation, testing standards, and effective release readiness practices.
Sitting within the Systems Reliability & Automation team, this role will help shape Cloud Direct’s testing strategy, build and maintain automated test suites, integrate testing into CI/CD pipelines, improve test coverage, and act as a shared service supporting reliable delivery across the wider Business Systems function.
The role also provides practical coordination of user acceptance testing, working with business stakeholders to ensure system changes are properly validated before release. It is not just a traditional manual testing role; it is a hands-on quality engineering role focused on making testing more repeatable, automated, structured, and useful.
Success in this role will be demonstrated by:
- Increased automated test coverage across key Business Systems platforms.
- Clearer, more consistent testing standards across Engineering, Operations, and Reliability teams.
- More reliable release readiness, with fewer defects reaching production.
- Better coordination and evidence of UAT for business-critical changes.
- Improved confidence in CI/CD pipelines, release gates, and regression testing.
- Stronger collaboration between technical teams and business stakeholders during testing and release activity.
What you’ll be doing?
Test Automation & Quality Engineering:
- Design, build and maintain automated test suites for Cloud Direct’s business systems, integrations and critical workflows.
- Integrate automated testing into CI/CD pipelines to provide meaningful quality gates across development and release processes.
- Continuously improve test coverage, test reliability and testing standards across the Business Systems estate.
- Define practical approaches for smoke, regression, release-gate and end-to-end testing.
Release Readiness & UAT:
- Support release readiness by ensuring changes are appropriately tested, evidenced and validated before production deployment.
- Coordinate user acceptance testing with business stakeholders, helping define test scenarios, capture outcomes and obtain sign-off.
- Improve the consistency of testing, release validation and quality assurance practices across Business Systems.
Continuous Improvement & Collaboration:
- Work with Systems Engineering, Systems Operations and Systems Reliability & Automation teams to improve quality processes, reduce change risk and support reliable delivery.
- Identify recurring quality issues and drive improvements to testing approaches, tooling, documentation and release practices.
- Contribute to a culture of engineering quality, continuous improvement and shared ownership of system reliability.
What you’ll need:
- 3+ years’ experience in software quality engineering, test automation, software testing, systems testing, or a similar technical quality role.
- Experience defining test strategies for smoke, regression, release-gate, and end-to-end testing.
- Proven experience building and maintaining automated test suites (using Playwright, Cypress or equivalent) for web applications, APIs, and integrations as part of complex business systems.
- Experience integrating automated testing into CI/CD pipelines (using Azure DevOps, GitHub or equivalent).
- Good technical understanding of modern development practices, including source control, branching, pull requests, and pipeline-based delivery.
- Deep understanding of software delivery lifecycle, release readiness, regression testing, and defect management.
- Ability to define clear test scenarios, acceptance criteria, test evidence, and sign-off expectations.
- Proven experience coordinating end user acceptance testing with business stakeholders.
- Strong analytical and evidence-driven problem-solving skills, with high attention to detail.
- Ability to communicate clearly with both technical and non-technical stakeholders, feeling comfortable to challenge their assumptions where needed.
- Strong documentation discipline, including test plans, test evidence, process guidance, and release notes.
- Must be able to work from our Cape Town office at least two days per week.
Experience and Skills Desired:
- Relevant testing or quality certifications, where supported by practical experience.
- Experience in writing and maintaining tests in languages such as C#, TypeScript, JavaScript and PowerShell.
- Experience testing APIs using tools such as Postman, REST clients, or automated API test frameworks.
- Experience with Azure Test Plans or similar test management tooling.
- Experience with test data management, mocking, fixtures, and repeatable test environments.
- Exposure to business-critical internal platforms such as ServiceNow, bespoke full stack applications built with SQL/C# .NET/React, and custom AI agents
What we’re offering:
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talent.
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.
“Can you lead high-pressure incidents with confidence, restore critical services, and keep customers informed every step of the way?”
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
We are looking for a Major Incident Manager to join our Operations team and take ownership of the end-to-end response to high-impact service incidents across Cloud Direct’s managed services portfolio.
The Major Incident Manager acts as the central point of accountability during a Major Incident, coordinating technical resolver groups, service delivery stakeholders, third-party suppliers, customer contacts, and leadership teams to restore service as quickly and safely as possible.
This role is responsible for assessing and declaring Major Incidents, leading the Major Incident process, maintaining clear and controlled communication, and ensuring that all actions, decisions, timelines, risks, impacts, and outcomes are accurately documented.
The role also contributes directly to Problem Management, Change Enablement, and selected elements of Request Management. This ensures that lessons learned from Major Incidents are converted into corrective actions, recurring issues are addressed, changes are assessed with operational risk in mind, and service requests that indicate risk or degradation are handled appropriately.
This is a highly visible service management role that requires calm leadership, strong customer focus, clear communication, and the ability to drive operational control during time-sensitive, high-pressure situations.
The successful candidate will help protect customer confidence, reduce business disruption, improve service resilience, and support continual improvement across Cloud Direct’s IT service management practices.
What You’ll Do:
Major Incident Management:
- Own and manage the Major Incident Management process from assessment and declaration through to service restoration, review, and closure.
- Assess incident priority, urgency, impact, customer sensitivity, business risk, and service criticality to determine whether an incident should be declared as major.
- Formally declare Major Incidents in line with Cloud Direct and customer-specific governance requirements.
- Act as the central point of contact for all information related to the Major Incident.
- Lead Major Incident bridge calls, technical war rooms, stakeholder calls, and internal coordination forums.
- Coordinate internal technical teams, Service Desk teams, Service Delivery Managers, Service Level Managers, third-party vendors, and customer representatives.
- Maintain operational control by ensuring clear ownership of actions, decisions, risks, dependencies, timescales, and escalation points.
- Drive rapid service restoration while balancing customer impact, technical risk, change risk, and business priorities.
- Ensure resolver groups provide timely updates, clear technical direction, accurate impact information, and realistic estimated restoration times.
- Escalate effectively to senior leadership, vendors, technical specialists, or customer stakeholders where progress, ownership, or risk requires intervention.
- Ensure Priority 1, high-severity, or customer-impacting incidents are managed consistently, professionally, and in line with agreed service management standards.
- Maintain accurate Major Incident records, including timelines, customer impact, actions taken, technical updates, communications, decisions, and closure evidence.
- Ensure Major Incident communications are issued at agreed intervals and are appropriate for technical, operational, customer, and executive audiences.
- Prepare customer-facing and executive-level Major Incident reports that clearly explain impact, timeline, resolution activity, root cause status, and recommended next steps.
- Facilitate Major Incident wash-up meetings, ensuring key stakeholders attend and meaningful lessons learned are captured.
- Support out-of-hours or on-call Major Incident activity where required by operational needs.
- Ensure Major Incident records are suitable for audit, service review, compliance, and continual improvement purposes.
Problem Management:
- Ensure all Major Incidents are reviewed for Problem Management qualification and ownership.
- Facilitate or contribute to Root Cause Analysis and Post Incident Review activities following Major Incidents.
- Work with technical resolver teams to identify root cause, contributing factors, failed controls, process gaps, and recurrence risks.
- Ensure corrective and preventative actions are clearly documented, assigned to owners, tracked, and followed through to completion.
- Identify recurring incidents, repeat failures, trend patterns, customer pain points, and known service risks.
- Support the creation and maintenance of Known Error Records, workaround documentation, and Knowledge Base articles.
- Provide input into Problem Review meetings, service improvement forums, operational governance meetings, and customer service reviews.
- Ensure lessons learned from Major Incidents are converted into practical improvements that reduce future service disruption.
Change Enablement:
- Contribute Major Incident and Problem Management insight into Change Enablement activities.
- Participate in Change Advisory Board, Emergency Change Advisory Board, and operational change review meetings where required.
- Review failed changes, emergency changes, and changes linked to Major Incidents to identify lessons learned and improvement actions.
- Support operational risk assessments by highlighting known incidents, recurring issues, affected services, vulnerable customers, or unresolved problems.
- Ensure emergency changes raised during Major Incident recovery are appropriately documented, approved, and reviewed after implementation.
- Promote stronger alignment between Incident, Problem, and Change practices to improve service stability and reduce avoidable disruption.
Request Management:
- Support selected Request Management activities where escalation, prioritisation, customer impact, or service restoration risk requires operational coordination.
- Assist with high-priority or sensitive service requests that may affect critical business services or customer outcomes.
- Identify service requests that may indicate underlying incidents, recurring problems, access issues, failed fulfilment, or process weaknesses.
- Work with Service Desk and Operations teams to improve request routing, fulfilment quality, escalation paths, and customer communication.
- Contribute to request fulfilment improvements where automation, knowledge articles, templates, or process clarification can reduce operational effort.
Governance, Reporting and Continual Service Improvement:
- Maintain Major Incident dashboards, service metrics, trend reports, and operational performance summaries.
- Track and report key measures such as Major Incident volume, Mean Time to Restore Service, communication timeliness, recurrence, root cause completion, action closure, and customer impact.
- Provide clear reporting for service reviews, operational governance forums, leadership updates, and customer meetings.
- Identify opportunities to improve process maturity, operational resilience, knowledge quality, communication templates, automation, and service reporting.
- Promote adherence to ITIL-aligned Incident, Major Incident, Problem, Change, and Request Management practices.
- Support audit, compliance, and evidence requirements where Major Incident, Problem, Change, or Request Management records are required.
- Build strong working relationships with customers, internal stakeholders, technical teams, vendors, and service delivery colleagues.
- Champion a culture of ownership, accountability, learning, customer focus, and continual service improvement.
What We’re Looking For:
- Minimum 3 years’ experience in Major Incident Management, Incident Management, Service Delivery, IT Operations, Service Management, or a similar operational role.
- Proven experience leading high-severity, high-priority, or business-critical incident response activities.
- ITIL Foundation certification is required or must be achieved within an agreed timeframe.
- ITIL Managing Professional modules, ITIL Specialist modules, SIAM, Service Desk Institute, or equivalent Service Management certifications are desirable.
- Strong understanding of ITIL-aligned Incident Management, Major Incident Management, Problem Management, Change Enablement, and Request Management practices.
- Experience chairing incident bridge calls, coordinating technical resolver groups, and maintaining control during high-pressure operational situations.
- Experience managing customer, stakeholder, and leadership communications during live incidents.
- Experience facilitating Root Cause Analysis, Post Incident Reviews, wash-up meetings, and lessons learned sessions.
- Experience documenting incident timelines, customer impact, action logs, decisions, escalations, resolution activity, and service improvement actions.
- Experience using ITSM platforms such as ServiceNow, HaloITSM, Jira Service Management, Freshservice, ManageEngine, or similar systems.
- Strong understanding of cloud services, infrastructure, networking, security, end-user computing, and modern workplace environments at a level sufficient to coordinate technical teams effectively.
- Experience working with third-party suppliers, vendors, and external support organisations during escalations.
- Ability to analyse operational metrics, identify trends, and convert findings into practical improvement actions.
Highly Desirable:
- Experience working within a Managed Service Provider environment supporting multiple customers, services, and technical towers.
- Experience supporting Microsoft Azure, Microsoft 365, Modern Workplace, Security, and Cloud Infrastructure services.
- Experience chairing or contributing to CAB, ECAB, Problem Review, Major Incident Review, operational governance, or customer service review meetings.
- Experience creating or improving Major Incident templates, communication templates, RCA templates, dashboards, service reports, and operational playbooks.
- Experience with ServiceNow reporting, dashboards, ticket history analysis, activity streams, related records, and audit evidence extraction.
- Knowledge of ISO 20000, ISO 27001, audit readiness, evidence management, and service governance best practices.
- Experience using Power BI, Power Automate, Excel dashboards, or other reporting and automation tools to improve operational visibility.
- Experience mentoring Service Desk, Incident Management, or Service Delivery colleagues in Major Incident process discipline and communication standards.
- Knowledge management experience, including the creation of Knowledge Base articles, known error documentation, operational guides, and lessons learned records.
- Exposure to enterprise-scale environments where incidents may affect revenue, operations, compliance, reputation, or customer-critical services.
- Comfortable working across globally distributed teams, customers, vendors, and time zones when required. -900
What We Offer:
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talent.
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.
“Can you turn cloud cost insights into measurable business value while driving customer adoption and long-term FinOps success?”
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
The FinOps Value & Adoption Manager helps customers convert FinOps insight into measurable value realisation from Cloud Direct’s FinOps services. The role owns the customer engagement lifecycle from onboarding through continual optimisation, ensuring identified opportunities become approved actions, implemented improvements, evidenced outcomes and validated value realised against agreed baselines.
Working across Azure FinOps today and future Microsoft 365 FinOps and AI FinOps services, the role acts as a trusted advisor to technical, operational, financial and executive stakeholders. Success is measured not by identifying optimisation opportunities alone, but by influencing customer adoption, accelerating implementation, evidencing value realisation, reporting validated outcomes and increasing customer FinOps maturity.
What You’ll Be Doing:
Customer Engagement & Relationship Management:
- Own a portfolio of FinOps customer engagements and act as the primary customer-facing value and adoption lead.
- Establish trusted advisor relationships with technical, operational, finance and executive stakeholders.
- Facilitate recurring service reviews, optimisation workshops and governance meetings.
- Develop customer-specific optimisation roadmaps aligned to business objectives, adoption priorities, value-realisation opportunities and FinOps maturity goals.
- Manage customer expectations, priorities and service outcomes.
Value Realisation & Adoption:
- Drive customer adoption of optimisation recommendations across Azure, Microsoft 365 and AI workloads.
- Convert identified opportunities into approved implementation plans.
- Maintain optimisation backlogs and track progress through identified, approved, implemented and value-realised stages.
- Remove blockers and coordinate stakeholder engagement to accelerate recommendation adoption.
- Validate and evidence savings and value outcomes against agreed baselines.
- Support gain-share measurement and reporting activities where applicable.
FinOps Advisory & Governance:
- Deliver FinOps best-practice guidance and coaching.
- Support customers in improving financial accountability, forecasting and governance maturity.
- Facilitate prioritisation of optimisation opportunities and investment decisions.
- Drive continual service improvement through structured governance processes.
- Escalate delivery risks, optimisation blockers and governance concerns when required.
Service Delivery & Reporting:
- Support onboarding activities including baseline assessments, stakeholder workshops and optimisation planning.
- Produce monthly service reports, optimisation dashboards and executive value reports.
- Coordinate cross-functional delivery activities involving Managed Services, Professional Services, Engineering and customer teams.
- Ensure customer actions, recommendations, decisions and outcomes are appropriately documented and tracked.
- Provide customer feedback and market insights to support evolution of Cloud Direct’s FinOps service portfolio.
Service Development & Continuous Improvement:
- Contribute to the development and enhancement of Azure, Microsoft 365 and AI FinOps services.
- Support the creation of repeatable engagement methodologies, reporting standards and governance frameworks.
- Identify opportunities to improve customer experience, adoption rates and service effectiveness.
- Act as a customer advocate within Cloud Direct to ensure services remain outcome-focused and value-led.
What We’re Looking For:
- Exceptional communication, presentation and stakeholder management skills.
- Demonstrable experience in customer-facing consulting, customer success, service delivery or advisory roles.
- Proven ability to influence stakeholders and drive action without direct authority.
- Strong relationship-building and trusted advisor capabilities.
- Commercial awareness and an ability to articulate technical value in business terms.
- Experience managing multiple customers, priorities and workstreams simultaneously.
- Strong organisational, planning and reporting skills.
- Technical literacy sufficient across Microsoft & FinOps optimisation platforms to engage credibly with customer Technical stakeholders and Internal engineering & support teams.
- Experience working with commercial, finance or operational stakeholders to translate optimisation opportunities into agreed actions.
- Working knowledge of FinOps principles, cloud cost optimisation, value tracking and financial accountability practices.
Highly Desirable:
- FinOps Foundation Practitioner, Professional or equivalent certification.
- Azure administration, engineering or architecture experience.
- Microsoft 365 licensing, administration or optimisation experience.
- Knowledge of Microsoft Copilot, AI services and consumption-based charging models.
- Experience delivering cloud optimisation, cost management or value realisation programmes.
- Managed Services, Professional Services or technology consultancy experience.
- Experience in governance, financial management or operational transformation initiatives.
- Executive reporting and business review facilitation experience.
What do we offer you?
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talent.
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.
“Can you troubleshoot, optimise, and support Azure cloud environments with confidence?”
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
We are looking for a Tier 2 Azure Cloud Associate Engineer to join our team in Cape Town.
In this role, you’ll provide advanced remote support while working closely with our project delivery and specialist engineering teams. As a Tier 2 Azure Cloud Associate Engineer, you’ll act as an escalation point for Tier 1 engineers, taking ownership of more complex technical issues and collaborating across Operations to deliver exceptional customer outcomes.
You’ll be responsible for ensuring service level agreements are consistently met by delivering timely, high-quality resolutions to customer incidents and requests received via our support channels, including phone, live chat, email, and self-service portals. Along the way, you’ll help customers maximise the value of their Azure environments through technical expertise, proactive problem-solving, and a customer-first approach..
What You’ll Do:
- Provide Level 2 technical support for Azure infrastructure and cloud platform services across multiple customer environments.
- Deploy, configure, and manage Azure compute resources, including Virtual Machines, storage, networking, and associated services.
- Support the administration of Azure subscriptions, resource groups, and resource configurations.
- Configure and manage Azure Virtual Networks, Network Security Groups (NSGs), and connectivity components.
- Monitor the health, performance, and availability of Azure environments using Azure Monitor and Log Analytics.
- Deploy, configure, and optimise Azure Virtual Machines to ensure security, performance, and cost efficiency.
- Perform routine backup, restore, and disaster recovery activities using Azure Backup and Azure Site Recovery.
- Support the management of Azure Storage services, including Azure Files, Blob Storage, Managed Disks, and snapshots.
- Assist with securing Azure resources by implementing Role-Based Access Control (RBAC) and following Azure security best practices.
- Support and maintain Azure Platform as a Service (PaaS) offerings, including App Services, Azure SQL Database, and Azure Key Vault.
- Assist customers with Azure cost optimisation by identifying opportunities to improve resource utilisation and reduce unnecessary spend.
- Develop and maintain PowerShell scripts and automation to improve operational efficiency and reduce manual tasks.
- Perform troubleshooting and root cause analysis for Azure infrastructure incidents, implementing permanent fixes where appropriate.
- Escalate complex technical issues to senior engineers while providing detailed troubleshooting and diagnostic information.
- Work across multiple customer Azure environments, ensuring service levels and operational standards are consistently met.
- Maintain accurate technical documentation, support records, and operational procedures.
- Collaborate effectively with engineers, consultants, and customer stakeholders to deliver high-quality technical support.
- Contribute to continuous service improvement by identifying opportunities to enhance processes, automation, and customer experience.
What We’re Looking For:
- At least 2 years of hands-on experience in a Microsoft Azure support role.
- Experience In the deployment, upkeep, and monitoring of Azure systems
- Strong understanding of the Microsoft Azure platform and its capabilities
- Two or more Azure cloud certifications: AZ-104, AZ-500, AZ-140, SC-200, SC-300, DP-900, Terraform
- Logical approach to problem solving
- Excellent inter-personal skills with good time management
- Proven customer service & communication skills
What We Offer:
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talent.
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.
“A hands-on leadership role driving system reliability, automation, and operational excellence”
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
The Systems Reliability & Automation Lead is responsible for applying software engineering and automation practices to improve the reliability, efficiency and scalability of Cloud Direct’s internal systems and managed service operations. This is a hands-on leadership role focused on reducing operational effort through automation, improving service quality through reliability engineering, and creating the capabilities required to support Cloud Direct’s continued growth.
The role works closely with Managed Services, Business Systems, and Security teams to identify operational bottlenecks, eliminate repetitive manual activities and implement reliable, scalable solutions. The core focus will be on elements such as reducing the workload caused by monitoring and alerting.
As the first hire in this function, you will be expected to establish the Systems Reliability & Automation capability, define its strategy, standards and roadmap, and build a future team as the function matures.
Success in this role will be demonstrated by:
- Reduction in manual operational effort across Managed Services.
- Measurable reduction in alert volumes through automation and optimisation.
- Increased percentage of operational processes that are automated.
- Improved availability and reliability of critical business systems.
- Reduction in recurring incidents through effective root cause analysis.
- Establishment of a Systems Reliability & Automation function with clear standards, tools and operating procedures.
What you’ll be doing?
Reliability Engineering:
- Improve the availability, performance and resilience of Cloud Direct’s business-critical systems.
- Establish monitoring, alerting and reliability standards across key platforms and services.
- Lead or support root cause analysis activities for significant incidents, identifying opportunities for automation, resilience improvements and permanent corrective actions.
- Define and report on key reliability and operational performance metrics.
Automation & Operational Excellence:
- Identify and eliminate repetitive manual work across Managed Services operations.
- Design and implement automation solutions that improve efficiency and reduce human error.
- Improve alert management through intelligent routing, automation and remediation.
- Develop self-service capabilities that reduce operational overhead and improve user experience.
- Continuously identify opportunities to streamline operational processes through technology.
Leadership & Strategy:
- Establish and lead the Systems Reliability & Automation function.
- Define standards, tooling and ways of working for the team.
- Build and maintain a roadmap of reliability and automation initiatives aligned to business priorities.
- Support future recruitment, mentoring and development of team members as the capability grows.
Collaboration:
- Work closely with Managed Services teams to understand operational challenges and identify improvement opportunities.
- Partner with Business Systems and Security teams to deliver reliable and scalable solutions.
- Provide leadership and guidance on reliability engineering and operational automation practices across the organisation.
What you’ll need:
- 5+ years experience in Reliability Engineering, Platform Engineering, DevOps, Automation Engineering, Systems Engineering or a similar technical leadership role.
- 2+ years leading technical initiatives, programmes or teams.
- Proven experience reducing operational effort/improving operational efficiency through automation in a Managed Services, IT Operations or enterprise technology environment.
- Strong understanding of incident management, problem management and operational best practices.
- Experience defining and measuring service reliability objectives and operational performance metrics.
- Strong scripting and automation skills using PowerShell, Python or similar technologies.
- Experience working with Microsoft Azure and cloud-hosted platforms.
- Experience implementing and improving monitoring, observability and alerting solutions using tools such as Azure Monitor, Log Analytics and Grafana.
- Experience integrating systems and services through APIs, workflows and automation tooling.
- Experience with CI/CD pipelines using Azure DevOps, GitHub or similar platforms.
- Ability to balance strategic planning with hands-on technical delivery.
- Ability to communicate effectively with both technical and operational stakeholders.
- Must be able to work in our Cape Town office 2 days a week.
Experience and Skills Desired:
- Experience with ServiceNow, or similar operational platforms.
- Experience with Microsoft Power Platform and Copilot technologies.
- Experience with Infrastructure as Code tooling such as Terraform or Bicep.
- Experience with Docker and Azure Container Apps.
- Experience with AIOps, intelligent automation or AI-assisted operations.
- Previous experience establishing a new technical capability or team.
What we’re offering:
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talen
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.
“Ready to be the first line of support, solve real problems, and keep our users productive every day?”
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
We’re looking for an ambitious and energetic Service Desk Analyst to join our rapidly expanding Cape Town office.
This role isn’t just about fixing IT issues—it’s about growing your career. You’ll be engaging with international clients, providing top-tier support, and working alongside a dynamic team that’s as committed to learning as they are to delivering excellent service. If you’re eager to work hard, level up your skills, and climb the support ladder, this is the place for you
What You’ll Do:
- Be the first point of contact for client support requests
- Provide remote technical support across a range of hardware and software
- Escalate complex issues to the senior team when needed
- Log and classify tickets, ensuring seamless case management
- Contribute to Standard Operating Procedures (SOPs) and knowledge-sharing
- Deliver a professional, positive experience for every client interaction
What We’re Looking For:
- A diploma/degree in Computer Science, Information Systems, or Engineering, OR at least one year of hands-on IT support experience
- Strong knowledge of Windows Desktop & Microsoft Operating Systems
- Experience with Microsoft Office Suites & Office 365
- Familiarity with Active Directory]
- Required certifications in MS-900 (Microsoft 365 Fundamentals), N+ and or CompTIA A+
Highly Desirable:
Certifications in the following:
- ITIL 4 Foundation
- AI-900
- CompTIA A+/N+ (if needed)
- SC-900
What We Offer:
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talent.
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.
“Can you build and lead an AI-native Security team that protects our people, our customers, and our reputation — while helping define the future of security services in a human-plus-AI world?”
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
Cloud Direct is evolving its security capability to protect our organisation and the customers that depend on us. As our Senior Security Lead, you will be the architect and operational owner of this capability — shaping detection and response, guiding the use of automation and AI across the Microsoft Security stack, and building a high-performing Security team. This role offers the opportunity to help define modern, AI-enabled security services that combine expert judgement with intelligent automation.
This is a hands-on leadership role. You will define detection logic, lead incident response, mentor analysts, and report directly to the CEO and leadership team. You will shape not only how we defend ourselves but how we bring modern security capabilities to market.
What You’ll Do:
Security Platform Architecture & Build:
- Design the end-to-end security monitoring and response capability using Microsoft Sentinel, Microsoft Defender, and the wider Microsoft Security stack.
- Architect the security platform and operating model so it can scale effectively across internal and customer environments over time.
- Assess the current Microsoft estate and identify opportunities to strengthen security outcomes through better use of existing capabilities, automation, and AI.
- Define and deploy log-ingestion strategy across endpoints, identity (Entra ID), email, and cloud workloads.
- Shape the use of complementary Microsoft Security capabilities to improve visibility, prioritisation, and response across the environment.
Detection Engineering & Threat Response:
- Develop and tune Sentinel analytics rules, KQL queries, and automated playbooks to detect high-priority threats across identity, endpoint, collaboration, and cloud workloads.
- Author and maintain investigation runbooks and standard operating procedures for all alert categories.
- Act as the primary escalation point for P1/P2 security incidents, coordinating containment, eradication, and recovery.
- Lead proactive threat hunting, purple-team collaboration, and continuous improvement activities to strengthen coverage and resilience.
Team Leadership & Mentoring:
- Lead and develop security analysts, creating clear operating rhythms, coaching, and capability growth across the team.
- Define a pragmatic coverage and escalation model that balances human expertise, automation, and intelligent assistance.
- Mentor team members in modern detection, investigation, response, and security engineering practices across the Microsoft ecosystem.
- Foster a culture of continuous learning through tabletop exercises, post-incident reviews, and knowledge sharing.
Operational Reporting & Governance:
- Produce regular security performance reporting for leadership, covering operational trends, incident themes, and opportunities for improvement.
- Integrate security workflows with ServiceNow for case management and Dynamics for commercial pipeline tracking.
- Own security-related compliance and audit readiness for UK GDPR (ICO) and South Africa POPIA.
Commercial Security Service Development:
- Partner with Sales and Pre-Sales to shape a modern managed security service aligned to customer needs and the Microsoft Security opportunity.
- Define service outcomes, onboarding approaches, and operating principles for customer-facing security services.
- Contribute to the evolution of Cloud Direct’s broader security services strategy and go-to-market proposition.
What We’re Looking For:
- Strong hands-on experience in security operations, incident response, detection engineering, or security engineering.
- Deep expertise with Microsoft Sentinel (KQL, analytics rules, playbooks, workbooks) and the Microsoft Defender suite.
- Proven experience building or significantly maturing a security operations capability — ideally within an MSP, MSSP, or multi-tenant environment.
- Strong knowledge of MITRE ATT&CK, common adversary TTPs targeting MSPs, and threat-hunting methodologies.
- Experience leading, mentoring, and developing junior security analysts.
- Excellent communication skills — able to translate technical findings into clear, actionable reports for senior leadership.
- Relevant certification: GIAC GCIH, Microsoft SC-200, or equivalent.
Highly Desirable:
- Experience designing or operating a commercial managed security or MDR offering.
- Familiarity with Microsoft’s extended detection, response, and security operations capabilities across endpoint, identity, email, and cloud.
- Working knowledge of ServiceNow (SecOps module), Entra ID, Intune, and Azure Arc.
- Understanding of UK GDPR/ICO and South Africa POPIA compliance requirements.
- Additional certifications: CISSP, CISM, GSOM, or Microsoft SC-100.
- Background in MSP toolchain security (RMM, remote access, PSA platforms).
What We Offer:
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talent.
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.
“Ready to optimize Azure environments, and engineer scalable systems that empower digital transformation from the ground up?”
As a frontier partner, we grow through great people, smart tech, and teamwork between humans and AI.
As an Azure Cloud Expert Engineer, you will provide high-level support to our customers, acting as an escalation point for Tier 2 support engineers and working alongside senior engineers to execute complex projects/onboardings.
You will ensure that service level agreements are met whilst delivering reliable resolution of customer technical issues when raised via Cloud Direct support lines, Live chat, e-mail or self-service portals.
What You’ll Do:
- Azure Infrastructure: Provision and configure VMs, containers, and resources using ARM templates, Terraform, Bicep, and scripting (PowerShell). Optimize for cost, performance, and security.
- Automation & IaC: Implement Infrastructure-as-Code using Azure DevOps, GitHub, and Visual Studio. Automate deployments and support processes.
- Security & Compliance: Enforce best practices with tools like Azure Security Center, AAD, and Sentinel. Manage RBAC, conduct audits, and address vulnerabilities.
- Monitoring & Optimization: Use Azure Monitor and related tools to track, troubleshoot, and improve infrastructure performance.
- Disaster Recovery: Design and manage DR and high availability solutions with Azure Backup and Site Recovery. Test and validate failover strategies.
- Azure Virtual Desktop: Deploy and manage AVD environments, handle user access, monitor performance, and resolve connectivity or user issues.
- Data & AI Support: Diagnose and fix pipeline failures, manage integrations, and ensure data accuracy and availability.
- Platform Management: Ensure backup integrity, monitor system health, resolve failures, and meet RTO/RPO objectives
What We’re Looking For:
- 3+ years of experience designing, implementing, and managing Azure cloud environments.
- Excellent communication and collaboration skills.
- Strong understanding of the Microsoft Azure platform and its capabilities
- Three or more Azure cloud certifications: AZ-104, DP-100, DP-300, AZ-500, AZ-305, SC-200, SC-300, AZ-140 customer service & communication skills
What We Offer:
- Responsible Time off (uncapped annual leave)
- Group Life Cover /Disability Income Cover/ Trauma Insurance Cover (Injury / Disability)
- Fitness Cash Contribution
- Pension Fund Contribution
- Medical Insurance Contribution
- Employee Assistance Programme
- Enhanced Maternity & Paternity Leave
- Endless Growth Opportunities: We provide ample opportunities for professional development, mentoring, and advancement.
- Culture of Excellence: We foster a high-performance culture that recognizes and rewards exceptional talent.
At Cloud Direct, we believe that diversity, equity, and inclusion are essential to our success. We are committed to creating a workplace where everyone feels valued, respected, and empowered. We welcome applicants from all backgrounds and strive to build a team that reflects the diverse communities we serve. We encourage candidates of all races, ethnicities, genders, sexual orientations, ages, abilities, and socioeconomic statuses to apply.