Senior DevOps Engineer
Job Description Key Responsibilities Cloud Architecture \& Provisioning - Lead the design, deployment, configuration, and maintenance of Client’s solutions on public cloud platforms (AWS, Azure, GCP) with a focus on high availability, scalability, and security. - Architect and implement standardized, reusable infrastructure patterns (e.g., templates, modules) to support Client’s PCE deployments at scale. - Own and enhance automated provisioning processes to accelerate deployment cycles, minimize manual intervention, and improve reliability. - Partner with development, QA, and platform teams to plan and provision environments for complex testing, performance, and release activities. Operations, Reliability \& Optimization - Serve as an escalation point for complex operational issues, performing advanced troubleshooting across application, infrastructure, and network layers. - Design and implement strategies for reliability, capacity planning, performance tuning, and cost optimization across multi-cloud environments. - Drive adoption of SRE/DevOps practices such as error budgets, SLIs/SLOs, and post incident reviews, and ensure follow through on corrective actions. - Implement and refine continuous monitoring and observability solutions to proactively identify, diagnose, and resolve issues. Security, Compliance \& Governance - Work closely with security, compliance, and governance teams to design, implement, and maintain robust security architectures and controls in cloud environments. - Ensure adherence to industry standards and regulatory requirements relevant to U.S. national security and critical infrastructure customers. - Lead regular security reviews, audits of access controls, configuration baselines, and infrastructure policies; drive remediation and hardening initiatives. - Champion securebydesign and DevSecOps practices within CI/CD pipelines and infrastructure automation. - Automation, Tooling \& CI/CD - Design and develop advanced automation frameworks, scripts, and tools to eliminate manual work, reduce risk, and improve operational efficiency. - Architect, implement, and maintain CI/CD pipelines for code deployment, configuration management, and infrastructure as code, including governance and quality gates. - Evaluate, recommend, and integrate new DevOps tools, platforms, and technologies that improve velocity, reliability, and security of deployments. - Ensure consistency, standardization, and best practices in the use of configuration management (e.g., Ansible, Terraform) across teams and environments. Technical Leadership, Collaboration \& Documentation - Act as a technical leader and subject matter expert for DevOps practices, cloud infrastructure, and automation within the team and across adjacent teams. - Mentor and coach junior and midlevel engineers, providing guidance on design, implementation, troubleshooting, and professional development. - Lead cross-functional initiatives with developers, QA engineers, system administrators, security, and operations teams to deliver complex projects. - Produce and maintain high quality documentation for reference architectures, runbooks, standard operating procedures, and troubleshooting guides. - Influence and help define team standards, best practices, and roadmaps for DevOps and cloud operations. Qualifications - Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field; or equivalent work experience. - Substantial hands-on experience (typically 7+ years) in DevOps, Site Reliability Engineering, or cloud infrastructure roles, including: - Proven track record deploying, managing, and optimizing solutions on public cloud platforms (AWS, Azure, GCP). - Strong expertise with scripting languages (e.g., Python, Bash) for automation and tooling. - Advanced experience with infrastructure as code and configuration management tools (e.g., Ansible, Terraform, CloudFormation, ARM templates). - Deep understanding of DevOps principles and practices, including CI/CD, infrastructure as code, continuous monitoring, and incident management. - Practical experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack or similar) and logging/tracing strategies. - Strong knowledge of security best practices for cloud environments, including identity and access management, network security, and encryption. - Demonstrated ability to lead complex technical initiatives, influence architecture decisions, and collaborate across diverse technical and business stakeholders. - Excellent analytical and problem solving skills, with the ability to perform under pressure in mission critical environments. - Strong written and verbal communication skills, including the ability to explain complex technical topics to both technical and nontechnical audiences. Preferred Qualifications - Professional certifications in one or more cloud platforms (e.g., AWS Certified DevOps Engineer – Professional, AWS Solutions Architect – Professional, Azure DevOps Engineer Expert, Google Cloud Professional DevOps Engineer). - Experience with SAP HANA, SAP Cloud Platform services, or SAP Private Cloud Edition implementations. - Knowledge of networking concepts and protocols (e.g., TCP/IP, DNS, VPN, load balancing, firewalls) in a cloud and hybrid context. - Experience with containerization and orchestration (e.g., Docker, Kubernetes) and related ecosystem tools. - Previous experience supporting government, defense, or other highly regulated environments, including exposure to compliance frameworks and accreditation processes. - Familiarity with Agile methodologies and working within cross-functional, multidisciplinary teams.