Senior Cloud Infrastructure Engineer. at MNC · Bengaluru (Bangalore) · 10 - 18 years · ₹5L - ₹24L / yr · Posted 23 Sep 2026

Cloud Infrastructure Engineer – BANG | 10+ Years
Location: Bangalore
Experience: 10+ Years
Job Description:
- Design, implement, and manage cloud infrastructure across AWS/Azure/GCP environments.
- Strong experience in cloud architecture, compute, storage, networking, and security.
- Manage VMs, VPC/VNet, load balancers, DNS, DHCP, firewalls, and IAM.
- Hands-on experience with Windows/Linux servers, VMware, virtualization, and infrastructure operations.
- Automate infrastructure provisioning and configuration using Terraform, Ansible, or similar tools.
- Monitor infrastructure performance, availability, and capacity using tools such as Grafana, Prometheus, or CloudWatch/Azure Monitor.
- Handle incident management, troubleshooting, disaster recovery, backup, and high-availability requirements.
- Work with cross-functional teams to support cloud migration, infrastructure upgrades, and production environments.
- Ensure infrastructure follows security, compliance, and operational best practices.
Must-Have Skills:
Cloud Infrastructure | AWS/Azure/GCP | Networking | Linux/Windows | VMware | Terraform | Ansible | DNS/DHCP | IAM | Monitoring | Backup & DR

Similar jobs (10)
WowPe is a leading fintech company revolutionizing the way businesses handle financial transactions. Our suite of innovative products includes a secure Payment Gateway for seamless online transactions, robust Payouts solutions to streamline bulk payments, and a versatile Point of Sale (POS) system for efficient in-store transactions. At WowPe, we’re dedicated to providing user-friendly, scalable, and reliable solutions that empower businesses to grow and succeed in today’s fast-paced digital economy.
We are looking for a highly skilled Cloud Infrastructure & Cloud Network Engineer to design, build, and manage secure, scalable hybrid cloud environments at WowPe. This role will focus on cloud networking, hybrid connectivity, infrastructure automation, security, reliability, and performance across on-prem and cloud platforms (AWS/Azure). You will play a critical role in ensuring high availability, security, and performance of our fintech platforms.
A Day in the Life
- Design and review cloud network architectures (VPC/VNet, routing, segmentation)
- Troubleshoot latency, connectivity, VPN, and performance issues
- Work with DevOps and application teams to support deployments and scalability
- Automate infrastructure provisioning and security guardrails
- Monitor network health, traffic flow, and system performance
- Ensure security, compliance, and disaster recovery readiness
- Support hybrid connectivity between on-prem data centers and cloud environments.
Key Responsibilities
Cloud Networking & Hybrid Architecture
- Design and implement VPC/VNet architectures, subnetting, routing tables, NAT, gateways, and secure segmentation
- Build and manage hybrid connectivity between on-prem data centers and cloud using Site-to-Site VPN, ExpressRoute, Direct Connect
- Configure and manage Layer 4 & Layer 7 load balancers for high availability and traffic distribution
- Architect DNS, CDN, and edge networking strategies for low-latency global access
Security & Zero-Trust
- Implement Zero-Trust networking using security groups, NACLs, identity-aware proxies, mTLS
- Design secure access controls and network isolation
- Work closely with security teams to ensure PCI, ISO, SOC compliance readiness
Automation, Platform & Reliability
- Automate infrastructure provisioning using Infrastructure as Code (Terraform/ARM/CloudFormation)
- Build self-healing, auto-scaling architectures with health checks and fault tolerance
- Support and manage Kubernetes clusters (EKS/AKS) including networking and service communication
- Integrate infra with CI/CD pipelines for safe and frequent deployments
Operations & Observability
- Perform advanced troubleshooting for latency, packet loss, MTU, routing loops
- Implement and manage monitoring, logging, and observability (Prometheus, Grafana, ELK, Azure Monitor, CloudWatch)
- Design and maintain disaster recovery architectures, multi-region networking, and replication strategies
- Ensure high availability, performance optimization, and cost efficiency
Basic Qualifications & Skills
- 4+ years of experience in Cloud Infrastructure & Cloud Networking
- Strong hands-on experience with AWS and/or Azure
- Deep understanding of VPC/VNet, routing, NAT, gateways, load balancers, DNS
- Experience with Hybrid connectivity (VPN, ExpressRoute, Direct Connect)
- Solid knowledge of network security, firewalls, access control, segmentation
- Hands-on experience with Infrastructure as Code (Terraform preferred)
- Experience in monitoring, troubleshooting, and incident handling
- Strong understanding of high availability, DR, and performance optimization
Preferred Qualifications
- Experience in fintech, BFSI, or high-compliance environments
- Exposure to Zero-Trust architecture and security best practices
- Experience with Kubernetes (EKS/AKS), service mesh, microservices networking
- Familiarity with CI/CD pipelines and DevOps practices
- Knowledge of CDN, edge networking, and global traffic management
- Experience with compliance frameworks (PCI-DSS, ISO 27001, SOC2)
- Ability to design large-scale, resilient, production-grade architectures
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Title: AWS Cloud Engineer (Terraform & Monitoring Tools)
Experience: 8–10 Years
Location: Bangalore / Hyderabad
Joining: Immediate Joiners Preferred - 20 day serving
Job Description:
We are looking for an experienced AWS Cloud Engineer with 8+ years of experience in cloud infrastructure, automation, and monitoring. The ideal candidate should have strong expertise in AWS, Terraform, and at least one monitoring/observability tool such as Splunk, AppDynamics, or Grafana.
Required Skills:
- 8–10 years of IT experience with strong AWS cloud expertise
- Hands-on experience with Terraform for Infrastructure as Code (IaC)
- Experience with Splunk, AppDynamics, or Grafana (any one mandatory)
- Knowledge of cloud monitoring, logging, and performance optimization
- Experience with CI/CD pipelines and cloud infrastructure automation
- Strong troubleshooting, analytical, and communication skills
Job Title : SDE 3 – Infrastructure Platform Engineer
Experience : 5.5 to 8.5 Years
Number of Positions : 2
Employment Type : C2H (Contract to Hire)
Work Mode : Remote during contractual period → 5 Days WFO after conversion
Contract Duration : 3 Months
Post-Conversion Location : Pune
Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred
(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)
Role Overview :
We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.
The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.
Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.
Key Responsibilities :
- Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
- Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
- Manage and optimize Docker and Kubernetes workloads.
- Build internal platform tools to improve developer productivity and engineering efficiency.
- Implement monitoring, logging, alerting, and observability solutions.
- Participate in incident response, RCA, postmortems, and reliability improvements.
- Design and improve CI/CD pipelines and deployment automation.
- Contribute to system design, architecture discussions, scalability, security, and cost optimization.
- Collaborate with application, data, and product engineering teams.
Required Skills :
- 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
- Strong hands-on experience with AWS, GCP, or Azure.
- Strong expertise in Terraform / CloudFormation.
- Experience with CI/CD, Docker, and Kubernetes.
- Strong programming skills in at least one of:
- Python, Go, Java, or Ruby.
- Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
- Experience working with production infrastructure and highly available systems.
- Strong troubleshooting and problem-solving skills.
Nice to Have :
- Experience with SRE practices and production on-call ownership.
- Experience in fintech, payments, banking, or transaction-heavy systems.
- Knowledge of cloud security, compliance, or FinOps/cost optimization.
- Experience building internal developer platforms or productivity tools.
- Previous product company experience.
Interview Process :
Round 1 : Take-Home Coding Assignment – Submit within 48 hours
Round 2 : Coding Assignment Discussion – 1 Hour
Round 3 : Technical Managerial Round – 30 Minutes
Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.
Ideal Candidate :
Strong Platform / Infrastructure Engineer with hands-on experience in :
Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations
Pure DevOps profiles without strong coding and platform engineering experience are not preferred.
Cloud Expertise(Azure):
• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.
• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.
• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).
• Understanding on policies and security aspects of cloud.
Kubernetes & Helm:
• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.
• Understand of Kubernetes templates and its deployment.
• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.
• Implement Kubernetes best practices, including security, networking, and scaling.
• Concepts of Docker and Containers
CI/CD & Programming:
• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).
• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.
• Proven skill in python programming and concepts.
Monitoring and Observability:
Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.
• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.
Job Summary
We are looking for an experienced AWS Cloud Engineer with strong expertise in AWS infrastructure, deployment, migration, and cloud operations. The candidate should have hands-on experience managing AWS services such as EC2, EBS, S3, EFS, and FSx, along with infrastructure provisioning, migration, troubleshooting, and optimization.
Key Responsibilities
- Design, deploy, configure, and manage AWS infrastructure environments.
- Perform application and infrastructure migration to AWS.
- Provision and manage EC2 instances, including configuration, scaling, patching, and troubleshooting.
- Manage EBS volumes, snapshots, backups, and storage performance.
- Configure and administer S3 buckets, storage policies, lifecycle management, and access controls.
- Manage EFS for scalable shared file storage.
- Implement and manage Amazon FSx file systems based on application requirements.
- Monitor AWS infrastructure performance, availability, and capacity.
- Troubleshoot infrastructure, networking, storage, and deployment-related issues.
- Implement security best practices including IAM, security groups, encryption, and access controls.
- Support backup, disaster recovery, high availability, and business continuity requirements.
- Optimize AWS resources for performance, scalability, reliability, and cost.
Mandatory Skills
- Strong hands-on experience in AWS Cloud Infrastructure.
- Expertise in EC2, EBS, S3, EFS, and FSx.
- Experience in AWS deployment and migration projects.
- Strong knowledge of AWS networking concepts such as VPC, Subnets, Route Tables, Security Groups, and Load Balancers.
- Experience with IAM and AWS security best practices.
- Good knowledge of AWS monitoring and troubleshooting.
- Experience with cloud infrastructure automation using Terraform or CloudFormation is preferred.
- Strong Linux administration and troubleshooting skills.
- Good understanding of backup, disaster recovery, and high-availability concepts
About the Team
SecurITe’s mission is to build an Agentic-AI driven security platform that protects critical infrastructure from modern cyber threats. Our focus is on delivering highly performant, resilient, and intelligent network security systems that help defenders stay ahead of adversaries.
About the Role
We’re looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.
As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premise environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.
You’ll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security.
This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.
Why This Role Matters
Cybersecurity is undergoing a fundamental shift. AI is no longer an enhancement—it’s becoming the core engine of how detection, investigation, and response are executed. As our Platform Engineer, you will architect and build the infrastructure, automation, deployment, and operational systems that make this transformation real.
Your work will directly influence the scalability, reliability, and security of our AI-driven cybersecurity platform across both cloud and enterprise on-premise deployments. You’ll help establish the operational backbone that enables rapid innovation, secure product delivery, and resilient large-scale deployments in mission-critical environments.
This is a chance to solve novel technical challenges involving distributed systems, hybrid infrastructure, observability, automation, and secure software delivery while shaping how defenders outpace modern attackers.
What You’ll Do
● Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, security groups)
● Administer and harden AlmaLinux VMs across production, staging, and dev environments
● Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling)
● Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch
● Lead incident response — diagnosis, RCA, and post-incident documentation — with no dedicated ops team to escalate to
● Make and document build-vs-buy and architecture decisions as the product and team scale
● Work directly with founders/engineering to translate ambiguous asks into scoped technical plans
Impact You’ll Have
● Accelerate engineering velocity through scalable developer platforms and automation
● Improve deployment reliability, platform uptime, and operational efficiency
● Enable secure and scalable AI-driven cybersecurity workloads
● Reduce operational overhead through infrastructure automation and self-service systems
● Help establish enterprise-grade cloud and on-premise deployment capabilities
● Enhance product resiliency, observability, and operational excellence
● Shape the long-term platform architecture powering next-generation cybersecurity products
● Enable rapid and secure delivery of critical security innovations to customers
Required Experience
● 4+ years hands-on Linux administration (RHEL-family strongly preferred — AlmaLinux, CentOS, RHEL)
● Deep Linux internals: systemd, networking, storage/LVM, process/resource management, kernel-level troubleshooting
● Real AWS architecture experience — not just operating existing infra, but designing it (VPC, EC2, IAM, security groups, networking)
● Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to
● Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts — built tooling that runs unattended
● Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one)
● Clear, proactive communicator — documents decisions and explains reasoning without being asked
Required Skills & Qualifications
● Strong Linux system administration and troubleshooting skills
● Redhat certifications
● Strong understanding of networking fundamentals, security, and distributed systems
● Proficiency with Docker, and container orchestration
● Experience with Terraform, Ansible, or similar infrastructure automation tools
● Strong scripting or programming skills in Python, Bash, or Go
● Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry
● Understanding of platform security best practices and secure infrastructure design
● Familiarity with virtualization technologies and hybrid infrastructure environments
● Strong problem-solving and debugging abilities
● Excellent communication and collaboration skills
● Ability to thrive in fast-paced startup environments
Nice to Have
● Configuration management/automation at scale (Ansible, AWX, Terraform, or similar)
● Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog)
● Container experience (Docker; Kubernetes a plus but not core to this VM-based stack)
● Experience in a startup or small-team environment where infra was built from scratch
● Security/compliance exposure (vulnerability remediation, hardening, SSO/access control)
The Mindset
Problem Solver
You thrive on complex, ambiguous challenges and engineer elegant solutions.
Ownership-Driven
You take initiative, move fast, and deliver outcomes without hand-holding.
Continuous Learner
You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.
Startup DNA
You excel in fast-moving environments where priorities evolve and impact is immediate.
Job Description:
- Infrastructure Management: Design, implement, and manage scalable, reliable, and secure cloud infrastructure using AWS, GCP, and/or Azure.
- CI/CD Pipelines: Develop and maintain continuous integration and continuous deployment (CI/CD) pipelines to streamline the development lifecycle.
- Automation: Automate infrastructure provisioning, configuration management, and application deployment processes.
- Monitoring and Performance: Implement monitoring, logging, and alerting solutions to ensure system health, performance, and reliability.
- Security: Ensure the security of cloud infrastructure and applications, including identity management and compliance with industry standards.
- Collaboration: Work closely with client and development teams to integrate DevOps practices and deliver high-quality software.
- Documentation: Maintain comprehensive documentation of infrastructure, configurations, and processes.
- Innovation: Stay current with emerging technologies and industry trends, integrating them into the DevOps strategy as appropriate.
Qualifications:
- Education: Bachelor's degree in Computer Science, Information Technology, or a related field.
- Experience: 7 - 10 years of overall experience with relevant experience of at least 7 years in DevOps and served as a lead or senior engineer.
Position: Cloud Infrastructure & Migration SME/Tech Lead (L3) | Total IT Experience: 8-12 Years
Relevant Cloud Experience: 6+ Years
Job Summary:
We are seeking a highly experienced Cloud Infrastructure & Migration SME (L3) to design, implement, secure, and optimize enterprise-scale cloud and hybrid infrastructure environments. The candidate will serve as the primary SME responsible for cloud transformation, datacentre migration, infrastructure modernization, governance, security, and operational excellence across cloud and on-premises platforms.
The ideal candidate will possess deep expertise in Azure cloud technologies along with exposure to other public cloud platforms, cloud infrastructure, migration strategies, landing zone implementation, security frameworks, hybrid networking, disaster recovery, and automation. The role requires strong experience in assessing on-premises infrastructure and applications, planning migration and modernization initiatives, and delivering secure, scalable, resilient, and cost-effective cloud solutions across business-critical environments.
Key Responsibilities:
- Design, implement, and support secure, scalable, and highly available cloud and hybrid infrastructure environments, primarily focused on Microsoft Azure with exposure to AWS.
- Design and implement Landing Zones, cloud governance frameworks, management hierarchies, security controls, compliance standards, and operational guardrails.
- Lead enterprise storage transformation initiatives including Azure Storage services, data lifecycle management, and large-scale NAS-to-Blob/File Share migration programs.
- Lead cloud transformation and datacentre migration initiatives, including discovery, assessments, dependency mapping, migration planning, P2V, Virtual-to-Cloud and On-Premises-to-Cloud migrations, validation, cutover, and stabilization.
- Design and implement enterprise security and identity controls covering Microsoft Entra ID, Active Directory, identity governance, Zero Trust principles, security hardening, and compliance requirements.
- Design and support cloud networking, hybrid connectivity, network segmentation, private connectivity, DNS, routing, and secure access architectures.
- Build and manage cloud infrastructure platforms including compute, backup, disaster recovery, business continuity, high availability, and performance optimization to ensure operational resilience.
- Enable and support Azure database PaaS implementations, including migration readiness, Private Endpoint connectivity, security, and operational handover to database teams.
- Design, automate, and manage infrastructure using Infrastructure-as-Code (IaC), cloud automation, and DevOps best practices using Terraform, PowerShell, and Azure-native tooling.
- Implement monitoring, observability, alerting, capacity planning, and operational reporting capabilities to maintain infrastructure health, service reliability, and operational excellence.
- Provide L3 support, lead major incident resolution and RCAs, and drive continuous service improvement initiatives across cloud and infrastructure platforms.
- Collaborate with application, database, infrastructure, networking, security, and cloud teams to deliver reliable, secure, and cost-effective cloud solutions.
- Create and maintain architecture documentation, migration runbooks, standards, operational procedures, governance artefacts, and infrastructure readiness assessments.
Experience and Qualification Requirements
- Bachelor’s degree in computer science, Information Technology, Engineering, or a related discipline.
- 8-12 years of IT experience, including 6+ years of hands-on Cloud Infrastructure and Migration experience, primarily within Microsoft Azure environments.
- Proven experience leading cloud transformation, datacentre migration, infrastructure modernization, and cloud adoption programs across enterprise environments.
- Strong expertise in enterprise cloud and hybrid infrastructure, cloud architecture, Landing Zones, governance, security, networking, identity management, disaster recovery, and operational excellence.
- Experience performing application and infrastructure assessments, dependency mapping, workload rationalization, and migration planning for cloud transformation initiatives.
- Strong experience with Windows Server, Active Directory, Microsoft Entra ID, virtualization platforms, hybrid infrastructure, and cloud operations.
- Experience delivering P2V, Virtual-to-Cloud, and On-Premises-to-Cloud migrations, including migration validation, cutover planning, and operational stabilization.
- Hands-on experience with Infrastructure-as-Code (Terraform, Bicep, ARM Templates), cloud automation, and operational tooling.
- Experience supporting Azure SQL services, cloud storage platforms, hybrid connectivity, monitoring, backup, and resilience frameworks.
- Exposure to AWS or other public cloud platforms and hybrid/multi-cloud environments is desirable.
- Experience working within managed services, cloud operations, or large-scale transformation programs, with strong stakeholder management, troubleshooting, and technical leadership capabilities.
Required Skills & Certifications
- Microsoft Azure Administrator Associate (AZ-104) – Mandatory
- Microsoft Azure Solutions Architect Expert (AZ-305) - mandatory
- AWS Certified Solution Architect - Preferred
- HashiCorp Terraform Associate - Optional
Key Competencies
- Strong system integration, analytical and troubleshooting skills
- Ability to manage business-critical production environments
- Leadership and mentoring capability
- Excellent communication and stakeholder management skills
- Proactive approach toward automation and continuous improvement
Work/ Roaster Requirement:
- Willingness to work in UK / US aligned shifts based on project requirements.
- Participation in on-call support and major incident response activities.
- Ability to operate in a fast-paced managed services environment.
Ideal Candidate Profile
A hands-on Cloud Infrastructure & Migration SME with expertise in cloud transformation, datacentre migration, hybrid infrastructure, governance, security, and operational excellence. The ideal candidate combines deep cloud infrastructure knowledge with strong migration and modernization experience to deliver secure, resilient, and cost-effective enterprise cloud solutions.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Role: DevOps Infrastructure
Location: Pune
Experience: 8–12 years
Job Summary:
We are looking for an experienced DevOps Infrastructure professional with strong expertise in Microsoft Azure, cloud architecture, automation, and infrastructure engineering.
Key Responsibilities:
- Design and manage scalable, secure Azure infrastructure and platform solutions.
- Implement CI/CD pipelines using Azure DevOps or GitHub Actions.
- Automate infrastructure using Terraform, Bicep, or ARM and scripting with PowerShell/Azure CLI/Python/Bash.
- Work with Azure networking, Entra ID, RBAC/PIM, storage, backup, and disaster recovery.
- Support containerized workloads using Docker and AKS.
- Define cloud standards, governance, reference architectures, and security best practices.
- Implement monitoring, observability, DevSecOps, and policy-as-code practices.
- Collaborate with architecture, security, and application teams on hybrid/multi-cloud initiatives.
Must-Have Skills:
- 3+ years hands-on experience with Microsoft Azure.
- Strong DevOps experience with CI/CD, Git, and Infrastructure as Code.
- 2+ years in infrastructure design, platform engineering, or architecture.
- Strong understanding of Azure Well-Architected Framework, networking, IAM, security, and governance.
- Experience with Azure DevOps, Terraform/Bicep, Docker, and preferably AKS.
Preferred: Azure/Azure DevOps certifications, Terraform, ITIL, CISSP, multi-cloud/hybrid cloud, GitOps, and DevSecOps experience.
Senior Cloud Site Reliability Engineer (CSRE) – Azure
About Searce:
Searce is an AI-native, engineering-led modern technology consultancy that empowers
clients to futurify their businesses by delivering real, intelligent business outcomes. As a
trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,
data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"
cultural mindset and our proprietary evlos problem-solving framework, we eliminate
bureaucratic fluff to build working prototypes fast and scale enterprise production
environments intelligently. We don't just fix systems; we leverage multi-cloud technologies
to transform client operations into distinct competitive advantages.
Position Overview:
We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to
architect, secure, and stabilize next-generation hybrid and multi-cloud environments.
Operating at the intersection of infrastructure design, security compliance, and production
operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,
and AWS.
Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a
massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For
the Lead path, you will couple this deep engineering toolkit with stakeholder management
and mentorship to drive an elite operational culture.
Experience & Level Expectation:
Years of Experience: 3 to 10 years of intensive, hands-on production operations
experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.
Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced
triaging of infrastructure failures, and ownership of the CI/CD and deployment
lifecycles.
Intermediate level (5-10 Years): Expected to take architectural ownership, serve as
primary Incident Commander for complex outages, design cross-cloud governance
frameworks, and act as a reliable bridge between technical teams and client
leadership.
Key Responsibilities & Role Expectations:
Multi-Cloud Platforms & Orchestration: Design, configure, and maintain
production-grade Kubernetes clusters across major platforms (AKS).
Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant
isolation.
Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable
infrastructure components using Terraform or Crossplane. Standardize automated
environment provisioning to eliminate configuration drift across multi-branch
environments.
Incident Management & Reliability (SRE): Own and optimize the production on-call
rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing
Mean Time to Recovery (MTTR) through centralized log and metric correlation.
Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to
identify core architectural vulnerabilities and establish long-term fixes preventing
recurrence.
Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,
operating system patching strategies (Linux/Windows), database lifecycle updates,
and multi-region Disaster Recovery (DR) failover drills.
Core Core Operations & Legacy Integration: Manage enterprise-level hybrid
networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP
configurations) while effectively connecting cloud native services to legacy
infrastructures like Active Directory.
Security & Governance: Embed Zero Trust policies, secure secrets management
(Secrets Manager/Key Vault), and continuous vulnerability patching into the
automated SDLC pipeline.
Required Technical Skills:
- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.
- Containers & Orchestration
- Production-level management of GKE, AKS, and EKS.
- Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption







