azure cloud engineer at MNC · Bengaluru (Bangalore), Hyderabad · 7 - 14 years · ₹2L - ₹15L / yr · Posted 23 Sep 2026

Job Description
We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.
Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and services.
- Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
- Configure dashboards, alerts, health rules, and monitoring metrics.
- Perform log analysis, troubleshooting, and root-cause analysis for production issues.
- Implement observability solutions for applications, cloud infrastructure, and services.
- Automate monitoring and operational activities using scripting.
- Support incident, problem, and change management processes.
- Collaborate with development, DevOps, and SRE teams to improve system reliability.
- Maintain monitoring standards, documentation, and operational procedures.
Primary Skills
- Microsoft Azure
- Splunk
- Grafana
- AppDynamics
- Cloud Monitoring & Observability
- Application Performance Monitoring (APM)
- Log Analysis & Troubleshooting
Secondary Skills
- Azure Monitor / Log Analytics
- Azure VMs, Storage, Networking
- Linux
- Python / PowerShell / Shell Scripting
- CI/CD
- Git
- ITIL / ServiceNow

Similar jobs (10)
Job Description
We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.
Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and services.
- Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
- Configure dashboards, alerts, health rules, and monitoring metrics.
- Perform log analysis, troubleshooting, and root-cause analysis for production issues.
- Implement observability solutions for applications, cloud infrastructure, and services.
- Automate monitoring and operational activities using scripting.
- Support incident, problem, and change management processes.
- Collaborate with development, DevOps, and SRE teams to improve system reliability.
- Maintain monitoring standards, documentation, and operational procedures.
Primary Skills
- Microsoft Azure
- Splunk
- Grafana
- AppDynamics
- Cloud Monitoring & Observability
- Application Performance Monitoring (APM)
- Log Analysis & Troubleshooting
Secondary Skills
- Azure Monitor / Log Analytics
- Azure VMs, Storage, Networking
- Linux
- Python / PowerShell / Shell Scripting
- CI/CD
- Git
- ITIL / ServiceNow
Cloud Expertise(Azure):
• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.
• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.
• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).
• Understanding on policies and security aspects of cloud.
Kubernetes & Helm:
• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.
• Understand of Kubernetes templates and its deployment.
• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.
• Implement Kubernetes best practices, including security, networking, and scaling.
• Concepts of Docker and Containers
CI/CD & Programming:
• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).
• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.
• Proven skill in python programming and concepts.
Monitoring and Observability:
Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.
• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.
Observability Engineer (AppDynamics)
Hyderabad
Exp: 8+years of exp
Mandate Skills: AppDynamics, Splunk, Python (Scripting knowledge)
Primary Skill Set:
- AppDynamics administration
- SPLOC administration
- Enterprise monitoring and observability
- Application performance monitoring
- Alerting and event management
- Monitoring strategy and design
- Platform configuration and governance
- Python scripting
- Monitoring automation
Secondary Skills:
- Glassbox monitoring
- Customer journey observability
- GenAI concepts for operations
- Log, metric, and trace telemetry
- Dashboarding and visualization
🚀 Hiring: Senior Azure Platform Engineer | Azure | Terraform | Kubernetes
📍 Location: India
💼 Employment: Full-Time | Long-Term
🎯 Experience: 10+ Years | 8+ Years Hands-on Azure
🔑 What We’re Looking For
▪️ Strong expertise in Azure Cloud Platform & Landing Zone Architecture
▪️ Advanced hands-on experience with Terraform & reusable IaC modules
▪️ Expertise in Kubernetes, Helm & container platforms
▪️ Strong understanding of Azure Networking, Entra ID, RBAC & Azure Policy
▪️ Experience with GitHub, GitHub Actions / Azure Pipelines & CI/CD
▪️ Hands-on GitOps & ArgoCD experience
▪️ Strong observability skills with Prometheus, Grafana & Azure Monitor
▪️ Experience designing and supporting microservices architectures
▪️ Knowledge of security, policy-as-code and IaC security scanning
▪️ Exposure to Azure Arc / Azure Local / Azure Stack HCI is a plus
- Strong hands-on experience in Microsoft Azure Cloud.
- Good understanding of Azure services such as Compute, Storage, Event Hub, Event Subscription, Storage Queue, and PaaS services.
- Basic understanding of Azure AI Foundry and AI-related Azure service setup.
- Good Azure networking basics: VNet, subnet, routing, and basic troubleshooting.
- Strong knowledge of Terraform, especially:
- Terraform state
- plan / apply
- troubleshooting failures
- migration risks
- Terraform Enterprise concepts
- Strong Python coding capability, not just basic scripting.
- Experience using Python for API integration, automation, JSON/YAML handling, and internal tooling.
- Good understanding of CI/CD pipelines.
- Ability to troubleshoot pipeline failures.
- Comfortable with YAML and JSON.
- Ability to troubleshoot Azure infrastructure/platform issues.
- Ability to collect logs/evidence and coordinate with network/app/Microsoft support teams.
- Basic awareness of agentic AI / LLM concepts.
- Awareness of security and cost best practices.
Good to Have Skills
- Hands-on experience with Harness.
- Hands-on experience with Terraform Enterprise.
- Exposure to LangGraph / LangChain.
- Exposure to agentic AI workflows or skill creation.
- Exposure to Claude or enterprise LLM integrations.
- Knowledge of Azure ML Workspace, model registry, and managed endpoints.
- MLOps / LLMOps knowledge.
- FinOps / Azure cost optimization experience.
- Azure certifications: AZ-104, AZ-305, AZ-400, AZ-500.
Screening Priority:
Azure Cloud + Terraform + Python Coding + CI/CD Troubleshooting + YAML/JSON + Basic Agentic AI Awareness
Work Mode: WORK FROM OFFICE
Interview Mode: Virtual
Role Descriptions:
Exp Range: 5 - 8 years
City Locations: Bengaluru
*Key Responsibilities*
Devops Developer (Azure)
Azure L2 Support | Specialist | Infrastructure Operations-260006F2- Certificate & DNS Management- Monitoring & Observability- Azure Operations & Compliance- Migration: Open Cloud TnC- Reliability & Operations Excellence (OPS)
Profile – 5 to 8+ years in Cloud/Infrastructure Operations- Strong hands-on Azure skills- Certificates/PKI experience- DNS management- Grafana| Health Checks| Azure Monitor- IaC automation (Terraform/Bicep)- Networking fundamentals Nice-to-Have- Azure security and compliance- AKS monitoring basics- FinOps practices- French reading proficiency- Certifications: AZ-104| AZ-305| etc.
Desire candidate
- Candidate should have valid PF.
Senior Cloud Site Reliability Engineer (CSRE) – Azure
About Searce:
Searce is an AI-native, engineering-led modern technology consultancy that empowers
clients to futurify their businesses by delivering real, intelligent business outcomes. As a
trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,
data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"
cultural mindset and our proprietary evlos problem-solving framework, we eliminate
bureaucratic fluff to build working prototypes fast and scale enterprise production
environments intelligently. We don't just fix systems; we leverage multi-cloud technologies
to transform client operations into distinct competitive advantages.
Position Overview:
We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to
architect, secure, and stabilize next-generation hybrid and multi-cloud environments.
Operating at the intersection of infrastructure design, security compliance, and production
operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,
and AWS.
Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a
massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For
the Lead path, you will couple this deep engineering toolkit with stakeholder management
and mentorship to drive an elite operational culture.
Experience & Level Expectation:
Years of Experience: 3 to 10 years of intensive, hands-on production operations
experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.
Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced
triaging of infrastructure failures, and ownership of the CI/CD and deployment
lifecycles.
Intermediate level (5-10 Years): Expected to take architectural ownership, serve as
primary Incident Commander for complex outages, design cross-cloud governance
frameworks, and act as a reliable bridge between technical teams and client
leadership.
Key Responsibilities & Role Expectations:
Multi-Cloud Platforms & Orchestration: Design, configure, and maintain
production-grade Kubernetes clusters across major platforms (AKS).
Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant
isolation.
Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable
infrastructure components using Terraform or Crossplane. Standardize automated
environment provisioning to eliminate configuration drift across multi-branch
environments.
Incident Management & Reliability (SRE): Own and optimize the production on-call
rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing
Mean Time to Recovery (MTTR) through centralized log and metric correlation.
Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to
identify core architectural vulnerabilities and establish long-term fixes preventing
recurrence.
Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,
operating system patching strategies (Linux/Windows), database lifecycle updates,
and multi-region Disaster Recovery (DR) failover drills.
Core Core Operations & Legacy Integration: Manage enterprise-level hybrid
networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP
configurations) while effectively connecting cloud native services to legacy
infrastructures like Active Directory.
Security & Governance: Embed Zero Trust policies, secure secrets management
(Secrets Manager/Key Vault), and continuous vulnerability patching into the
automated SDLC pipeline.
Required Technical Skills:
- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.
- Containers & Orchestration
- Production-level management of GKE, AKS, and EKS.
- Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption
Required Technical Skill Set
MS Azure, M365, PowerShell
No of Requirements
1
Desired Experience Range
5 to 10 Years
Location of Requirement
Bangalore/Hydrabad
Desired Competencies (Technical/Behavioral Competency)
Must-Have
· Hands-on experience in Automating solutions in Power shell at L3 level
· Experience in monitoring, troubleshooting, and executing Azure Runbooks for operational automation
· Working knowledge of Microsoft 365 Groups management and support
· String Experience in M365 License Management
· Strong understanding of Agile methodologies, with experience working in Scrum or Kanban environments
· Knowledge of Azure monitoring and diagnostics tools (Log Analytics, Application Insights, Alerts)
· Good understanding of application security, reliability, and performance optimization in cloud environments
· Excellent problem-solving, collaboration, and communication skills
· Proficiency in DevOps practices, including CI/CD pipeline implementation and maintenance (Azure DevOps or GitHub Actions)
Good-to-Have
· AZ-204
· Good understanding of M365 Security and Compliance features, including retention, audit, and data governance policies
Azure DevOps Engineer
Experience
5-10 years of hands-on experience in Platform Engineering, Cloud Engineering, SRE, DevOps or Infrastructure Engineering roles.
Priority 1 – Must Have (Hands-On)
Azure Cloud Platform
· Strong hands-on experience supporting Azure workloads in production environments.
· Experience designing, building and supporting Azure infrastructure using Terraform.
· Good understanding of Azure networking and connectivity patterns.
· Experience supporting:
Ø AKS
Ø Application Gateway
Ø Azure Traffic Manager
Ø Key Vault
Ø Azure Monitor / Log Analytics
Ø Managed Identities
Ø Service Principals
Ø Private Endpoints
Ø VNets, NSGs and Route Tables
Kubernetes / AKS
- Strong practical experience operating and supporting AKS.
- Ability to troubleshoot:
Ø Pod failures
Ø Ingress issues
Ø DNS issues
Ø SSL/TLS certificate issues
Ø Network routing issues
Ø Performance and availability incidents
- Experience with:
Ø Helm
Ø Ingress Controllers
Ø Cluster upgrades
Ø Scaling
Ø Monitoring
Terraform
· Strong hands-on experience writing and maintaining Terraform.
· Experience creating reusable modules.
· Experience managing:
· State files
· Remote backends
· Environment promotion
· Infrastructure lifecycle
Linux & Scripting
· Strong Linux administration fundamentals.
· Practical experience troubleshooting production issues.
· Bash scripting mandatory.
· Python desirable.
Application Support / Troubleshooting
Must be comfortable supporting business applications end-to-end.
Priority 2 – Highly Desirable
GitHub & DevOps Platform
Hands-on experience with:
· GitHub Enterprise
· GitHub Actions
· Shared workflows
· Reusable pipelines
· Repository onboarding
· Branch protections
· GitHub security features
Experience supporting:
· Runner issues
· Disk space issues
· Network connectivity issues
· Dependency failures
· Self-hosted runners lifecycle management
API Management
Pipeline failures
GitOps
Experience with:
· ArgoCD
· GitOps deployment models
· Kubernetes deployment automation
Monitoring & Observability
Experience working with:
· Prometheus
· Grafana
· Azure Monitor
· Log Analytics
· Application Insights
We are looking for a hands-on Azure DevOps & Infrastructure Engineer to manage enterprise Azure environments, automate infrastructure, build CI/CD pipelines, and support cloud modernization initiatives.
🔑 Key Responsibilities
• Design, implement & support CI/CD pipelines and deployment automation
• Perform TFS / Azure DevOps Server → Azure DevOps Services migration
• Build & manage Azure infrastructure using Infrastructure as Code (IaC)
• Support Azure cloud infrastructure, networking, security, monitoring & production environments
• Drive cloud migration, application modernization & platform transformation initiatives
• Work with Terraform, containerization, DevSecOps & cloud automation
• Troubleshoot infrastructure/deployment issues and support incident & on-call operations
• Create technical documentation, runbooks and deployment standards
🎯 Mandatory Requirements
✅ AZ-204 – Azure Developer Associate – Mandatory
✅ 5–8 years total IT experience
✅ 3+ years hands-on Azure DevOps / Azure Infrastructure experience
✅ Strong experience in CI/CD & Azure DevOps
✅ Hands-on experience with TFS / Azure DevOps Server to Azure DevOps Services migration
✅ Strong understanding of Azure cloud infrastructure & cloud operations
✅ Experience with IaC, provisioning & infrastructure automation
✅ Exposure to cloud/data migration and application modernization
✅ Experience supporting Azure networking, security & monitoring
✅ Experience in production support / managed services / cloud operations
✅ Willingness to work UK/US shifts and participate in on-call support
⭐ Preferred / Optional
• AZ-400 – Microsoft Certified DevOps Engineer Expert
• AZ-104 – Azure Administrator Associate
• HashiCorp Terraform Associate
• Container platforms, orchestration & DevSecOps experience
🎓 Education: Bachelor's degree in Computer Science / IT / Engineering or related field.







