Cutshort logo
For Employers
Searce Inc logo
Sr Engineer - Cloud Reliability Engineer
Sr Engineer - Cloud Reliability Engineer

Sr Engineer - Cloud Reliability Engineer at Searce Inc · Coimbatore · 3 - 8 years · Profitable · Posted 24 Sep 2026

Searce Inc's logo

Sr Engineer - Cloud Reliability Engineer

Mohammed Rabidheen's profile picture
Posted by Mohammed Rabidheen
3 - 8 yrs
Best in industry
Coimbatore
Skills
Windows Azure
AKS
DevOps
Microsoft Windows Azure

Senior Cloud Site Reliability Engineer (CSRE) – Azure


About Searce:

Searce is an AI-native, engineering-led modern technology consultancy that empowers

clients to futurify their businesses by delivering real, intelligent business outcomes. As a

trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,

data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"

cultural mindset and our proprietary evlos problem-solving framework, we eliminate

bureaucratic fluff to build working prototypes fast and scale enterprise production

environments intelligently. We don't just fix systems; we leverage multi-cloud technologies

to transform client operations into distinct competitive advantages.

Position Overview:

We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to

architect, secure, and stabilize next-generation hybrid and multi-cloud environments.

Operating at the intersection of infrastructure design, security compliance, and production

operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,

and AWS.

Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a

massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For

the Lead path, you will couple this deep engineering toolkit with stakeholder management

and mentorship to drive an elite operational culture.


Experience & Level Expectation:

Years of Experience: 3 to 10 years of intensive, hands-on production operations

experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.

Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced

triaging of infrastructure failures, and ownership of the CI/CD and deployment

lifecycles.

Intermediate level (5-10 Years): Expected to take architectural ownership, serve as

primary Incident Commander for complex outages, design cross-cloud governance

frameworks, and act as a reliable bridge between technical teams and client

leadership.


Key Responsibilities & Role Expectations:

Multi-Cloud Platforms & Orchestration: Design, configure, and maintain

production-grade Kubernetes clusters across major platforms (AKS).

Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant

isolation.

Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable

infrastructure components using Terraform or Crossplane. Standardize automated

environment provisioning to eliminate configuration drift across multi-branch

environments.

Incident Management & Reliability (SRE): Own and optimize the production on-call

rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing

Mean Time to Recovery (MTTR) through centralized log and metric correlation.

Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to

identify core architectural vulnerabilities and establish long-term fixes preventing

recurrence.

Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,

operating system patching strategies (Linux/Windows), database lifecycle updates,

and multi-region Disaster Recovery (DR) failover drills.

Core Core Operations & Legacy Integration: Manage enterprise-level hybrid

networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP

configurations) while effectively connecting cloud native services to legacy

infrastructures like Active Directory.

Security & Governance: Embed Zero Trust policies, secure secrets management

(Secrets Manager/Key Vault), and continuous vulnerability patching into the

automated SDLC pipeline.


Required Technical Skills:

- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.

- Containers & Orchestration

  • Production-level management of GKE, AKS, and EKS.
  • Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Searce Inc

Founded :
2004
Type :
Products & Services
Size :
1000-5000
Stage :
Profitable

About

What is ‘searce’


Searce means ‘a fine sieve’ & indicates ‘to refine, to analyze, to improve’. It signifies our way of working: To improve to the finest degree of excellence, ‘solving for better’ every time. Searcians are passionate improvers & solvers who love to question the status quo.

The primary purpose of all of us, at Searce, is driving intelligent, impactful & futuristic business outcomes using new-age technology. This purpose is driven passionately by HAPPIER people who aim to become better, everyday.

What we do

Searce is a modern tech consulting firm that empowers clients to futurify their businesses, leveraging Cloud, AI & Analytics.


  1. We are a category defining niche’ cloud-native technology consulting company, specializing in modernizing (improve, automate & transform) the full-scope of infra, app, process & work
  2. We partner with clients in their ‘beyond x’ journey to drive intelligent, impactful & futuristic business outcomes
  3. We are the most preferred tech partner of choice when it comes to ‘solving for better’ for the new-age tech startups & digital enterprises, leading disruption in their industries
  4. Our Service Offerings: We offer Advanced Cloud, Data & App Modernization, Cloud Consulting, Management & Improvement (DevOps, SysOps & Cloud Managed Services), Applied AI & Analytics services 
  5. As one of the top 5 niche’ full scope global partners for Google Cloud & a preferred partner for AWS, we are the most preferred ‘engineering-led’ tech company of choice when it comes to solving complex business problems.


Who we are


We are passionate improvers, solvers & futurists. Driven by our engineering excellence mindset, we care most about delivering intelligent, impactful & futuristic business outcomes. Searcians are motivated by continuous improvement & solving for better in everything we do.

At the core, a Searcian is self-driven to become better, everyday. In passionate pursuit of the finest degree of excellence we drive exceptional outcomes in everything we do.

We believe that trust is the most important value. We also believe that we need to ‘earn the trust’, everytime one engages with us. And earning trust for us is far more important than anything else. We aim to be the *most trusted* tech consulting partner for our clients.

We are HAPPIER at heart. Humble, Adaptable, Positive, Passionate, Innovative, Excellence focused, & Responsible. We live the HAPPIER Culture Code.

Being HAPPIER.



How we work


  1. Customers. Partners. Our aim is to build relationships with customers for life. And meaningfully improve the life of every customer.
  2. We do what we say. We say what we do. We are uncomfortably honest and transparent. Being genuine wins trust & makes people happier.
  3. Mistakes are encouraged. We make mistakes. Tons of those. Everyday. And we don’t mind apologizing to our juniors, peers or superiors. We are no ego-doers.
  4. Underpromise. Overdeliver. We work with a deep desire to go above and beyond in everything we do. Everytime.

So, If you are passionate about tech, future & what you read above (we really are!), apply here to experience the ‘Art of Possible’

Read more

Connect with the team

Profile picture
Hardik Parekh
Profile picture
Vishal Jarsania
Profile picture
Vamsi Krishna

Company social profiles

bloglinkedintwitterfacebook

Similar jobs (10)

company logo
Ganesh Ram
Posted by Ganesh Ram
Bengaluru (Bangalore), Mumbai, Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Hyderabad, Pune
7 - 10 yrs
₹15L - ₹20L / yr
CI/CD
skill iconKubernetes
helm
Terraform
yaml

Cloud Expertise(Azure):

• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.

• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.

• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).

• Understanding on policies and security aspects of cloud.


Kubernetes & Helm:

• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.

• Understand of Kubernetes templates and its deployment.

• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.

• Implement Kubernetes best practices, including security, networking, and scaling.

• Concepts of Docker and Containers


CI/CD & Programming:

• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).

• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.

• Proven skill in python programming and concepts.


Monitoring and Observability:

Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.

• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.

 

Read more
company logo
Pune
7 - 12 yrs
Best in industry
Google Cloud Platform (GCP)
Terraform
skill iconKubernetes
GKE
Reliability engineering
+1 more

About Searce

Searce is a global, AI-native, engineering-led technology consultancy and a Premier Google

Cloud Partner — recognized as the Google Cloud Workplace AI Transformation Partner of the

Year, APAC (2026). With 20+ years of experience and 3,000+ clients across 10+ countries, we

help businesses stay ahead of the cloud curve.


The Role

We're looking for a Lead Cloud Security & Reliability Engineer with deep GCP expertise to own

end-to-end cloud reliability and security forAPAC enterprise clients. As Lead, you'll set the architectural direction, mentor your squad, and drive measurable client outcomes across multi-

cloud environments.


What You'll Do

Own Client Delivery — Lead 24x7 GCP cloud operations forAPAC clients. Define SLO frameworks and ensure adherence.


Architect Solutions — Design scalable, secure GCP-primary architectures with multi-cloud awareness.


Drive Reliability — Lead incident response, RCA, and long-term remediation across production systems.


Mentor & Elevate — Coach and grow a squad of Senior CSREs.


Drive FinOps — Own cloud cost governance and optimization with quantified impact.


Be the Expert — Represent Searce's technical depth in global client conversations.


What We're Looking For

Experience

7–12 years total with 5+ years on GCP cloud infrastructure

Strong background in Cloud Managed Services / MSP environments

Proven experience leading a team in client-facing delivery

Multi-cloud exposure (AWS/Azure secondary) preferred


Technical Skills (Must-Have)

  • GCP: GKE, IAM, VPC, Cloud Monitoring, Stackdriver, KMS — demonstrated in work
  • experience
  • Kubernetes: GKE — production cluster management, Helm
  • IaC: Terraform — module-level, reusable frameworks
  • Observability: Prometheus, Grafana, Thanos or equivalent
  • Security: IAM, Zero-trust, DevSecOps, CSPM tools
  • Scripting: Python or Go
  • FinOps: GCP cost governance demonstrated


Nice to Have

  • GCP Professional Cloud Architect / Pro DevOps Engineer certification
  • AWS / Azure secondary experience
  • CKA (Certified Kubernetes Administrator)
  • ITIL / change management awareness
  • APAC client delivery experience


Why Searce?

🏆 Google Cloud Partner of the Year — APAC 2026

🌍 Work with APAC enterprise clients across multiple industries

🤖 AI-first, engineering-led culture

📈 Lead-level ownership with real career growth

🤝 HAPPIER values — Humble, Adaptable, Positive, Passionate, Innovative, Excellence,

Responsible

Read more
company logo
Carol Rangreji
Posted by Carol Rangreji
Bengaluru (Bangalore)
4 - 6 yrs
₹7L - ₹10L / yr
Microsoft Windows Azure
CI/CD
Google Cloud Storage
Cloud Computing

WowPe is a leading fintech company revolutionizing the way businesses handle financial transactions. Our suite of innovative products includes a secure Payment Gateway for seamless online transactions, robust Payouts solutions to streamline bulk payments, and a versatile Point of Sale (POS) system for efficient in-store transactions. At WowPe, we’re dedicated to providing user-friendly, scalable, and reliable solutions that empower businesses to grow and succeed in today’s fast-paced digital economy.


We are looking for a highly skilled Cloud Infrastructure & Cloud Network Engineer to design, build, and manage secure, scalable hybrid cloud environments at WowPe. This role will focus on cloud networking, hybrid connectivity, infrastructure automation, security, reliability, and performance across on-prem and cloud platforms (AWS/Azure). You will play a critical role in ensuring high availability, security, and performance of our fintech platforms.


A Day in the Life

  • Design and review cloud network architectures (VPC/VNet, routing, segmentation)
  • Troubleshoot latency, connectivity, VPN, and performance issues
  • Work with DevOps and application teams to support deployments and scalability
  • Automate infrastructure provisioning and security guardrails
  • Monitor network health, traffic flow, and system performance
  • Ensure security, compliance, and disaster recovery readiness
  • Support hybrid connectivity between on-prem data centers and cloud environments.


Key Responsibilities

Cloud Networking & Hybrid Architecture

  • Design and implement VPC/VNet architectures, subnetting, routing tables, NAT, gateways, and secure segmentation
  • Build and manage hybrid connectivity between on-prem data centers and cloud using Site-to-Site VPN, ExpressRoute, Direct Connect
  • Configure and manage Layer 4 & Layer 7 load balancers for high availability and traffic distribution
  • Architect DNS, CDN, and edge networking strategies for low-latency global access

Security & Zero-Trust

  • Implement Zero-Trust networking using security groups, NACLs, identity-aware proxies, mTLS
  • Design secure access controls and network isolation
  • Work closely with security teams to ensure PCI, ISO, SOC compliance readiness

Automation, Platform & Reliability

  • Automate infrastructure provisioning using Infrastructure as Code (Terraform/ARM/CloudFormation)
  • Build self-healing, auto-scaling architectures with health checks and fault tolerance
  • Support and manage Kubernetes clusters (EKS/AKS) including networking and service communication
  • Integrate infra with CI/CD pipelines for safe and frequent deployments

Operations & Observability

  • Perform advanced troubleshooting for latency, packet loss, MTU, routing loops
  • Implement and manage monitoring, logging, and observability (Prometheus, Grafana, ELK, Azure Monitor, CloudWatch)
  • Design and maintain disaster recovery architectures, multi-region networking, and replication strategies
  • Ensure high availability, performance optimization, and cost efficiency


Basic Qualifications & Skills

  • 4+ years of experience in Cloud Infrastructure & Cloud Networking
  • Strong hands-on experience with AWS and/or Azure
  • Deep understanding of VPC/VNet, routing, NAT, gateways, load balancers, DNS
  • Experience with Hybrid connectivity (VPN, ExpressRoute, Direct Connect)
  • Solid knowledge of network security, firewalls, access control, segmentation
  • Hands-on experience with Infrastructure as Code (Terraform preferred)
  • Experience in monitoring, troubleshooting, and incident handling
  • Strong understanding of high availability, DR, and performance optimization


Preferred Qualifications

  • Experience in fintech, BFSI, or high-compliance environments
  • Exposure to Zero-Trust architecture and security best practices
  • Experience with Kubernetes (EKS/AKS), service mesh, microservices networking
  • Familiarity with CI/CD pipelines and DevOps practices
  • Knowledge of CDN, edge networking, and global traffic management
  • Experience with compliance frameworks (PCI-DSS, ISO 27001, SOC2)
  • Ability to design large-scale, resilient, production-grade architectures
Read more
MNC
MNC
Agency job
via by aafia parveen
Bengaluru (Bangalore)
7 - 11 yrs
₹15L - ₹18L / yr
skill iconAmazon Web Services (AWS)
skill iconKubernetes
Linux/Unix
openshift

AWS / Kubernetes / OpenShift / Linux –

Location: Bangalore

Experience: 7–10 Years

Mandatory Skills:

  • Strong hands-on experience with AWS
  • Experience in Kubernetes & OpenShift
  • Strong knowledge of Linux administration
  • Experience with Docker & containerization
  • Knowledge of CI/CD pipelines and DevOps practices
  • Troubleshooting, monitoring, and deployment experience

Role: Cloud/DevOps Engineer – AWS, Kubernetes & OpenShift

Read more
MNC
MNC
Agency job
via by Gauri Naik
Remote only
10 - 18 yrs
Best in industry
Landing page optimization
skill iconKubernetes
Terraform
CI/CD

🚀 Hiring: Senior Azure Platform Engineer | Azure | Terraform | Kubernetes


📍 Location: India

💼 Employment: Full-Time | Long-Term

🎯 Experience: 10+ Years | 8+ Years Hands-on Azure


🔑 What We’re Looking For

▪️ Strong expertise in Azure Cloud Platform & Landing Zone Architecture

▪️ Advanced hands-on experience with Terraform & reusable IaC modules

▪️ Expertise in Kubernetes, Helm & container platforms

▪️ Strong understanding of Azure Networking, Entra ID, RBAC & Azure Policy

▪️ Experience with GitHub, GitHub Actions / Azure Pipelines & CI/CD

▪️ Hands-on GitOps & ArgoCD experience

▪️ Strong observability skills with Prometheus, Grafana & Azure Monitor

▪️ Experience designing and supporting microservices architectures

▪️ Knowledge of security, policy-as-code and IaC security scanning

▪️ Exposure to Azure Arc / Azure Local / Azure Stack HCI is a plus

Read more
company logo
Remote only
10 - 15 yrs
₹30L - ₹35L / yr
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Generative AI
Implementation
System deployment
+9 more

Job Description: Lead - Cloud Engineering (AWS / Azure)

Role Title: Lead - Cloud Engineering

Experience Level: 10+ Years

Domain Focus: Healthcare AI & Cloud Infrastructure

Location: Remote

Job Overview

We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.

The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.

Key Responsibilities

Cloud Architecture & Migration

  • Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
  • Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
  • Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.

Security & Healthcare Compliance

  • Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
  • Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.

Leadership & Team Management

  • Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
  • Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
  • Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.

Operations & FinOps

  • Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
  • Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.

Key Requirements

  • Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
  • Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
  • DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
  • Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
  • AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
  • Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).


Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
company logo
Daniel Castellanos
Posted by Daniel Castellanos
Remote only
1 - 10 yrs
₹10L - ₹50L / yr
skill iconKubernetes
ArgoCD
Exim
Postfix

We are looking for an experienced DevOps Engineer to take ownership of production infrastructure, cloud environments, Kubernetes platforms, and infrastructure automation. This is a hands-on role for someone who enjoys solving complex infrastructure challenges and is comfortable being responsible for systems in production.


Key Responsibilities

  • Own and operate production infrastructure, including participating in an on-call rotation and responding to production incidents.
  • Design, operate, and continuously improve Kubernetes clusters in production.
  • Manage and automate infrastructure using Infrastructure as Code, primarily with Terraform.
  • Build, maintain, and optimise cloud infrastructure across AWS, GCP, or Azure.
  • Work extensively with Linux, including system administration, networking, troubleshooting, and system-level configuration.
  • Manage production deployment and GitOps workflows using ArgoCD.
  • Improve infrastructure reliability, scalability, security, monitoring, and operational efficiency.
  • Troubleshoot complex production issues and drive problems through to resolution.
  • Develop automation and processes that reduce manual operational work.

Essential Requirements

  • 4+ years of hands-on experience operating production infrastructure, with personal ownership and responsibility for live systems, including on-call experience.
  • Deep, hands-on Kubernetes experience — you must have operated and managed Kubernetes clusters, rather than simply deploying applications onto clusters managed by another team.
  • Strong experience with Infrastructure as Code, with Terraform strongly preferred. Experience with Pulumi or CloudFormation is also considered.
  • Strong experience with at least one major cloud platform, ideally AWS. Strong GCP or Azure experience is also welcome, provided you are willing to work with AWS.
  • Strong Linux skills and confidence working from the command line, including networking, troubleshooting, system configuration, and performance issues.
  • Production experience with ArgoCD and GitOps-based deployment workflows.
  • Strong troubleshooting and problem-solving skills, with the ability to take ownership of production incidents and infrastructure issues.


Nice to Have

Experience with email infrastructure would be a strong advantage, particularly:

  • Exim
  • IMAP / SMTP
  • Postfix
  • Dovecot
  • General mail server administration and maintenance


Read more
company logo
Ashish Singh
Posted by Ashish Singh
Remote only
0 - 1 yrs
₹12000 - ₹18000 / mo
CI/CD
DevOps

Build production-grade cloud infrastructure that powers enterprise applications with cutting-edge DevOps practices.


What you'll do:

  • Design CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI)
  • Containerize apps with Docker, deploy on Kubernetes clusters
  • Manage infrastructure as code (Terraform, CloudFormation)
  • Set up monitoring (Prometheus, Grafana, ELK Stack)
  • Cloud migrations (AWS EC2, EKS, RDS → GCP equivalent)
  • Optimize costs and performance for live production systems

What we need:

  • Basic Python/Bash scripting
  • Docker basics, Git workflows
  • Cloud exposure (AWS/GCP/Azure free tier projects)
  • Problem-solving mindset, eagerness to learn

Real impact:

  • Deploy apps used by 1000+ daily users
  • Work with senior DevOps engineers on client deliverables
  • Build portfolio for FAANG-level interviews


Read more
company logo
Priya Rawat
Posted by Priya Rawat
Gurugram
4 - 5 yrs
₹8L - ₹10L / yr
RCA
SLA
skill icongrafana
ELKI
SOP

About the Role


We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.


Key Responsibilities


  • Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
  • Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
  • Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
  • Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
  • Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
  • Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
  • Participate in incident/severity calls, ensuring clear communication and coordination
  • Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting


Required Skills & Experience


  • Strong understanding of Linux/Unix systems for application support
  • Hands-on experience troubleshooting applications in staging and production environments
  • Ability to monitor system performance and identify root causes using logs and metrics
  • Experience working with Kubernetes and microservices-based architectures
  • Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
  • Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
  • Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
  • Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
  • Experience with ticketing and documentation tools such as Jira and Confluence
  • Minimum 4+ years of experience in application support or reliability engineering


Education & Certifications


  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Relevant certifications (Cloud, Kubernetes, Microservices) are a plus


Work Schedule


  • Willingness to work in a 24x7 environment, including weekends and on-call rotations
Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos