Cutshort logo
For Employers
Fynd logo
Site Reliability Engineer - 2/3
at Fynd
Site Reliability Engineer - 2/3

Site Reliability Engineer - 2/3 at Fynd · Mumbai · 3 - 8 years · Profitable · Posted 23 Jun 2025

Fynd's logo

Site Reliability Engineer - 2/3

Akshata Kadam's profile picture
Posted by Akshata Kadam
3 - 8 yrs
Best in industry
Mumbai
Skills
Google Cloud Platform (GCP)
skill iconAmazon Web Services (AWS)

Fynd is India’s largest omnichannel platform and a multi-platform tech company specializing in retail technology and products in AI, ML, big data, image editing, and the learning space. It provides a unified platform for businesses to seamlessly manage online and offline sales, store operations, inventory, and customer engagement. Serving over 2,300 brands, Fynd is at the forefront of retail technology, transforming customer experiences and business processes across various industries.


What will you do at Fynd?

  • Run the production environment by monitoring availability and taking a holistic view of system health.
  • Improve reliability, quality, and time-to-market of our suite of software solutions
  • Be the 1st person to report the incident.
  • Debug production issues across services and levels of the stack.
  • Envisioning the overall solution for defined functional and non-functional requirements, and being able to define technologies, patterns and frameworks to realise it.
  • Building automated tools in Python / Java / GoLang / Ruby etc.
  • Help Platform and Engineering teams gain visibility into our infrastructure.
  • Lead design of software components and systems, to ensure availability, scalability, latency, and efficiency of our services.
  • Participate actively in detecting, remediating and reporting on Production incidents, ensuring the SLAs are met and driving Problem Management for permanent remediation.
  • Participate in on-call rotation to ensure coverage for planned/unplanned events.
  • Perform other task like load-test & generating system health reports.
  • Periodically check for all dashboards readiness.
  • Engage with other Engineering organizations to implement processes, identify improvements, and drive consistent results.
  • Working with your SRE and Engineering counterparts for driving Game days, training and other response readiness efforts.
  • Participate in the 24x7 support coverage as needed Troubleshooting and problem-solving complex issues with thorough root cause analysis on customer and SRE production environments
  • Collaborate with Service Engineering organizations to build and automate tooling, implement best practices to observe and manage the services in production and consistently achieve our market leading SLA.
  • Improving the scalability and reliability of our systems in production.
  • Evaluating, designing and implementing new system architectures.


Some specific Requirements:

  • B.E./B.Tech. in Engineering, Computer Science, technical degree, or equivalent work experience
  • At least 3 years of managing production infrastructure. Leading / managing a team is a huge plus.
  • Experience with cloud platforms like - AWS, GCP.
  • Experience developing and operating large scale distributed systems with Kubernetes, Docker and and Serverless (Lambdas)
  • Experience in running real-time and low latency high available applications (Kafka, gRPC, RTP)
  • Comfortable with Python, Go, or any relevant programming language.
  • Experience with monitoring alerting using technologies like Newrelic / zybix /Prometheus / Garafana / cloudwatch / Kafka / PagerDuty etc.
  • Experience with one or more orchestration, deployment tools, e.g. CloudFormation / Terraform / Ansible / Packer / Chef.
  • Experience with configuration management systems such as Ansible / Chef / Puppet.
  • Knowledge of load testing methodologies, tools like Gating, Apache Jmeter.
  • Work your way around Unix shell.
  • Experience running hybrid clouds and on-prem infrastructures on Red Hat Enterprise Linux / CentOS
  • A focus on delivering high-quality code through strong testing practices.


What do we offer?


Growth

Growth knows no bounds, as we foster an environment that encourages creativity, embraces challenges, and cultivates a culture of continuous expansion. We are looking at new product lines, international markets and brilliant people to grow even further. We teach, groom and nurture our people to become leaders. You get to grow with a company that is growing exponentially.

Flex University: We help you upskill by organising in-house courses on important subjects

Learning Wallet: You can also do an external course to upskill and grow, we reimburse it for you.


Culture

Community and Team building activities

Host weekly, quarterly and annual events/parties.


Wellness

Mediclaim policy for you + parents + spouse + kids

Experienced therapist for better mental health, improve productivity & work-life balance 


We work from the office 5 days a week to promote collaboration and teamwork. Join us to make an impact in an engaging, in-person environment!

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Fynd

Founded :
2012
Type :
Product
Size :
1000-5000
Stage :
Profitable

About

Fynd, India’s largest omni channel platform and multi-platform tech company, pioneers retail tech and products in AI, ML, big data ops, gaming+crypto, image editing, and the learning space. Founded in 2012 by three IIT Bombay alumni: Farooq Adam, Harsh Shah, and Sreeraman MG, Fynd is headquartered in Mumbai. With over 1000 brands under management, more than 10k stores, and servicing 23k+ pin codes, Fynd collaborates with major retail giants like Reliance, working on projects for Jio, Reliance Retail, and Reliance Digital, among others

Read more

Tech stack

skill iconPython
skill iconNodeJS (Node.js)
skill iconJavascript
skill iconMongoDB
skill iconReact.js
MySQL
Google Cloud Platform (GCP)

Connect with the team

Profile picture
Akshata Kadam
Profile picture
Kushan Shah
Profile picture
Farooq Adam

Company social profiles

bloginstagrampinterestlinkedintwitterfacebook

Similar jobs (10)

Gurugram
3 - 9 yrs
Best in industry
DevOps

Site Reliability Engineer (SRE) / DevOps Engineer (Walk-In Drive)

Location: Gurgaon Experience: 3–6 Years


About the Role

We are looking for a hands-on Site Reliability Engineer (SRE) / DevOps Engineer with strong programming and automation skills.

The role will initially involve development and automation work, helping the engineer build a strong understanding of the applications and platform. Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, observability, production reliability, incident management, and operational automation.

The ideal candidate should be comfortable working with both application code and production systems and should use engineering and automation to improve reliability and reduce manual effort.

Key Responsibilities

· Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.

· Write clean, maintainable, testable, and production-ready code.

· Participate in code reviews, debugging, testing, and technical discussions.

· Build, maintain, and improve CI/CD pipelines and automated deployment processes.

· Work with Docker and Kubernetes for application deployment and operations.

· Support on prem and cloud-based application and infrastructure deployments.

· Maintain reliable, scalable, secure, and highly available production environments.

· Implement and manage monitoring, logging, alerting, and observability solutions.

· Contribute to defining and tracking SLIs, SLOs, and error budgets.

· Troubleshoot application and production issues and perform Root Cause Analysis (RCA).

· Identify recurring operational problems and address them through automation and engineering improvements.

· Support incident response, change management, deployment governance, and disaster recovery practices.

· Maintain runbooks, SOPs, incident documentation, and technical documentation.

· Collaborate with Engineering, Product, Platform, Security, Operations, and external teams.

Technical Skills

· Strong hands-on experience with Python for development and automation.

· Experience developing scripts, APIs, integrations, utilities, or backend services.

· Good understanding of software engineering principles, debugging, logging, testing, and exception handling.

· Experience with REST APIs, JSON, Git, pull requests, and code reviews.

· Strong knowledge of Linux/Unix environments and basic Windows administration.

· Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.

· Experience with at least one cloud platform: AWS, Azure, or GCP.

· Hands-on experience with Docker and Kubernetes.

· Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.

· Familiarity with Infrastructure-as-Code tools such as Terraform is preferred.

· Experience with monitoring and observability tools such as Grafana, Prometheus, Power BI, or equivalent.

· Ability to analyze logs, metrics, alerts, and traces for troubleshooting.

· Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.

· Experience with JIRA, ServiceNow, and Confluence is desirable.

Preferred Experience

· 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, or related roles.

· Strong Python development or automation experience.

· Experience supporting applications across development, deployment, and production environments.

· Exposure to cloud-native, distributed, or production-grade systems.

· Understanding of security and compliance best practices.

· Familiarity with AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools.

Soft Skills

· Strong analytical and troubleshooting skills.

· Engineering and automation mindset.

· Good written and verbal communication skills.

· Effective cross-functional collaboration.

· Ownership-driven approach to problem solving.

· Ability to remain structured during production incidents.

Read more
company logo
Bhattacharjee Akash
Posted by Bhattacharjee Akash
Bengaluru (Bangalore), Chennai, Mumbai, Hyderabad, Pune, Gurugram
3 - 10 yrs
₹12L - ₹35L / yr
Linux/Unix
skill iconKubernetes
Monitoring
skill iconDocker
skill iconAmazon Web Services (AWS)
+4 more



We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.



Key Responsibilities

  • Own reliability, availability, and performance of production services, including on-call rotation
  • Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
  • Automate deployments, scaling, and operational tasks to reduce manual toil
  • Manage containerized workloads on Kubernetes and cloud infrastructure
  • Design and maintain CI/CD pipelines for safe, frequent releases
  • Lead incident response and drive blameless postmortems with clear follow-ups
  • Perform capacity planning, performance tuning, and cost optimization
  • Define and track SLIs/SLOs and error budgets with product teams


Requirements

  • 3+ years in SRE, DevOps, or production-focused engineering
  • Strong Linux administration and hands-on Kubernetes experience
  • Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
  • Cloud experience with AWS, GCP, or Azure
  • CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
  • Proficient scripting in Python and/or Bash


Nice to have

  • Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
  • Background in high-traffic or distributed systems
Read more
company logo
Harsha Mehrotra
Posted by Harsha Mehrotra
Remote only
3 - 5 yrs
Best in industry
CI/CD
skill iconPostgreSQL
skill iconPython
skill iconAmazon Web Services (AWS)
Artificial Intelligence (AI)
+2 more

SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.


Core responsibilities include:

  • Production monitoring and debugging of live systems.
  • Incident investigation, troubleshooting, and problem resolution.
  • AWS cloud infrastructure support and maintenance.
  • Deployment and operational support activities.
  • Supporting a 24x7 production environment.
  • Working with GitHub-based development workflows.
  • Technical debt remediation and platform improvements.
  • Customer issue investigation and support.
  • Security and compliance-related work, including FedRAMP initiatives.


Preferred Skills:

AWS (especially S3 and EC2)

Strong debugging and troubleshooting skills

Site Reliability Engineering (SRE) experience

GitHub experience

Basic software development skills

TypeScript/JavaScript knowledge

C# preferred

AI experience is a plus.


Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
company logo
Priya Rawat
Posted by Priya Rawat
Gurugram
4 - 5 yrs
₹8L - ₹10L / yr
RCA
SLA
skill icongrafana
ELKI
SOP

About the Role


We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.


Key Responsibilities


  • Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
  • Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
  • Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
  • Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
  • Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
  • Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
  • Participate in incident/severity calls, ensuring clear communication and coordination
  • Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting


Required Skills & Experience


  • Strong understanding of Linux/Unix systems for application support
  • Hands-on experience troubleshooting applications in staging and production environments
  • Ability to monitor system performance and identify root causes using logs and metrics
  • Experience working with Kubernetes and microservices-based architectures
  • Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
  • Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
  • Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
  • Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
  • Experience with ticketing and documentation tools such as Jira and Confluence
  • Minimum 4+ years of experience in application support or reliability engineering


Education & Certifications


  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Relevant certifications (Cloud, Kubernetes, Microservices) are a plus


Work Schedule


  • Willingness to work in a 24x7 environment, including weekends and on-call rotations
Read more
company logo
Agency job
via by Soundarya Valli Chintapalli
Hyderabad
3 - 8 yrs
₹8L - ₹18L / yr
Linux/Unix
  • Linux troubleshooting
  • Hands-on AWS
  • Production/Application Support
  • Bash/Shell/Python
  • Monitoring/log analysis
  • Incident resolution
  • Application deployment/support
  • Basic networking and database knowledge
  • Production/batch support exposure
  • Willingness for rotational weekend/critical production support


Read more
company logo
Mayank Choudhary
Posted by Mayank Choudhary
Hyderabad
7 - 10 yrs
₹27L - ₹30L / yr
skill iconPython
skill iconAmazon Web Services (AWS)

Strong Python Developer profile with robust AWS exposure

2

Mandatory (Experience 1): Must have 7+ years of hands-on software development experience with at least the recent 4+ years in Python and strong hands on knowledge of AWS

3

Mandatory (Tech skill 1): Must have strong working knowledge of Python.

4

Mandatory (Tech skill 2): Must have good understanding of AWS services including EC2, S3, Lambda, IAM, CloudWatch, and ECS or ECR

5

Mandatory (Tech skill 3): Must be able to write and understand REST APIs

6

Mandatory (Tech skill 4): Must be comfortable with version control tools such as Git, GitHub, Bitbucket, or GitLab

7

Mandatory (Tech skill 5): Must have good understanding of databases such as PostgreSQL, MySQL, or DynamoDB

8

Mandatory (Tech skill 6): Must have familiarity with Linux commands and shell scripting

9

Mandatory (Skill): Must have good debugging and problem-solving skills, with the ability to read existing code and make changes independently with light guidance

10

Mandatory (Skill 2): Must have strong communication skills

11

Preferred (Tech skill 1): Experience with AWS CodeCommit, and exposure to AWS CodeBuild, CodeDeploy, CodePipeline, GitHub Actions, Jenkins, or similar CI/CD tools

12

Preferred (Tech skill 2): Basic understanding of CI/CD pipelines

13

Preferred (Tech skill 3): Experience with Docker or container-based applications

14

Preferred (Tech skill 4): Basic knowledge of infrastructure-as-code tools such as Terraform or AWS CloudFormation

Read more
Product Based Co
Product Based Co
Agency job
via by Rishika Teja
Hyderabad
18 - 25 yrs
₹70L - ₹80L / yr
SRE
skill iconAmazon Web Services (AWS)

Hiring SRE - Director


Exp : 18 - 25 yrs

Edu : BE/B.Tech

Work Location : Hyd


Must Have Skills :


Must be from product SaaS based companies.


18+ years in Software Engineering, SRE, or reliability roles; 5+ years in leadership(Director). 


Proven ability to leverage software engineering principles and practices to solve reliability and operational challenges.


Expertise in SLI/SLO and monitoring.


Expertise in CI/CD, observability, and incident response. 


Strong AWS knowledge and experience with container orchestration. 


Proven ability to lead reliability programs across multiple SaaS products. 


Experience architecting applications or infrastructure for high-growth cloud platforms. 


Experience in B2B SaaS environments involving large-scale distributed systems. 



Read more
Gurugram
4 - 10 yrs
₹4L - ₹10L / yr
DevOps
Site Reliability Engineer (SRE)
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
+14 more

🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience Level : 4+ Years

Location : Gurugram Sector 48, Haryana (On-site)

Employment Type : Full Time Opportunity


About the Role :

We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.

In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).


Mandatory Skills :

AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux


Key Responsibilities :

1. Cloud Infrastructure & Infrastructure as Code (IaC) :

  • Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
  • Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
  • Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.

2. CI/CD, Automation & Development :

  • Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
  • Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.

3. Containerization & Orchestration :

  • Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
  • Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.

4. Observability, SRE & Incident Management :

  • Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
  • Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
  • Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
  • Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).


Required Qualifications & Skills :

  • Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
  • Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
  • Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
  • CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
  • Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
  • Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
  • Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
  • Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
  • Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Read more
company logo
Mohammed Rabidheen
Posted by Mohammed Rabidheen
Coimbatore
3 - 8 yrs
Best in industry
Windows Azure
AKS
DevOps
Microsoft Windows Azure

Senior Cloud Site Reliability Engineer (CSRE) – Azure


About Searce:

Searce is an AI-native, engineering-led modern technology consultancy that empowers

clients to futurify their businesses by delivering real, intelligent business outcomes. As a

trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,

data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"

cultural mindset and our proprietary evlos problem-solving framework, we eliminate

bureaucratic fluff to build working prototypes fast and scale enterprise production

environments intelligently. We don't just fix systems; we leverage multi-cloud technologies

to transform client operations into distinct competitive advantages.

Position Overview:

We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to

architect, secure, and stabilize next-generation hybrid and multi-cloud environments.

Operating at the intersection of infrastructure design, security compliance, and production

operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,

and AWS.

Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a

massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For

the Lead path, you will couple this deep engineering toolkit with stakeholder management

and mentorship to drive an elite operational culture.


Experience & Level Expectation:

Years of Experience: 3 to 10 years of intensive, hands-on production operations

experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.

Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced

triaging of infrastructure failures, and ownership of the CI/CD and deployment

lifecycles.

Intermediate level (5-10 Years): Expected to take architectural ownership, serve as

primary Incident Commander for complex outages, design cross-cloud governance

frameworks, and act as a reliable bridge between technical teams and client

leadership.


Key Responsibilities & Role Expectations:

Multi-Cloud Platforms & Orchestration: Design, configure, and maintain

production-grade Kubernetes clusters across major platforms (AKS).

Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant

isolation.

Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable

infrastructure components using Terraform or Crossplane. Standardize automated

environment provisioning to eliminate configuration drift across multi-branch

environments.

Incident Management & Reliability (SRE): Own and optimize the production on-call

rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing

Mean Time to Recovery (MTTR) through centralized log and metric correlation.

Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to

identify core architectural vulnerabilities and establish long-term fixes preventing

recurrence.

Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,

operating system patching strategies (Linux/Windows), database lifecycle updates,

and multi-region Disaster Recovery (DR) failover drills.

Core Core Operations & Legacy Integration: Manage enterprise-level hybrid

networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP

configurations) while effectively connecting cloud native services to legacy

infrastructures like Active Directory.

Security & Governance: Embed Zero Trust policies, secure secrets management

(Secrets Manager/Key Vault), and continuous vulnerability patching into the

automated SDLC pipeline.


Required Technical Skills:

- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.

- Containers & Orchestration

  • Production-level management of GKE, AKS, and EKS.
  • Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption
Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos