Cutshort logo
For Employers
Olacabs.com logo
DevOps Engineer
DevOps Engineer

DevOps Engineer at Olacabs.com · Bengaluru (Bangalore) · 5 - 9 years · ₹8L - ₹21L / yr · Posted 8 Nov 2021

Olacabs.com's logo

DevOps Engineer

Roshni Pillai's profile picture
Posted by Roshni Pillai
5 - 9 yrs
₹8L - ₹21L / yr
Bengaluru (Bangalore)
Skills
DevOps
skill iconAmazon Web Services (AWS)
skill iconKubernetes
Linux/Unix
We are looking for a Site Reliability Engineer/Sr. Site Reliability Engineer to help us build and enhance platforms to achieve availability, scalability and operational effectiveness. The right individual will embrace the opportunity to tackle challenging problems and use their influence to drive continual improvement. You will also work on the cutting edge of technology, leveraging Kong, Repose, Docker, Mesos/Kubernetes, Jenkins, Chef, HaProxy, Nginx, GitLab, MySQL, Scylla, Aerospike, Service Mesh ( Istio/Linkerd), Prometheus etc.

Roles and Responsibilities
● Managing Availability, Performance, Capacity of infrastructure and applications.
● Building and implementing observability for applications health/performance/capacity.
● Optimizing On-call rotations and processes.
● Documenting “tribal” knowledge.
● Managing Infra-platforms like
- Mesos/Kubernetes
- CICD
- Observability(Prometheus/New Relic/ELK)
- Cloud Platforms ( AWS/ Azure )
- Databases
- Data Platforms Infrastructure
● Providing help in onboarding new services with the production readiness review process.
● Providing reports on services SLO/Error Budgets/Alerts and Operational Overhead.
● Working with Dev and Product teams to define SLO/Error Budgets/Alerts.
● Working with the Dev team to have an in-depth understanding of the application architecture and its bottlenecks.
● Identifying observability gaps in product services, infrastructure and working with stake owners to fix it.
● Managing Outages and doing detailed RCA with developers and identifying ways to avoid that situation.
● Managing/Automating upgrades of the infrastructure services.
● Automate toil work.

Experience & Skills
● 3+ Years of experience as an SRE/DevOps/Infrastructure Engineer on large scale microservices and infrastructure.
● A collaborative spirit with the ability to work across disciplines to influence, learn, and deliver.
● A deep understanding of computer science, software development, and networking principles.
● Demonstrated experience with languages, such as Python, Java, Golang etc.
● Extensive experience with Linux administration and good understanding of the various linux kernel subsystems (memory, storage, network etc).
● Extensive experience in DNS, TCP/IP, UDP, GRPC, Routing and Load Balancing.
● Expertise in GitOps, Infrastructure as a Code tools such as Terraform etc.. and Configuration Management Tools such as Chef, Puppet, Saltstack, Ansible.
● Expertise of Amazon Web Services (AWS) and/or other relevant Cloud Infrastructure solutions like Microsoft Azure or Google Cloud.
● Experience in building CI/CD solutions with tools such as Jenkins, GitLab, Spinnaker, Argo etc.
● Experience in managing and deploying containerized environments using Docker,
Mesos/Kubernetes is a plus.
● Experience with multiple datastores is a plus (MySQL, PostgreSQL, Aerospike,
Couchbase, Scylla, Cassandra, Elasticsearch).
● Experience with data platforms tech stacks like Hadoop, Hive, Presto etc is a plus
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Olacabs.com

Founded :
2010
Type :
Products & Services
Size
Stage

About

Ola is India’s largest mobility platform and one of the world’s largest ride-hailing companies, serving 250+ cities across India, Australia, New Zealand, and the UK. The Ola app offers mobility solutions by connecting customers to drivers and a wide range of vehicles across bikes, auto-rickshaws, metered taxis, and cabs, enabling convenience and transparency for hundreds of millions of consumers and over 1.5 million driver-partners. Ola’s core mobility offering in India is supplemented by its electric-vehicle arm, Ola Electric; India’s largest fleet management business, Ola Fleet Technologies and Ola Skilling, that aims to enable millions of livelihood opportunities for India's youth. With its acquisition of Ridlr, India’s leading public transportation app and investment in Vogo, a dockless scooter sharing solution, Ola is looking to build mobility for the next billion Indians. Ola also extends its consumer offerings like micro-insurance and credit led payments through Ola Financial Services and a range of owned food brands through India’s largest network of kitchens under its Food business. Ola’s core mobility offering in India is supplemented by its electric-vehicle arm, Ola Electric; India’s largest fleet management business, Ola Fleet Technologies and Ola Skilling, that aims to enable millions of livelihood opportunities for India's youth. With its acquisition of Ridlr, India’s leading public transportation app and investment in Vogo, a dockless scooter sharing solution, Ola is looking to build mobility for the next billion Indians. Ola also extends its consumer offerings like micro-insurance and credit led payments through Ola Financial Services and a range of owned food brands through India’s largest network of kitchens under its Food business.
Read more

Connect with the team

Profile picture
Athul ps
Profile picture
Shuhaib I
Profile picture
Roshni Pillai
Profile picture
Supriya Singh
Profile picture
Shivani Kukreja
Profile picture
Pradeep Kumaar

Company social profiles

linkedintwitterfacebook

Similar jobs (10)

EDM NETWORK
AHLOUCHE AHLOUCHE
Posted by AHLOUCHE AHLOUCHE
Remote only
1 - 10 yrs
$10K - $20K / yr (ESOP available)
User Experience (UX) Design

Key Responsibilities

Platform Monitoring and Reliability

  • Monitor the health and performance of EDM’s advertiser, publisher, broker, and internal platforms.
  • Maintain monitoring, alerting, logging, and system-health dashboards.
  • Investigate platform outages, degraded performance, failed transactions, delayed data, and integration errors.
  • Respond to production incidents and coordinate resolutions with the appropriate engineers and vendors.
  • Perform root-cause analysis and document corrective and preventive actions.
  • Help maintain defined uptime, response-time, recovery-time, and system-reliability targets.
  • Identify recurring problems and recommend permanent solutions.

Cloud Infrastructure and Systems Operations

  • Maintain and support cloud infrastructure, servers, databases, networks, storage, and production environments.
  • Support development, staging, and production environments.
  • Assist with infrastructure scaling, system upgrades, patching, backups, and disaster recovery.
  • Monitor cloud usage and help control infrastructure and technology costs.
  • Maintain access controls, service accounts, certificates, domain configurations, and environment variables.
  • Ensure production systems are properly documented and recoverable.

Deployment and Release Support

  • Support safe and consistent application deployments.
  • Maintain or improve continuous integration and deployment workflows.
  • Coordinate release schedules, deployment validation, rollback procedures, and post-release monitoring.
  • Help engineering teams identify configuration or infrastructure problems before releases reach production.
  • Maintain deployment documentation, technical checklists, and change logs.
  • Reduce manual deployment work through automation.

API and Integration Support

  • Monitor and troubleshoot third-party APIs, webhooks, postbacks, dialer connections, CRM integrations, payment systems, tracking platforms, and compliance services.
  • Investigate failed lead deliveries, missing postbacks, duplicate records, delayed reporting, and authentication problems.
  • Support ping-post, real-time bidding, call-routing, SIP, and data-transfer workflows.
  • Work with advertisers, publishers, and vendors to diagnose technical integration problems.
  • Create clear documentation for common integration methods and troubleshooting procedures.
  • Develop alerts that identify integration failures before clients report them.

Call and Lead Operations

  • Monitor call-routing, tracking, recording, attribution, and disposition systems.
  • Investigate calls that fail to route, connect, record, track, or report correctly.
  • Support number provisioning, routing rules, caps, schedules, geographic restrictions, buyer availability, and failover logic.
  • Validate that leads, calls, and transactions are properly attributed to the correct advertiser, publisher, campaign, and payout.
  • Assist with discrepancies involving call duration, billable events, conversions, payouts, and reporting.
  • Help protect revenue by identifying technical leakage and delivery failures.

Data and Reporting Support

  • Monitor data pipelines, scheduled jobs, reporting processes, and database performance.
  • Investigate discrepancies between platform reporting, billing records, payment records, and third-party systems.
  • Write and maintain database queries for troubleshooting, validation, and operational reporting.
  • Assist with data corrections using controlled and documented procedures.
  • Support dashboards and operational alerts for revenue, margin, consumption, conversion, and platform activity.
  • Maintain appropriate controls around production data access and modification.

Security and Access Management

  • Support role-based access controls, multifactor authentication, audit logging, encryption, and secure system configuration.
  • Provision and remove employee, contractor, client, and vendor access.
  • Monitor suspicious activity and report potential security incidents.
  • Assist with vulnerability remediation, security reviews, access audits, and incident-response procedures.
  • Protect consumer, advertiser, publisher, employee, and company information.
  • Follow company policies for credentials, production access, sensitive data, and change management.

Automation and Process Improvement

  • Automate repetitive operational tasks using scripts, workflows, APIs, and infrastructure tools.
  • Reduce manual work associated with monitoring, deployments, reporting, reconciliation, onboarding, and support.
  • Build internal tools that improve visibility and response times.
  • Identify operational bottlenecks that affect revenue, margin, client satisfaction, or employee productivity.
  • Maintain clear runbooks and standard operating procedures for recurring technical tasks.

Technical Support and Documentation

  • Serve as an escalation point for complex platform and integration issues.
  • Translate technical problems into clear explanations for nontechnical teams.
  • Create and maintain architecture diagrams, system inventories, runbooks, troubleshooting guides, and incident reports.
  • Track incidents and technical requests through completion.
  • Document known issues, temporary workarounds, permanent resolutions, and system dependencies.
  • Participate in an on-call rotation for urgent production incidents.

First 90-Day Priorities

The successful candidate will be expected to:

  • Learn EDM’s platforms, infrastructure, integrations, call-routing systems, reporting processes, and revenue workflows.
  • Document critical systems, dependencies, credentials ownership, vendor contacts, and escalation procedures.
  • Review existing monitoring, alerting, backups, access controls, and deployment procedures.
  • Establish baseline metrics for uptime, incident volume, response time, recovery time, deployment success, and integration failures.
  • Resolve high-priority recurring production and integration issues.
  • Improve alerting for call-routing failures, API errors, delayed data, failed jobs, and reporting discrepancies.
  • Create runbooks for the company’s most common and highest-risk technical incidents.
  • Identify at least three meaningful automation or cost-saving opportunities.
  • Participate in production support and demonstrate ownership of incidents through resolution.

Performance Expectations

Success will be measured by:

  • Platform uptime and reliability
  • Mean time to acknowledge and resolve incidents
  • Reduction in recurring production problems
  • Deployment success and rollback rates
  • API, postback, webhook, and call-routing reliability
  • Reporting and data accuracy
  • Backup and recovery readiness
  • Quality of technical documentation
  • Security and access-control compliance
  • Reduction in manual operational work
  • Infrastructure costs relative to platform volume
  • Responsiveness to internal teams, clients, and technical partners

Required Qualifications

  • Three or more years of experience in operations engineering, DevOps, site reliability engineering, cloud infrastructure, systems administration, or production support.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Experience supporting Linux-based production environments.
  • Working knowledge of networking, DNS, SSL certificates, firewalls, load balancing, and application security.
  • Experience with relational databases and SQL.
  • Experience troubleshooting REST APIs, webhooks, authentication, and third-party integrations.
  • Familiarity with monitoring, logging, alerting, and incident-management tools.
  • Experience with scripting languages such as Python, Bash, JavaScript, or PowerShell.
  • Understanding of source control, deployment pipelines, and release management.
  • Strong troubleshooting, documentation, and communication skills.
  • Ability to prioritize incidents based on business and revenue impact.
  • Availability to participate in an on-call rotation.

Preferred Qualifications

  • Experience in ad-tech, mar-tech, affiliate marketing, lead generation, pay-per-call, telecommunications, or SaaS.
  • Familiarity with SIP, VoIP, dialers, call-tracking platforms, routing systems, and phone-number provisioning.
  • Experience with containers, infrastructure as code, and automated deployment tools.
  • Experience with Docker, Kubernetes, Terraform, GitHub Actions, or similar technologies.
  • Familiarity with payment processing, usage-based billing, reconciliation, and commission systems.
  • Experience working with real-time bidding, ping-post, lead distribution, or high-volume transactional systems.
  • Understanding of TCPA-related controls, consent records, DNC suppression, data privacy, or regulated marketing environments.
  • Experience with security audits, disaster-recovery testing, and compliance documentation.

Ideal Candidate

The ideal candidate:

  • Takes ownership instead of waiting for someone else to fix the problem.
  • Remains calm and methodical during high-impact incidents.
  • Understands the difference between applying a temporary fix and eliminating a root cause.
  • Communicates technical problems clearly and directly.
  • Recognizes that production reliability is a business and revenue responsibility.
  • Automates repetitive work whenever practical.
  • Documents systems so the company is not dependent on one person’s memory.
  • Protects security and stability without creating unnecessary bureaucracy.
  • Is comfortable working in a fast-moving entrepreneurial environment.
  • Can manage competing priorities while maintaining attention to detail.


Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Farook sharief
Hyderabad
7 - 11 yrs
₹2L - ₹15L / yr
Network
Reliability engineering
skill iconAmazon Web Services (AWS)
Windows Azure

Job Summary

We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in networking, DNS, load balancing, and hybrid cloud environments. The candidate will be responsible for maintaining the reliability, availability, performance, and scalability of production infrastructure and services.

The ideal candidate should have strong troubleshooting skills and experience working across network, cloud, infrastructure, and application environments.

Key Responsibilities

  • Monitor and maintain the availability and reliability of production systems and services.
  • Troubleshoot complex network, infrastructure, and application connectivity issues.
  • Manage and troubleshoot DNS services, DNS resolution, records, and configuration issues.
  • Configure, manage, and troubleshoot Load Balancers and traffic routing.
  • Work with Layer 4 and Layer 7 networking and understand TCP/IP, HTTP/HTTPS, routing, and network connectivity.
  • Support hybrid cloud environments involving on-premises infrastructure and public cloud platforms.
  • Troubleshoot connectivity between on-premises data centers and cloud environments.
  • Participate in production incidents, troubleshooting, root cause analysis (RCA), and problem management.
  • Develop automation scripts and tools to reduce manual operational activities.
  • Configure and maintain monitoring, alerting, and observability solutions.
  • Work closely with Network, Cloud, DevOps, Security, and Application teams.
  • Participate in on-call support and resolve production issues within defined SLAs.
  • Document infrastructure, troubleshooting procedures, incident reports, and operational processes.
  • Identify opportunities to improve system reliability, performance, and scalability.
Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen
Hyderabad
7 - 11 yrs
₹2L - ₹15L / yr
Reliability engineering
skill iconAmazon Web Services (AWS)
Windows Azure
Network
DNS

SRE – Network / DNS / Load Balancer

Experience: 7–11 Years

Location: Hyderabad

Work Mode: WFO

Availability: Immediate Joiner

Job Description:

  • Strong experience in Site Reliability Engineering (SRE) with focus on infrastructure and application reliability.
  • Hands-on experience with Network, DNS and Load Balancer troubleshooting and administration.
  • Monitor system performance, availability, latency and infrastructure health.
  • Troubleshoot network connectivity, DNS resolution, routing and load-balancing issues.
  • Experience with load balancers such as F5, BIG-IP, HAProxy or similar technologies.
  • Good understanding of TCP/IP, HTTP/HTTPS, LAN/WAN, SSL/TLS and networking concepts.
  • Experience with DNS technologies such as BIND, Infoblox or equivalent.
  • Work on incident management, root cause analysis and problem resolution.
  • Collaborate with application, network, cloud and infrastructure teams to resolve production issues.
  • Experience with monitoring and alerting tools such as Splunk, Grafana, Prometheus, AppDynamics or similar tools.
  • Strong troubleshooting, production support and communication skills.
  • Willingness to work from Hyderabad office (WFO) and join immediately. 


Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via Unique Occupational by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
Vy Systems
at Vy Systems
1 recruiter
Kalki K
Posted by Kalki K
Mumbai
6 - 8 yrs
₹15L - ₹20L / yr
SLA

We are looking for an experienced SRE with DevOps expertise to ensure the reliability, availability, scalability, and performance of production systems. The candidate should have hands-on experience in cloud infrastructure, CI/CD automation, monitoring, incident management, and infrastructure as code.

Responsibilities

  • Monitor production environments and ensure high availability, reliability, and performance.
  • Build and maintain CI/CD pipelines using Jenkins, GitLab CI, or GitHub Actions.
  • Automate infrastructure provisioning and configuration using Terraform, Ansible, or similar tools.
  • Manage cloud infrastructure on AWS, Azure, or GCP.
  • Work with Docker and Kubernetes for containerization and orchestration.
  • Implement monitoring, logging, and alerting using Prometheus, Grafana, Splunk, ELK, or Datadog.
  • Handle incident management, root cause analysis, troubleshooting, and post-incident reviews.
  • Develop automation scripts using Python, Bash, or Shell scripting.
  • Improve system reliability by reducing manual tasks, toil, and recurring incidents.
  • Manage deployments, releases, disaster recovery, and production support activities.
  • Define and monitor SLIs, SLOs, and SLAs to improve service reliability.

Required Skills

  • Strong experience in SRE and DevOps practices.
  • Hands-on experience with AWS / Azure / GCP.
  • CI/CD tools: Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
  • Infrastructure as Code: Terraform, Ansible, or CloudFormation.
  • Containerization: Docker and Kubernetes.
  • Scripting: Python, Bash, or Shell.
  • Monitoring and observability: Prometheus, Grafana, Splunk, ELK, or Datadog.
  • Strong Linux/Unix administration and troubleshooting skills.
  • Experience in production support, incident management, and root cause analysis.
  • Good understanding of networking, system performance, high availability, and disaster recovery.


Read more
Hyderabad
5 - 10 yrs
₹4L - ₹15L / yr
Production support
Application server
skill iconKubernetes
Terraform
skill iconAmazon Web Services (AWS)
+2 more

ob Summary

We are looking for an experienced Site Reliability Engineer (SRE) / Production Support Engineer with strong hands-on experience in application and production support, incident management, monitoring, cloud operations, automation, and infrastructure technologies.

The ideal candidate will be responsible for ensuring the availability, reliability, performance, and stability of production applications and infrastructure. The role involves troubleshooting critical production issues, monitoring applications and infrastructure, supporting deployments, managing incidents, and driving automation and operational improvements.

Key Responsibilities

  • Provide L2/L3 Application and Production Support for critical business applications.
  • Monitor production applications, infrastructure, batch jobs, and system health.
  • Handle and troubleshoot critical production incidents, ensuring timely resolution and minimal business impact.
  • Participate in Incident, Problem, and Change Management processes.
  • Perform root-cause analysis (RCA) for recurring and major production issues.
  • Troubleshoot issues related to applications, networks, load balancers, databases, operating systems, and infrastructure.
  • Monitor application and infrastructure performance using Splunk, APM, and other monitoring tools.
  • Create and maintain Splunk queries, dashboards, alerts, and operational monitoring.
  • Support production deployments, including Blue-Green and Canary deployment strategies.
  • Work with cloud infrastructure and perform day-to-day Cloud Operations activities.
  • Manage and troubleshoot containerized applications using Docker and Kubernetes.
  • Work with Terraform / Infrastructure as Code (IaC) for infrastructure provisioning and automation.
  • Support Linux and Windows server administration.
  • Develop and maintain Shell scripts and Python automation scripts to reduce manual operational activities.
  • Monitor and analyze SLIs, SLOs, Error Budgets, and Burn Rates.
  • Identify reliability risks and proactively implement solutions to improve system availability and performance.
  • Collaborate with Development, DevOps, Infrastructure, Network, Database, and Cloud teams during production incidents.
  • Leverage GenAI tools such as GitHub Copilot, Claude, or similar tools to improve troubleshooting, automation, documentation, and operational efficiency.
  • Participate in on-call/shift-based production support as required by business and customer needs.

Mandatory / Key Skills

  • Production / Application Support
  • Incident Management
  • Production Monitoring & Batch Monitoring
  • Splunk – Queries, Dashboards & Monitoring
  • APM / Application Performance Monitoring
  • SLI / SLO / Error Budget / Burn Rate
  • Cloud Operations
  • Kubernetes
  • Docker
  • Terraform / Infrastructure as Code
  • Linux & Windows Administration
  • Shell Scripting
  • Python Scripting / Automation
  • Production Deployment Support
  • Blue-Green & Canary Deployments
  • Network, Load Balancing & Database Troubleshooting

​

Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Akilandeswari Panneerselvam
Hyderabad
5 - 10 yrs
₹6L - ₹10L / yr
SRE
Production support
Cloud Computing
skill iconKubernetes
Linux/Unix
+1 more

Site Reliability Engineer (SRE) / Production Support Engineer

Experience: 5–10 Years

Location: Hyderabad

Work Mode: Face-to-Face Drive

Shift: Rotational Shifts

Job Description

Looking for an experienced SRE / Production Support Engineer with strong experience in application and production support, incident management, monitoring, troubleshooting, and cloud operations.

Key Skills

Production Support, Incident Management, Splunk, APM, SLI/SLO, Cloud, Kubernetes, Docker, Terraform, Linux/Windows Administration, Shell Scripting and Python.

Good understanding of production deployments, batch monitoring, network/load balancing, and troubleshooting is required.

Candidates from SRE, Production Support, Application Support, Cloud Operations, or DevOps backgrounds are preferred.

Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen
Mumbai
6 - 8 yrs
₹1L - ₹15L / yr
DevOps
Reliability engineering
skill iconKubernetes
skill iconDocker



Job Summary

We are looking for an experienced SRE / DevOps Engineer with strong expertise in production support, site reliability engineering, cloud infrastructure, automation, and monitoring. The candidate should have experience maintaining highly available systems, troubleshooting production issues, and improving infrastructure reliability and performance.

Mandatory Skills

  • Site Reliability Engineering (SRE)
  • DevOps Engineering
  • Production Support
  • Linux Administration
  • Cloud Infrastructure (AWS / Azure / GCP)
  • Kubernetes and Docker
  • CI/CD pipelines
  • Monitoring and Observability
  • Scripting using Python / Shell
  • Incident Management and Root Cause Analysis

Key Responsibilities

  • Monitor and maintain the availability, reliability, and performance of applications and infrastructure.
  • Handle production incidents, troubleshoot issues, and perform root cause analysis.
  • Manage and automate infrastructure deployment and operational tasks.
  • Work with CI/CD pipelines and container orchestration platforms.
  • Implement monitoring, alerting, and observability solutions.
  • Collaborate with development and operations teams to improve system reliability.
  • Support cloud infrastructure, Linux environments, and production deployments.
  • Identify recurring issues and implement preventive measures to reduce downtime.
Read more
MNC
MNC
Agency job
Mumbai
4 - 8 yrs
₹1L - ₹10L / yr
SRE
DevOps
CI/CD
skill iconKubernetes

Required Skills:

  • Strong experience in Site Reliability Engineering (SRE) and DevOps practices.
  • Hands-on experience with CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI/CD.
  • Strong knowledge of AWS, Azure, or GCP cloud platforms.
  • Experience with Docker, Kubernetes, and container orchestration.
  • Hands-on experience with Terraform, Ansible, or other Infrastructure as Code (IaC) tools.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, ELK, or Datadog.
  • Good scripting skills in Python, Bash, or Shell scripting.
  • Understanding of SLI, SLO, SLA, error budgets, and service reliability.
  • Experience in incident management, troubleshooting, root-cause analysis (RCA), and production support.
  • Knowledge of system performance monitoring, capacity planning, high availability, and disaster recovery.
  • Experience with Linux/Unix administration and networking fundamentals.
  • Familiarity with Agile methodologies, Git, and automated deployment practices.


Read more
Deltek
Remote only
3 - 5 yrs
Best in industry
CI/CD
skill iconPostgreSQL
skill iconPython
skill iconAmazon Web Services (AWS)
Artificial Intelligence (AI)
+2 more

SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.


Core responsibilities include:

  • Production monitoring and debugging of live systems.
  • Incident investigation, troubleshooting, and problem resolution.
  • AWS cloud infrastructure support and maintenance.
  • Deployment and operational support activities.
  • Supporting a 24x7 production environment.
  • Working with GitHub-based development workflows.
  • Technical debt remediation and platform improvements.
  • Customer issue investigation and support.
  • Security and compliance-related work, including FedRAMP initiatives.


Preferred Skills:

AWS (especially S3 and EC2)

Strong debugging and troubleshooting skills

Site Reliability Engineering (SRE) experience

GitHub experience

Basic software development skills

TypeScript/JavaScript knowledge

C# preferred

AI experience is a plus.


Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos