Sre with Devops at MNC · Mumbai · 6 - 8 years · ₹1L - ₹15L / yr · Posted 9 Oct 2026

Job Summary
We are looking for an experienced SRE / DevOps Engineer with strong expertise in production support, site reliability engineering, cloud infrastructure, automation, and monitoring. The candidate should have experience maintaining highly available systems, troubleshooting production issues, and improving infrastructure reliability and performance.
Mandatory Skills
- Site Reliability Engineering (SRE)
- DevOps Engineering
- Production Support
- Linux Administration
- Cloud Infrastructure (AWS / Azure / GCP)
- Kubernetes and Docker
- CI/CD pipelines
- Monitoring and Observability
- Scripting using Python / Shell
- Incident Management and Root Cause Analysis
Key Responsibilities
- Monitor and maintain the availability, reliability, and performance of applications and infrastructure.
- Handle production incidents, troubleshoot issues, and perform root cause analysis.
- Manage and automate infrastructure deployment and operational tasks.
- Work with CI/CD pipelines and container orchestration platforms.
- Implement monitoring, alerting, and observability solutions.
- Collaborate with development and operations teams to improve system reliability.
- Support cloud infrastructure, Linux environments, and production deployments.
- Identify recurring issues and implement preventive measures to reduce downtime.

Similar jobs (10)
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
Site Reliability Engineer (SRE) / Production Support Engineer
Experience: 5–10 Years
Location: Hyderabad
Work Mode: Face-to-Face Drive
Shift: Rotational Shifts
Job Description
Looking for an experienced SRE / Production Support Engineer with strong experience in application and production support, incident management, monitoring, troubleshooting, and cloud operations.
Key Skills
Production Support, Incident Management, Splunk, APM, SLI/SLO, Cloud, Kubernetes, Docker, Terraform, Linux/Windows Administration, Shell Scripting and Python.
Good understanding of production deployments, batch monitoring, network/load balancing, and troubleshooting is required.
Candidates from SRE, Production Support, Application Support, Cloud Operations, or DevOps backgrounds are preferred.
Required Skills:
- Strong experience in Site Reliability Engineering (SRE) and DevOps practices.
- Hands-on experience with CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI/CD.
- Strong knowledge of AWS, Azure, or GCP cloud platforms.
- Experience with Docker, Kubernetes, and container orchestration.
- Hands-on experience with Terraform, Ansible, or other Infrastructure as Code (IaC) tools.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, ELK, or Datadog.
- Good scripting skills in Python, Bash, or Shell scripting.
- Understanding of SLI, SLO, SLA, error budgets, and service reliability.
- Experience in incident management, troubleshooting, root-cause analysis (RCA), and production support.
- Knowledge of system performance monitoring, capacity planning, high availability, and disaster recovery.
- Experience with Linux/Unix administration and networking fundamentals.
- Familiarity with Agile methodologies, Git, and automated deployment practices.

Job Title: Senior Site Reliability Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 6+ years
About Compnay
It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.
Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.
We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.
Role Overview
We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.
Key Responsibilities
- Ensure high availability, performance, and reliability of production systems
- Design, implement, and manage scalable infrastructure solutions
- Build and maintain CI/CD pipelines for efficient software delivery
- Monitor system health using observability tools and respond to incidents proactively
- Automate operational processes using scripting and Infrastructure as Code (IaC)
- Manage containerized environments using Docker and Kubernetes
- Collaborate with cross-functional teams to improve system architecture and resilience
- Participate in on-call rotations and incident management processes
- Continuously optimize cloud infrastructure for cost, performance, and scalability
Required Qualifications & Skills
- Bachelor’s degree in Computer Science, IT, or related field
- 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
- Strong experience with containerization (Docker) and orchestration (Kubernetes)
- Proficiency in Linux administration, networking, and system security
- Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
- Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
- Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
- Proficiency in scripting languages (Python, Bash, or PowerShell)
- Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
- Solid understanding of system architecture, microservices, and SaaS/PaaS models
- Strong analytical and problem-solving skills
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience Level : 4+ Years
Location : Gurugram Sector 48, Haryana (On-site)
Employment Type : Full Time Opportunity
About the Role :
We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.
In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).
Mandatory Skills :
AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux
Key Responsibilities :
1. Cloud Infrastructure & Infrastructure as Code (IaC) :
- Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
- Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
- Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.
2. CI/CD, Automation & Development :
- Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
- Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.
3. Containerization & Orchestration :
- Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
- Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.
4. Observability, SRE & Incident Management :
- Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
- Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
- Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
- Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).
Required Qualifications & Skills :
- Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
- Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
- Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
- CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
- Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
- Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
- Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
- Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
- Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
ob Summary
We are looking for an experienced Site Reliability Engineer (SRE) / Production Support Engineer with strong hands-on experience in application and production support, incident management, monitoring, cloud operations, automation, and infrastructure technologies.
The ideal candidate will be responsible for ensuring the availability, reliability, performance, and stability of production applications and infrastructure. The role involves troubleshooting critical production issues, monitoring applications and infrastructure, supporting deployments, managing incidents, and driving automation and operational improvements.
Key Responsibilities
- Provide L2/L3 Application and Production Support for critical business applications.
- Monitor production applications, infrastructure, batch jobs, and system health.
- Handle and troubleshoot critical production incidents, ensuring timely resolution and minimal business impact.
- Participate in Incident, Problem, and Change Management processes.
- Perform root-cause analysis (RCA) for recurring and major production issues.
- Troubleshoot issues related to applications, networks, load balancers, databases, operating systems, and infrastructure.
- Monitor application and infrastructure performance using Splunk, APM, and other monitoring tools.
- Create and maintain Splunk queries, dashboards, alerts, and operational monitoring.
- Support production deployments, including Blue-Green and Canary deployment strategies.
- Work with cloud infrastructure and perform day-to-day Cloud Operations activities.
- Manage and troubleshoot containerized applications using Docker and Kubernetes.
- Work with Terraform / Infrastructure as Code (IaC) for infrastructure provisioning and automation.
- Support Linux and Windows server administration.
- Develop and maintain Shell scripts and Python automation scripts to reduce manual operational activities.
- Monitor and analyze SLIs, SLOs, Error Budgets, and Burn Rates.
- Identify reliability risks and proactively implement solutions to improve system availability and performance.
- Collaborate with Development, DevOps, Infrastructure, Network, Database, and Cloud teams during production incidents.
- Leverage GenAI tools such as GitHub Copilot, Claude, or similar tools to improve troubleshooting, automation, documentation, and operational efficiency.
- Participate in on-call/shift-based production support as required by business and customer needs.
Mandatory / Key Skills
- Production / Application Support
- Incident Management
- Production Monitoring & Batch Monitoring
- Splunk – Queries, Dashboards & Monitoring
- APM / Application Performance Monitoring
- SLI / SLO / Error Budget / Burn Rate
- Cloud Operations
- Kubernetes
- Docker
- Terraform / Infrastructure as Code
- Linux & Windows Administration
- Shell Scripting
- Python Scripting / Automation
- Production Deployment Support
- Blue-Green & Canary Deployments
- Network, Load Balancing & Database Troubleshooting
Job Description
The engineer will provide direct support to service development teams using the platform and maintain/develop platform components across CI pipelines, tool integrations, deployment architecture, monitoring, documentation, and self-service initiatives.
Responsibilities
- Support onboarding and technical discussions with development teams.
- Provide technical guidance and documentation.
- Support development engineers on CI/CD pipelines, deployments, logging and monitoring.
- Improve platform code, processes and documentation.
- Provide production infrastructure/operations support, including on-call support.
- Research product requirements and new technology rollouts.
Core Mandatory Skills
Kubernetes / Amazon EKS, Terraform, Terragrunt, AWS Cloud, Docker, CI/CD Pipelines, GitOps, Python / Go / Java, Cloud Operations / DevOps, Logging / Monitoring / Tracing, Incident Management
Job Summary :
We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.
Key Responsibilities :
- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.
- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.
- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management
- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.
- Collaborate with developers to ensure smooth and frequent deployments to production.
- Manage versioning and rollback strategies for critical deployments.
- Containerization & Orchestration using Kubernetes.
- Containerize applications using Docker, and manage them using Kubernetes.
- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.
- Develop utilities and tools to enhance operational efficiency and reliability.
- Monitoring & Incident Management
- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.
- Optimize application and system performance through proactive monitoring and configuration tuning.
Desired Skills and Experience :
- Experience Required - 6+ yrs.
- Hands-on experience on cloud services like AWS, EKS etc.
- Ability to design a good cloud solution.
- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.
- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.
- Use knowledge and research to constantly modernize our applications and infrastructure stacks.
- Be a team player and strong problem-solver to work with a diverse team.
- Having good communication skills.
Role Overview
We are looking for an experienced **Site Reliability Engineer (SRE) / Production Support Engineer** with strong expertise in application/production support, incident management, monitoring, cloud operations, automation, and infrastructure technologies.
The candidate will be responsible for ensuring the **availability, reliability, and performance of production applications**, troubleshooting critical issues, monitoring systems, and supporting production deployments.
#### Key Responsibilities & Skills
- Strong experience in **Application / Production Support**.
- Hands-on experience in **Production Issue Handling and Incident Management**.
- Experience in **Batch Monitoring** and production monitoring activities.
- Good understanding of **Network, Load Balancing, and Database-related issues**.
- Strong understanding of **SLI, SLO, Error Budget, and Burn Rate concepts**.
- Hands-on experience with **Splunk queries and dashboard creation/monitoring**.
- Experience with **APM (Application Performance Monitoring) tools**.
- Experience supporting **Production Deployments**, including **Blue-Green and Canary deployments**.
- Good knowledge of **Cloud Operations**.
- Hands-on experience with **Docker, Kubernetes, and Terraform/IaC**.
- Strong knowledge of **Linux/Windows administration and shell scripting**.
- Experience with **Windows administration and scripting**.
- Knowledge of **Python scripting/automation**.
- Exposure to **GenAI tools such as GitHub Copilot, Claude, or similar tools** for day-to-day operational activities.
- Strong troubleshooting, problem-solving, and incident-resolution skills.
Shift Timings
The role involves rotational shifts:
As a DevOps Engineer at YOYO, you'll own the infrastructure and delivery backbone that keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and observability that let a small, fast-moving team ship confidently and you'll keep our AI and data workloads reliable and affordable at scale. This is a hands-on role with real ownership: you won't be maintaining someone else's setup, you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make deployment boring, incidents rare, and scaling a non-event.
If you are Interested DM me on LinkedIn - Saquib Mundagnur
What You'll Own
CI/CD & developer experience - Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible automated testing, rollbacks, and release controls.Cloud infrastructure & IaC - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
Containers & orchestration - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch processing for the speech pipeline.
Reliability & observability (SRE) - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
Data & pipeline infrastructure - Support the infrastructure behind large-scale, edge-to-cloud data movement and processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
Security & compliance - Bake security into the platform: secrets management, IAM/least-privilege, encryption in transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security requirements) for a product that handles sensitive customer conversations.
Cost & scale - Own cloud cost visibility and optimization; make scaling decisions that balance reliability and spend.
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production systems at meaningful scale.
- Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and Infrastructure-as-Code (**Terraform** or equivalent).
- Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong automation-first mindset.
- Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, OpenTelemetry, or similar) and running incident response / on-call.
- A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team











