Senior Devops Engineer 1 top b2c product company only at Talent Pro · Bengaluru (Bangalore) · 4 - 6 years · ₹30L - ₹37L / yr · Bootstrapped · Posted 12 Jan 2026

Candidate must be from a product-based company with experience handling large-scale production traffic.
2. Candidate must have strong Linux expertise with hands-on production troubleshooting and working knowledge of databases and middleware (Mongo, Redis, Cassandra, Elasticsearch, Kafka).
3. Candidate must have solid experience with Kubernetes.
4. Candidate should have strong knowledge of configuration management tools like Ansible, Terraform, and Chef / Puppet. Add on- Prometheus & Grafana etc.
5. Candidate must be an individual contributor with strong ownership.
6. Candidate must have hands-on experience with DATABASE MIGRATIONS and observability tools such as Prometheus and Grafana.
7. Candidate must have working knowledge of Go/Python and Java.
8. Candidate should have working experience on Cloud platform - AWS
9. Candidate should have Minimum 1.5 years stability per organization, and a clear reason for relocation

Similar jobs (10)
Roles & Responsibilities
- Own end-to-end customer deployments across cloud, BYOC, air-gapped, and data-center environments
- Deploy and operate production environments using Kubernetes, Helm, and Terraform
- Troubleshoot complex Kubernetes, networking, infrastructure, application, and deployment issues
- Build and maintain observability, monitoring, alerting, and reliability
- Handle customer escalations and drive issues to resolution with a "fix first, optimize later" mindset
- Automate deployment and operational workflows
- Work closely with Engineering, Product, and Customer teams to resolve production challenges
- Participate in on-call and shift rotations, including critical customer escalations outside standard working hours
Ideal Candidate
1.Strong Hands-On DevOps Engineer Profile with deep Kubernetes and multi-cloud experience
2.Mandatory (Experience 1): Must have 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.
3.Mandatory (Experience 2): Must be able to own end-to-end customer deployments across cloud and enterprise environments (the role covers cloud, BYOC, air-gapped, and data-center deployments).
4.Mandatory (Tech skill 1): Must have strong hands-on Kubernetes experience — troubleshooting, networking, workloads, storage, RBAC, and security
5.Mandatory (Tech skill 2): Must have hands-on Helm experience — deployment, templating, and troubleshooting.
6.Mandatory (Tech skill 3): Must have hands-on Terraform experience — infrastructure provisioning and automation
7.Mandatory (Tech skill 4): Must have strong Linux and cloud infrastructure troubleshooting skills
8.Mandatory (Tech skill 5): Must have hands-on experience with observability — metrics, logs, traces, and alerting
9.Mandatory (Tech skill 6): Must have hands-on experience with one or more of AWS / GCP / Azure / OCI
10.Mandatory (Tech skill 7): Must have a strong incident-management and production-troubleshooting mindset, with a "fix first, optimize later" approach to customer escalations.
11.Mandatory (Company): B2B SaaS product companies
12.Preferred (Enterprise deployments): Prior experience with BYOC, air-gapped, or restricted-network environments.
13.Preferred (Observability tools): Prometheus / Grafana / OpenTelemetry / ELK / OpenSearch / Datadog
14.Preferred (Other): GitOps / CI-CD; Python or Golang for automation; multi-tenant SaaS infrastructure scaling; exposure to AI/ML pipeline deployments or iPaaS / reverse ETL connectors; SRE concepts (SLIs/SLOs, DR, high availability).
Build production-grade cloud infrastructure that powers enterprise applications with cutting-edge DevOps practices.
What you'll do:
- Design CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI)
- Containerize apps with Docker, deploy on Kubernetes clusters
- Manage infrastructure as code (Terraform, CloudFormation)
- Set up monitoring (Prometheus, Grafana, ELK Stack)
- Cloud migrations (AWS EC2, EKS, RDS → GCP equivalent)
- Optimize costs and performance for live production systems
What we need:
- Basic Python/Bash scripting
- Docker basics, Git workflows
- Cloud exposure (AWS/GCP/Azure free tier projects)
- Problem-solving mindset, eagerness to learn
Real impact:
- Deploy apps used by 1000+ daily users
- Work with senior DevOps engineers on client deliverables
- Build portfolio for FAANG-level interviews
We are looking for an experienced DevOps Engineer to take ownership of production infrastructure, cloud environments, Kubernetes platforms, and infrastructure automation. This is a hands-on role for someone who enjoys solving complex infrastructure challenges and is comfortable being responsible for systems in production.
Key Responsibilities
- Own and operate production infrastructure, including participating in an on-call rotation and responding to production incidents.
- Design, operate, and continuously improve Kubernetes clusters in production.
- Manage and automate infrastructure using Infrastructure as Code, primarily with Terraform.
- Build, maintain, and optimise cloud infrastructure across AWS, GCP, or Azure.
- Work extensively with Linux, including system administration, networking, troubleshooting, and system-level configuration.
- Manage production deployment and GitOps workflows using ArgoCD.
- Improve infrastructure reliability, scalability, security, monitoring, and operational efficiency.
- Troubleshoot complex production issues and drive problems through to resolution.
- Develop automation and processes that reduce manual operational work.
Essential Requirements
- 4+ years of hands-on experience operating production infrastructure, with personal ownership and responsibility for live systems, including on-call experience.
- Deep, hands-on Kubernetes experience — you must have operated and managed Kubernetes clusters, rather than simply deploying applications onto clusters managed by another team.
- Strong experience with Infrastructure as Code, with Terraform strongly preferred. Experience with Pulumi or CloudFormation is also considered.
- Strong experience with at least one major cloud platform, ideally AWS. Strong GCP or Azure experience is also welcome, provided you are willing to work with AWS.
- Strong Linux skills and confidence working from the command line, including networking, troubleshooting, system configuration, and performance issues.
- Production experience with ArgoCD and GitOps-based deployment workflows.
- Strong troubleshooting and problem-solving skills, with the ability to take ownership of production incidents and infrastructure issues.
Nice to Have
Experience with email infrastructure would be a strong advantage, particularly:
- Exim
- IMAP / SMTP
- Postfix
- Dovecot
- General mail server administration and maintenance
DevOps / Infrastructure Engineer
Location: Chennai
Experience: 5+ Years
Role: DevOps / Infrastructure Engineer
Job Description
We are looking for an experienced DevOps / Infrastructure Engineer with strong hands-on experience in Linux administration, containerization, Kubernetes, automation, monitoring, and troubleshooting.
Mandatory Skills
- 5+ years of experience in DevOps / Infrastructure Administration
- Strong hands-on experience with Linux Administration
- Experience with Docker and Kubernetes
- Monitoring tools: AppDynamics, Prometheus, Grafana
- Strong Shell Scripting / Python Scripting skills
- Hands-on experience with Ansible 4.1
- Strong troubleshooting and problem-solving skills
- Experience in infrastructure/application monitoring and production support
- Good understanding of DevOps practices and automation
Key Responsibilities
- Manage and support Linux-based infrastructure and production environments.
- Deploy, manage, and troubleshoot applications using Docker and Kubernetes.
- Develop and maintain automation scripts using Shell/Python.
- Automate infrastructure and configuration management using Ansible.
- Monitor applications and infrastructure using AppDynamics, Prometheus, and Grafana.
- Perform root-cause analysis and resolve infrastructure/application issues.
- Handle incidents, troubleshoot performance issues, and ensure system availability.
- Collaborate with development and operations teams to improve deployment and operational processes.
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
Strong Hands-On DevOps Engineer Profile with deep Kubernetes and multi-cloud experience
2
Mandatory (Experience 1): Must have 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.
3
Mandatory (Experience 2): Must be able to own end-to-end customer deployments across cloud and enterprise environments (the role covers cloud, BYOC, air-gapped, and data-center deployments).
4
Mandatory (Tech skill 1): Must have strong hands-on Kubernetes experience — troubleshooting, networking, workloads, storage, RBAC, and security
5
Mandatory (Tech skill 2): Must have hands-on Helm experience — deployment, templating, and troubleshooting.
6
Mandatory (Tech skill 3): Must have hands-on Terraform experience — infrastructure provisioning and automation
7
Mandatory (Tech skill 4): Must have strong Linux and cloud infrastructure troubleshooting skills
8
Mandatory (Tech skill 5): Must have hands-on experience with observability — metrics, logs, traces, and alerting
9
Mandatory (Tech skill 6): Must have hands-on experience with one or more of AWS / GCP / Azure / OCI
10
Mandatory (Tech skill 7): Must have a strong incident-management and production-troubleshooting mindset, with a "fix first, optimize later" approach to customer escalations.
11
Mandatory (Company): B2B SaaS product companies
12
Preferred (Enterprise deployments): Prior experience with BYOC, air-gapped, or restricted-network environments.
13
Preferred (Observability tools): Prometheus / Grafana / OpenTelemetry / ELK / OpenSearch / Datadog
14
Preferred (Other): GitOps / CI-CD; Python or Golang for automation; multi-tenant SaaS infrastructure scaling; exposure to AI/ML pipeline deployments or iPaaS / reverse ETL connectors; SRE concepts (SLIs/SLOs, DR, high availability
🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience Level : 4+ Years
Location : Gurugram Sector 48, Haryana (On-site)
Employment Type : Full Time Opportunity
About the Role :
We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.
In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).
Mandatory Skills :
AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux
Key Responsibilities :
1. Cloud Infrastructure & Infrastructure as Code (IaC) :
- Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
- Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
- Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.
2. CI/CD, Automation & Development :
- Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
- Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.
3. Containerization & Orchestration :
- Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
- Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.
4. Observability, SRE & Incident Management :
- Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
- Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
- Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
- Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).
Required Qualifications & Skills :
- Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
- Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
- Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
- CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
- Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
- Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
- Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
- Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
- Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Key Responsibilities
- Automate application deployments from Bitbucket to servers using CI/CD pipelines.
- Design and manage scalable, highly available AWS infrastructure.
- Implement Auto Scaling, ELB, and Route 53 for traffic management and high availability.
- Work with AWS services including IAM, RDS, DynamoDB, EC2, and other cloud services.
- Build and manage Docker containers and server images.
- Deploy and manage applications using Kubernetes.
- Implement Infrastructure as Code using Terraform, CloudFormation, or Ansible.
- Develop automation scripts using Python and Bash.
- Implement monitoring and logging using tools such as Prometheus, Grafana, and ELK.
- Integrate security and compliance practices into CI/CD pipelines.
- Optimize infrastructure for security, scalability, performance, and cost.
Required Skills
- 3+ years of experience in DevOps or a similar role.
- Strong knowledge of AWS beyond EC2.
- Hands-on experience with Jenkins or similar CI/CD tools.
- Experience with Docker and Kubernetes.
- Good understanding of Terraform/IaC and automation.
- Proficiency in Python and/or Bash scripting.
- Knowledge of DevSecOps, security, and compliance best practices.
- Strong troubleshooting and problem-solving skills.
Job Description:
Pre-requisite skills required for a DevOps Engineer role include:
- 6+yrs exp in DevOps
- Experience working on Linux based infrastructure
- knowledge in AWS, docker, CI/CD tools
- Hands on exp in Python/shell scripting language
- hands on exp in AWS and Azure
- Work exp in Docker, Terraform, Ansible, Kubernetes, LINUX
- Excellent understanding of Ruby, Python, Perl, and Java
- Configuration and managing databases such as MySQL, Mongo
- Excellent troubleshooting
- Working knowledge of various tools, open-source technologies, and cloud services
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 4+ Years
Location : Gurugram, Sector 48, Haryana (On-site)
Employment Type : Full-Time
Working Days : Monday to Saturday (1st & 3rd Saturday Off)
About the Role :
We are looking for a hands-on DevOps Engineer / Site Reliability Engineer (SRE) with strong experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, and production application deployments.
The ideal candidate should have real-world production experience, excellent troubleshooting skills, and the ability to manage both infrastructure and application-level issues.
Mandatory Skills :
Linux, AWS, Docker, Kubernetes, Terraform, Ansible, Jenkins, GitHub Actions, GitLab CI/CD, CI/CD, Infrastructure as Code (IaC), Python, Bash, Git, Grafana, Prometheus, ELK, CloudWatch, New Relic, SRE (SLI/SLO/SLA), Networking (DNS, HTTP/HTTPS, TCP/IP, Load Balancer), Production Application Deployment & Troubleshooting
Key Responsibilities :
- Manage and maintain AWS cloud infrastructure.
- Build and optimize CI/CD pipelines using Jenkins, GitHub Actions, or GitLab CI.
- Deploy, monitor, and troubleshoot applications across production environments.
- Automate infrastructure using Terraform and Ansible.
- Manage Docker containers and Kubernetes clusters.
- Monitor systems using Grafana, Prometheus, ELK, CloudWatch, and New Relic.
- Perform Linux server administration and troubleshooting.
- Handle production incidents, Root Cause Analysis (RCA), and improve system reliability.
- Collaborate with development teams to support application releases and automation.
Required Qualifications :
- Bachelor's degree in Computer Science or related field.
- 4+ years of hands-on experience in DevOps / SRE.
- Strong Linux administration and production troubleshooting skills.
- Experience with AWS and modern DevOps toolchains.
- Hands-on experience with application deployment and production support.
What We're Looking For :
- Strong practical Linux and cloud knowledge.
- Real production experience with application deployments.
- Ability to troubleshoot both infrastructure and application issues.
- Experience handling live production incidents.
- Excellent communication and problem-solving skills.
- Candidates should be comfortable with scenario-based technical discussions and demonstrate genuine hands-on expertise.
Interview Process :
- HR Screening
- Technical Round
- Client Technical Round
- Final Discussion
Note : The interview will focus on practical hands-on experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, application deployment, production troubleshooting, and real-world DevOps scenarios.






