Sr. DevOps Engineer L3 · Mohali · 4 - 15 years · ₹10L - ₹20L / yr · Posted 25 Aug 2022

Sr. DevOps Engineer L3
Hands on experience in:
- Deploying, managing, securing and patching enterprise applications on large scale in Cloud preferably AWS.
- Experience leading End-to-end DevOps projects with modern tools encompassing both Applications and Infrastructure
- AWS Code deploy, Code build, Jenkins, Sonarqube.
- Incident management and root cause analysis.
- Strong understanding of immutable infrastructure and infrastructure as code concepts. Participate in capacity planning and provisioning of new resources. Importing already deployed infra into IaaC.
- Utilizing AWS cloud services such as EC2, S3, IAM, Route53, RDS, VPC, NAT/IG Gateway, LAMBDA, Load Balancers, CloudWatch, API Gateway are some of them.
- AWS ECS managing multi cluster container environments (ECS with EC2 and Fargate with service discovery using Route53)
- Monitoring/analytics tools like Nagios/DataDog and logging tools like LogStash/SumoLogic
- Simple Notification Service (SNS)
- Version Control System: Git, Gitlab, Bitbucket
- Participate in Security Audit of Cloud Infrastructure.
- Exceptional documentation and communication skills.
- Ready to work in Shift
- Knowledge of Akamai is Plus.
- Microsoft Azure is Plus
- Adobe AEM is plus.
- AWS Certified DevOps Professional is plus

Similar jobs (10)
Job Title : DevOps Engineer – Linux, AWS & Infrastructure
Experience : 3+ Years
Location : Sector 48, Gurgaon
Work Mode : 6 Days WFO – Monday to Saturday
Week Off : 01st & 03rd Saturday Off
Employment Type : Full-Time
Job Summary :
We are looking for a DevOps Engineer with strong hands-on experience in Linux, Networking, Server Administration, AWS, CI/CD, Containers, Kubernetes, and Infrastructure Automation. The ideal candidate should have strong troubleshooting skills, production ownership, and the ability to manage and automate infrastructure reliably.
Key Responsibilities :
- Manage and troubleshoot Linux servers, bare-metal infrastructure, and server environments.
- Perform system administration, networking, performance monitoring, and production troubleshooting.
- Design, maintain, and optimize Jenkins-based CI/CD pipelines and deployment workflows.
- Manage Docker containers and Kubernetes environments.
- Work with AWS Cloud services, infrastructure, security, and deployment environments.
- Monitor system health, application performance, logs, and infrastructure using appropriate monitoring and logging tools.
- Implement and maintain Infrastructure as Code (IaC) using tools such as Terraform or CloudFormation.
- Automate repetitive operational tasks using Python, Bash, Shell scripting, or similar technologies.
- Implement infrastructure and application security, access controls, patching, and hardening.
- Investigate production incidents, perform root-cause analysis (RCA), and drive issues to resolution.
- Take end-to-end ownership of infrastructure reliability, availability, and operational issues.
- Collaborate with development and other engineering teams to improve deployment, scalability, and system reliability.
Mandatory Skills :
Linux & Networking | Bare Metal / Server Administration | Jenkins / CI-CD | Docker / Containers | AWS Cloud | Kubernetes | Git | Monitoring & Logging | Security | Infrastructure as Code (IaC) | Automation / Scripting | Production Troubleshooting & Ownership
Preferred Skills :
- Strong understanding of TCP/IP, DNS, HTTP/HTTPS, SSH, load balancing, and networking fundamentals.
- Hands-on experience with Terraform / CloudFormation.
- Experience with Prometheus, Grafana, ELK / EFK, CloudWatch, or similar monitoring / logging tools.
- Good understanding of Linux performance troubleshooting, processes, memory, disk, networking, and file systems.
- Experience handling production incidents, RCA, deployments, and system reliability.
- Exposure to cloud security, IAM, secrets management, and server hardening.
What We’re Looking For :
- 3+ years of hands-on experience in DevOps / SRE / Infrastructure Engineering.
- Strong practical knowledge rather than certification-based/theoretical understanding.
- Good troubleshooting and analytical skills.
- Strong sense of ownership and accountability for production systems.
- Comfortable working in a 6-day work-from-office environment.
About the Role
We are looking for an experienced AWS DevOps Engineer with around 4 years of hands-on experience to join our infrastructure/DevOps team. The ideal candidate will independently design, deploy, and maintain cloud infrastructure primarily on AWS, drive automation initiatives, mentor junior engineers, and work closely with cross-functional teams to build scalable, secure, and highly available systems. Exposure to Azure is a strong plus.
Key Responsibilities
· Design, deploy, and maintain robust, scalable, and secure infrastructure on AWS
· Architect and manage core AWS services such as EC2, S3, VPC, IAM, RDS, Lambda, ECS/EKS, Route 53, and CloudFront
· Build, own, and optimize CI/CD pipelines (e.g., CodePipeline, Jenkins, GitLab CI, GitHub Actions) to enable fast and reliable deployments
· Design and implement Infrastructure as Code (IaC) using Terraform / AWS CloudFormation
· Set up and manage monitoring, logging, and alerting solutions (CloudWatch, ELK, Prometheus, Grafana, Datadog, etc.)
· Implement and enforce security best practices including IAM policies, security groups, NACLs, KMS, Secrets Manager, and compliance standards
· Lead troubleshooting and root cause analysis for infrastructure, deployment, and production incidents
· Drive backup, disaster recovery, high-availability, and cost-optimization strategies (Reserved Instances, Savings Plans, right-sizing)
· Containerize applications and manage orchestration using Docker, ECS, and/or Kubernetes (EKS)
· Automate repetitive operational tasks through scripting and tooling
· Support any hybrid or multi-cloud initiatives involving Azure services, where applicable
· Mentor junior engineers and review their work, providing technical guidance
· Collaborate with development, QA, security, and product teams to support and streamline application deployments
· Maintain comprehensive documentation of architecture, configurations, processes, and runbooks
· Participate in on-call rotations and incident response as needed
Required Skills & Qualifications
· Bachelor's degree in Computer Science, IT, or a related field (or equivalent practical experience)
· 4+ years of hands-on experience working with AWS cloud services in a production environment
· Strong expertise in core AWS services: EC2, S3, VPC, IAM, RDS, Lambda, CloudWatch, ECS/EKS, Route 53, ELB/ALB, Auto Scaling
· Solid understanding of networking concepts (subnets, routing, security groups, load balancers, VPNs, VPC peering, Direct Connect)
· Strong scripting/programming skills in Python, Bash, or PowerShell for automation
· Hands-on experience with Infrastructure as Code tools such as Terraform or AWS CloudFormation
· Proven experience with Linux and/or Windows server administration
· Strong understanding of CI/CD pipelines, Git-based version control, and branching strategies
· Solid experience with containerization and orchestration (Docker, ECS, or EKS/Kubernetes)
· Experience with configuration management tools (Ansible, Chef, or Puppet) is a plus
· Application Server Management — strong knowledge of networking, firewalls, load balancers, Nginx, Apache, etc.
· Ability to independently read, interpret AWS documentation, and troubleshoot complex issues
· Experience with cost optimization, security audits, and compliance frameworks (e.g., ISO, SOC2) is a plus
· AWS certification (Solutions Architect Associate/Professional, DevOps Engineer Professional) preferred
Good to Have
· Working knowledge of Microsoft Azure services (Virtual Machines, VNets, Azure DevOps, Azure Storage, Azure AD/Entra ID, AKS)
· Experience with multi-cloud or hybrid-cloud environments
· Familiarity with Azure Resource Manager (ARM) templates or Bicep
· Any Azure certification (AZ-104, AZ-400, etc.)
Soft Skills
· Strong analytical, problem-solving, and decision-making ability
· Excellent verbal and written communication skills
· Ability to mentor and guide junior team members
· Proactive, ownership-driven approach to infrastructure and incident management
· Strong collaboration skills in a cross-functional, team-oriented environment
· High attention to detail with a focus on reliability and scalability
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 4+ Years
Location : Gurugram, Sector 48, Haryana (On-site)
Employment Type : Full-Time
Working Days : Monday to Saturday (1st & 3rd Saturday Off)
About the Role :
We are looking for a hands-on DevOps Engineer / Site Reliability Engineer (SRE) with strong experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, and production application deployments.
The ideal candidate should have real-world production experience, excellent troubleshooting skills, and the ability to manage both infrastructure and application-level issues.
Mandatory Skills :
Linux, AWS, Docker, Kubernetes, Terraform, Ansible, Jenkins, GitHub Actions, GitLab CI/CD, CI/CD, Infrastructure as Code (IaC), Python, Bash, Git, Grafana, Prometheus, ELK, CloudWatch, New Relic, SRE (SLI/SLO/SLA), Networking (DNS, HTTP/HTTPS, TCP/IP, Load Balancer), Production Application Deployment & Troubleshooting
Key Responsibilities :
- Manage and maintain AWS cloud infrastructure.
- Build and optimize CI/CD pipelines using Jenkins, GitHub Actions, or GitLab CI.
- Deploy, monitor, and troubleshoot applications across production environments.
- Automate infrastructure using Terraform and Ansible.
- Manage Docker containers and Kubernetes clusters.
- Monitor systems using Grafana, Prometheus, ELK, CloudWatch, and New Relic.
- Perform Linux server administration and troubleshooting.
- Handle production incidents, Root Cause Analysis (RCA), and improve system reliability.
- Collaborate with development teams to support application releases and automation.
Required Qualifications :
- Bachelor's degree in Computer Science or related field.
- 4+ years of hands-on experience in DevOps / SRE.
- Strong Linux administration and production troubleshooting skills.
- Experience with AWS and modern DevOps toolchains.
- Hands-on experience with application deployment and production support.
What We're Looking For :
- Strong practical Linux and cloud knowledge.
- Real production experience with application deployments.
- Ability to troubleshoot both infrastructure and application issues.
- Experience handling live production incidents.
- Excellent communication and problem-solving skills.
- Candidates should be comfortable with scenario-based technical discussions and demonstrate genuine hands-on expertise.
Interview Process :
- HR Screening
- Technical Round
- Client Technical Round
- Final Discussion
Note : The interview will focus on practical hands-on experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, application deployment, production troubleshooting, and real-world DevOps scenarios.
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 5+ Years
Location : Gurugram, Haryana
Work Mode : On-site (Full-time)
About the Role :
We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.
Mandatory Skills :
AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux
Key Responsibilities :
- Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
- Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
- Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
- Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
- Automate operational tasks using Python, Bash, and Groovy scripting.
- Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.
Required Skills & Qualifications :
- Bachelor's degree in Computer Science, IT, Electronics, or a related field.
- 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
- Strong expertise in AWS, with exposure to Azure/GCP.
- Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
- Strong scripting skills in Python and Bash.
- Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Good understanding of Linux, networking, SQL, and cloud security best practices.
Preferred Skills :
- Experience with multi-cloud environments and DevSecOps practices.
- Knowledge of disaster recovery, automation, and microservices architecture.
- Strong troubleshooting, communication, and problem-solving skills.
Job Description:
- Infrastructure Management: Design, implement, and manage scalable, reliable, and secure cloud infrastructure using AWS, GCP, and/or Azure.
- CI/CD Pipelines: Develop and maintain continuous integration and continuous deployment (CI/CD) pipelines to streamline the development lifecycle.
- Automation: Automate infrastructure provisioning, configuration management, and application deployment processes.
- Monitoring and Performance: Implement monitoring, logging, and alerting solutions to ensure system health, performance, and reliability.
- Security: Ensure the security of cloud infrastructure and applications, including identity management and compliance with industry standards.
- Collaboration: Work closely with client and development teams to integrate DevOps practices and deliver high-quality software.
- Documentation: Maintain comprehensive documentation of infrastructure, configurations, and processes.
- Innovation: Stay current with emerging technologies and industry trends, integrating them into the DevOps strategy as appropriate.
Qualifications:
- Education: Bachelor's degree in Computer Science, Information Technology, or a related field.
- Experience: 7 - 10 years of overall experience with relevant experience of at least 7 years in DevOps and served as a lead or senior engineer.
Key Responsibilities
- Automate application deployments from Bitbucket to servers using CI/CD pipelines.
- Design and manage scalable, highly available AWS infrastructure.
- Implement Auto Scaling, ELB, and Route 53 for traffic management and high availability.
- Work with AWS services including IAM, RDS, DynamoDB, EC2, and other cloud services.
- Build and manage Docker containers and server images.
- Deploy and manage applications using Kubernetes.
- Implement Infrastructure as Code using Terraform, CloudFormation, or Ansible.
- Develop automation scripts using Python and Bash.
- Implement monitoring and logging using tools such as Prometheus, Grafana, and ELK.
- Integrate security and compliance practices into CI/CD pipelines.
- Optimize infrastructure for security, scalability, performance, and cost.
Required Skills
- 3+ years of experience in DevOps or a similar role.
- Strong knowledge of AWS beyond EC2.
- Hands-on experience with Jenkins or similar CI/CD tools.
- Experience with Docker and Kubernetes.
- Good understanding of Terraform/IaC and automation.
- Proficiency in Python and/or Bash scripting.
- Knowledge of DevSecOps, security, and compliance best practices.
- Strong troubleshooting and problem-solving skills.
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience Level : 4+ Years
Location : Gurugram Sector 48, Haryana (On-site)
Employment Type : Full Time Opportunity
About the Role :
We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.
In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).
Mandatory Skills :
AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux
Key Responsibilities :
1. Cloud Infrastructure & Infrastructure as Code (IaC) :
- Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
- Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
- Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.
2. CI/CD, Automation & Development :
- Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
- Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.
3. Containerization & Orchestration :
- Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
- Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.
4. Observability, SRE & Incident Management :
- Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
- Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
- Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
- Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).
Required Qualifications & Skills :
- Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
- Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
- Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
- CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
- Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
- Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
- Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
- Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
- Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Mactores is a trusted leader among businesses in providing modern data platform solutions. Since 2008, Mactores have been enabling businesses to accelerate their value through automation by providing End-to-End Data Solutions that are automated, agile, and secure. We collaborate with customers to strategize, navigate, and accelerate an ideal path forward with a digital transformation via assessments, migration, or modernization.
You will be part of the DevOps engineers' team, managing large customer deployments including Linux and Windows Administration, Large Enterprise Application, and Big Data Workloads. You will have broad business and technology expertise coupled with a background in professional services and client-facing skills. You are passionate about the best practices of cloud deployment and ensuring the customer expectation is set and met appropriately. You will help us build scalable, efficient cloud infrastructure.
You’ll implement monitoring for automated system health checks. Lastly, you’ll build our CI pipeline, and train and guide the team in DevOps practices. If you love to solve problems using your skills, then come join the Team Mactores. We have a casual and fun office environment that actively steers clear of rigid "corporate" culture, focuses on productivity and creativity, and allows you to be part of a world-class team while still being yourself.
What you will do?
- Application migration projects from on-premises to AWS.
- Database (RDBMS, NoSQL, DW, Hadoop) migration projects from on-premises to AWS.
- Automate operational and server provisioning workflows using AWS CFT on AWS.
- Share the responsibility for deploying releases and conducting other operations maintenance.
- Enhance operations infrastructures such as Jenkins clusters, Bitbucket, monitoring tools (Consul), and metrics tools such as Graphite and Grafana.
- Provide operational support for the rest of the Engineering team.
- Help migrate our remaining dedicated hardware infrastructure to the cloud.
- Establish and maintain operational best practices.
What do you have?
- 2+ years of experience in using Terraform for IaaC.
- 2+ years of configuration management and engineering for large scale customers, ideally supporting an Agile development process.
- 2+ years of Linux Administration experience.
- Deep understanding of version control systems (git), including branching and merging strategies.
- Experience working with cloud platforms (AWS/EC2/ ECS/ RDS/ CloudFormation, Cloudwatch, etc.) and cloud automation tools (Ansible, Chef).
- Experience with software build tools (Maven, Gradle) and continuous integration tools (Jenkins).
- Must have supported Java-based applications in a production environment.
- Experience with Linux environments and scripting languages - bash, python, Groovy.
- Experience in supporting Node.js in production is a plus.
- Knowledge of service discovery tools such as Consul is a plus.
- Comfortable working late evening hours, which is when most patching occurs.
- You are extremely proactive at identifying ways to improve things and make them more reliable.
You will be preferred if
- You are AWS DevOps Pro or AWS SA Pro Certified
Strong Hands-On DevOps Engineer Profile with deep Kubernetes and multi-cloud experience
2
Mandatory (Experience 1): Must have 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.
3
Mandatory (Experience 2): Must be able to own end-to-end customer deployments across cloud and enterprise environments (the role covers cloud, BYOC, air-gapped, and data-center deployments).
4
Mandatory (Tech skill 1): Must have strong hands-on Kubernetes experience — troubleshooting, networking, workloads, storage, RBAC, and security
5
Mandatory (Tech skill 2): Must have hands-on Helm experience — deployment, templating, and troubleshooting.
6
Mandatory (Tech skill 3): Must have hands-on Terraform experience — infrastructure provisioning and automation
7
Mandatory (Tech skill 4): Must have strong Linux and cloud infrastructure troubleshooting skills
8
Mandatory (Tech skill 5): Must have hands-on experience with observability — metrics, logs, traces, and alerting
9
Mandatory (Tech skill 6): Must have hands-on experience with one or more of AWS / GCP / Azure / OCI
10
Mandatory (Tech skill 7): Must have a strong incident-management and production-troubleshooting mindset, with a "fix first, optimize later" approach to customer escalations.
11
Mandatory (Company): B2B SaaS product companies
12
Preferred (Enterprise deployments): Prior experience with BYOC, air-gapped, or restricted-network environments.
13
Preferred (Observability tools): Prometheus / Grafana / OpenTelemetry / ELK / OpenSearch / Datadog
14
Preferred (Other): GitOps / CI-CD; Python or Golang for automation; multi-tenant SaaS infrastructure scaling; exposure to AI/ML pipeline deployments or iPaaS / reverse ETL connectors; SRE concepts (SLIs/SLOs, DR, high availability







