DevOps Engineering Manager at IDFC Bank · Chennai · 8 - 12 years · ₹50L - ₹75L / yr · Posted 25 Nov 2024

Role & Responsibilities:
AWS Cloud Management:
• Lead the design, deployment, and management of AWS cloud infrastructure to ensure
scalability, security, and reliability.
• Oversee the implementation of best practices for cloud resource utilization.
Automated Provisioning:
• Drive the development and maintenance of automated provisioning processes for
infrastructure deployment, leveraging tools such as Terraform and Packer.
• Continuously enhance deployment workflows to optimize efficiency.
Financial Operations (FinOps):
• Implement and champion FinOps practices to optimize cloud costs and resource
utilization.
• Conduct regular cost analysis and identify opportunities for cost savings without
compromising performance.
Infrastructure as Code (IaC):
• Collaborate with teams to implement and maintain IaC scripts for infrastructure
configuration and deployment.
• Ensure version control and consistency in infrastructure code across projects.
Team Leadership:
• Lead and mentor a team of DevOps engineers, providing technical guidance and
support.\
• Foster a collaborative and innovative team culture focused on continuous improvement.
Continuous Integration/Continuous Deployment (CI/CD):
• Drive the implementation and maintenance of CI/CD pipelines to automate software
delivery processes.
• Ensure seamless and reliable application deployments across environments.
Monitoring and Optimization:
• Implement monitoring solutions for cloud resources and applications.
• Proactively identify and address performance bottlenecks, ensuring optimal system
performance.

About IDFC Bank
About
Similar jobs (10)
We are looking for a hands-on Senior AWS Cloud Engineer to lead the infrastructure build, optimization, automation, and production deployment of a Multi-Agent AI Chatbot Platform hosted on AWS. The development environment is already in place, and the successful candidate will drive the solution through testing, integrations, and production go-live.
Key Responsibilities
- Review, validate, and optimize existing Terraform code and AWS infrastructure.
- Establish and manage integrations with enterprise platforms such as ServiceNow, Workday, and other third-party systems.
- Design, build, and support secure, scalable, and highly available AWS environments.
- Implement and automate CI/CD pipelines and Infrastructure-as-Code practices.
- Lead infrastructure testing, performance tuning, and production readiness activities.
- Drive deployment and operationalization of the platform in the Production environment.
- Implement cloud governance, security, monitoring, and FinOps best practices.
- Troubleshoot and resolve complex cloud infrastructure issues.
Required Skills & Experience
- 10+ years of IT experience with strong expertise in AWS Cloud Engineering.
- Proven experience designing, deploying, and managing AWS production environments.
- Strong hands-on experience with Terraform and Infrastructure-as-Code.
- Experience with CI/CD pipeline automation and DevOps practices.
- Expertise in AWS services including VPC, IAM, EC2, S3, Lambda, CloudWatch, and networking.
- Experience in performance optimization, reliability, and cloud cost management (FinOps).
- Strong scripting and automation skills.
- Experience integrating enterprise applications through APIs and secure connectivity patterns.
Job Title: DevOps Manager
Experience: 8+ Years
Location: Pune
Employment Type: Full-time
About the Role
We are looking for an experienced DevOps Manager with strong hands-on technical expertise and leadership capabilities to manage cloud infrastructure, automation, CI/CD, deployments, and platform reliability. The ideal candidate should have strong experience in Shell Scripting, cloud platforms, Kubernetes, Terraform, and production DevOps environments.
Key Responsibilities
• Lead and mentor the DevOps engineering team.
• Design, implement, and optimize CI/CD pipelines and deployment processes.
• Manage scalable and highly available cloud infrastructure across AWS / Azure / GCP.
• Drive Infrastructure as Code (IaC) using Terraform and configuration automation using Ansible.
• Manage Docker and Kubernetes based environments.
• Develop and maintain automation scripts using Shell/Bash Scripting and Python.
• Implement monitoring, logging, alerting, and observability.
• Handle production incidents, troubleshooting, root cause analysis, and preventive actions.
• Drive infrastructure security, performance, scalability, and cloud cost optimization.
• Collaborate with Development, QA, Security, Architecture, and Delivery teams.
Required Skills & Qualifications
• 8+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or related roles.
• Strong hands-on experience in Shell / Bash Scripting.
• Experience with AWS / Azure / GCP cloud platforms.
• Strong knowledge of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
• Hands-on experience with Docker, Kubernetes, and Helm.
• Strong experience with Terraform and Ansible.
• Experience with Linux administration, Git, networking, IAM, and cloud security.
• Experience with monitoring tools such as Prometheus, Grafana, ELK, or equivalent.
• Strong troubleshooting, communication, stakeholder management, and team leadership skills.
Preferred Skills (Nice to Have)
• Python scripting experience.
• Cloud / Kubernetes / Terraform certifications.
• Experience with SRE, observability, high-availability architecture, and cloud cost optimization.
This is a senior role in our Application & Database Modernization pillar, on a specific mission: leading a large-scale VMware-to-AWS migration and landing it in production. VMware estates are exactly where modernization programs go to stall — sprawling dependency graphs, undocumented workloads, and a hypervisor bill that grows while the migration deck gathers dust. Our customers need that estate moved, cut over, and retired, on a date in the contract.
You'll lead the design and implementation of the target infrastructure: Landing Zones built to AWS best practices, EKS-based container platforms, Infrastructure as Code across the stack with Terraform and CloudFormation, and HA, DR, security, and backup strategies that hold up in production. Aedeon absorbs the discovery, dependency-mapping, and validation grind that would otherwise consume the project's first two quarters. You own the judgment calls, architecture tradeoffs, cutover sequencing, risk decisions and you carry them through to production. You'll also mentor engineers and shape long-term infrastructure strategy with customer stakeholders. If you want your migration experience to end in retired VMware clusters rather than revised project plans, this is the role.
What you will do?
- Lead the design and implementation of infrastructure for a large-scale VMware-to-AWS migration — from discovery through production cutover.
- Architect and build secure, scalable, and highly available AWS environments, including Landing Zones that follow AWS best practices.
- Design and implement containerized application platforms on Amazon EKS; exposure to ECS is a plus.
- Implement Infrastructure as Code using Terraform and AWS CloudFormation.
- Define and enforce best practices across security, backups, high availability (HA), disaster recovery (DR), monitoring, and operations.
- Enable configuration management using tools such as Ansible, Chef, or similar.
- Build CI/CD pipelines and manage the complete build and release lifecycle for customer applications.
- Drive automation across provisioning, deployment, and operational workflows.
- Retire legacy infrastructure and land cloud-native architectures in its place.
- Improve and maintain DevOps platforms: Jenkins, Git repositories, monitoring, and observability stacks.
- Provide L3-level support for complex infrastructure and platform issues.
- Work with stakeholders on technical strategy and long-term architecture.
- Mentor engineers, contribute to talent evaluation, and support team development.
What are we looking for?
- 6+ years of hands-on DevOps experience, with strong expertise in designing and managing cloud infrastructure.
- Strong experience in VMware-to-AWS migration projects or large-scale infrastructure migrations.
- 4+ years of Terraform and CloudFormation for Infrastructure as Code (IaC).
- 4+ years in configuration management, systems engineering, and managing production-grade infrastructure.
- Solid Linux and/or Windows administration background.
- Deep understanding of AWS services, including VPC, EC2, IAM, EKS, ECS, RDS, S3, Backup, CloudWatch, etc.
- Experience with Kubernetes (EKS), ECS, and Docker in production environments.
- Hands-on experience designing HA, DR, security, and backup strategies.
- Experience with Landing Zone setup, multi-account strategy, and AWS governance frameworks.
- Proficiency in Git with a strong understanding of branching and merging strategies.
- Experience with CI/CD pipelines, automation, and operational tooling.
- A problem-solving mindset that's proactive and oriented toward reliability, performance, and scalability.
You will be preferred if
- Experience across multiple cloud platforms (AWS, Azure, GCP).
- AWS certifications such as Solutions Architect Associate/Professional or DevOps Engineer Professional.
- Exposure to data platforms such as Amazon EMR, Redshift, Lake Formation, and SageMaker.
- Experience designing cloud architectures at an L3/Architect level.
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 5+ Years
Location : Gurugram, Haryana
Work Mode : On-site (Full-time)
About the Role :
We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.
Mandatory Skills :
AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux
Key Responsibilities :
- Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
- Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
- Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
- Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
- Automate operational tasks using Python, Bash, and Groovy scripting.
- Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.
Required Skills & Qualifications :
- Bachelor's degree in Computer Science, IT, Electronics, or a related field.
- 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
- Strong expertise in AWS, with exposure to Azure/GCP.
- Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
- Strong scripting skills in Python and Bash.
- Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Good understanding of Linux, networking, SQL, and cloud security best practices.
Preferred Skills :
- Experience with multi-cloud environments and DevSecOps practices.
- Knowledge of disaster recovery, automation, and microservices architecture.
- Strong troubleshooting, communication, and problem-solving skills.
The Role
As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that
keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and
observability that let a small, fast-moving team ship confidently — and you'll keep our AI and
data workloads reliable and affordable at scale.
This is a hands-on role with real ownership: you won't be maintaining someone else's setup,
you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make
deployment boring, incidents rare, and scaling a non-event. ---
What You'll Own
**CI/CD & developer experience**
- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with
confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible
automated testing, rollbacks, and release controls.
**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or
similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and
resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch
processing for the speech pipeline.
**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting,
on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and
processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in
transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security
requirements) for a product that handles sensitive customer conversations.
**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability
and spend. ---
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production
systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and
Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI,
Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong
automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog,
OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team.
Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch
pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data
pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud
ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. ---
Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real
scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data,
and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an
As a DevOps Engineer at YOYO, you'll own the infrastructure and delivery backbone that keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and observability that let a small, fast-moving team ship confidently and you'll keep our AI and data workloads reliable and affordable at scale. This is a hands-on role with real ownership: you won't be maintaining someone else's setup, you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make deployment boring, incidents rare, and scaling a non-event.
If you are Interested DM me on LinkedIn - Saquib Mundagnur
What You'll Own
CI/CD & developer experience - Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible automated testing, rollbacks, and release controls.Cloud infrastructure & IaC - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
Containers & orchestration - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch processing for the speech pipeline.
Reliability & observability (SRE) - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
Data & pipeline infrastructure - Support the infrastructure behind large-scale, edge-to-cloud data movement and processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
Security & compliance - Bake security into the platform: secrets management, IAM/least-privilege, encryption in transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security requirements) for a product that handles sensitive customer conversations.
Cost & scale - Own cloud cost visibility and optimization; make scaling decisions that balance reliability and spend.
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production systems at meaningful scale.
- Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and Infrastructure-as-Code (**Terraform** or equivalent).
- Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong automation-first mindset.
- Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, OpenTelemetry, or similar) and running incident response / on-call.
- A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team
Job Title : SDE 3 – Infrastructure Platform Engineer
Experience : 5.5 to 8.5 Years
Number of Positions : 2
Employment Type : C2H (Contract to Hire)
Work Mode : Remote during contractual period → 5 Days WFO after conversion
Contract Duration : 3 Months
Post-Conversion Location : Pune
Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred
(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)
Role Overview :
We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.
The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.
Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.
Key Responsibilities :
- Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
- Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
- Manage and optimize Docker and Kubernetes workloads.
- Build internal platform tools to improve developer productivity and engineering efficiency.
- Implement monitoring, logging, alerting, and observability solutions.
- Participate in incident response, RCA, postmortems, and reliability improvements.
- Design and improve CI/CD pipelines and deployment automation.
- Contribute to system design, architecture discussions, scalability, security, and cost optimization.
- Collaborate with application, data, and product engineering teams.
Required Skills :
- 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
- Strong hands-on experience with AWS, GCP, or Azure.
- Strong expertise in Terraform / CloudFormation.
- Experience with CI/CD, Docker, and Kubernetes.
- Strong programming skills in at least one of:
- Python, Go, Java, or Ruby.
- Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
- Experience working with production infrastructure and highly available systems.
- Strong troubleshooting and problem-solving skills.
Nice to Have :
- Experience with SRE practices and production on-call ownership.
- Experience in fintech, payments, banking, or transaction-heavy systems.
- Knowledge of cloud security, compliance, or FinOps/cost optimization.
- Experience building internal developer platforms or productivity tools.
- Previous product company experience.
Interview Process :
Round 1 : Take-Home Coding Assignment – Submit within 48 hours
Round 2 : Coding Assignment Discussion – 1 Hour
Round 3 : Technical Managerial Round – 30 Minutes
Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.
Ideal Candidate :
Strong Platform / Infrastructure Engineer with hands-on experience in :
Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations
Pure DevOps profiles without strong coding and platform engineering experience are not preferred.
About the Role
The non-negotiable is deep, hands-on infrastructure expertise spanning hybrid cloud and self-hosted systems. You will architect and maintain our unique infrastructure combining AWS CDN, bare metal servers, and Kubernetes clusters. This is not just maintenance work: you will establish company-wide DevOps policy, build security-hardened environments, and create the documentation and processes that scale with us. You will design our CI/CD pipelines, implement zero-trust networking, and guide the technical team on how to keep our production systems robustly online. This role is hands-on, autonomous, and sets the standard for how we approach infrastructure as the company grows.
What You'll Build
- Hybrid Infrastructure Management: Architect and maintain our unique infrastructure spanning AWS CDN, bare metal servers, and Kubernetes clusters for computationally intensive facial analysis workloads.
- CI/CD Pipeline Architecture: Design and implement CircleCI or Jenkins pipelines with comprehensive build testing, versioning, and change logging.
- Zero-Trust Networking: Build and maintain mesh topology networks using Tailscale or Wireguard to securely connect our hybrid infrastructure.
- Security-First Culture: Establish and enforce security policies including key rotation, access controls, compliance frameworks, and employee security management.
- Infrastructure as Code: Document and codify all infrastructure decisions, creating repeatable, auditable deployments.
- Containerization Strategy: Implement and optimize Docker/K8s deployments for our AI/ML workloads.
- Cost Optimization: Continue our approach of strategic compute placement using owned, rented or borrowed infrastructure where it makes financial sense without sacrificing security or reliability.
- Observability & Monitoring: Implement comprehensive logging, monitoring, and alerting across our distributed systems.
What We're Looking For
- 5+ years of DevOps or infrastructure engineering experience, with a track record of building from scratch
- Hybrid infrastructure expertise: experience managing both cloud (AWS) and self-hosted infrastructure, understanding the tradeoffs and security risks of each
- Kubernetes production experience: deep knowledge of cluster design, operations, and scaling
- Networking mastery: strong understanding of VPCs, mesh networks, VPNs, and zero-trust architectures
- Security-first mindset: experience with security compliance, key management, IAM policies, and hardening production systems
- CI/CD expertise: hands-on experience building robust pipelines for build testing before deployment
- Infrastructure as Code: proficiency with Terraform, Ansible, or similar tools
- Scripting and automation: strong Python, Bash, or Go skills for tooling and automation
- Policy and documentation: ability to establish best practices and document them clearly for team adoption
- Leadership mentality: comfortable setting standards and directing technical decisions, not just executing them
Nice to Have
- Experience architecting infrastructure for AI/ML workloads
- Background in a fast-moving startup or scale-up environment
- H ands-on experience with cost optimization across cloud and on-premises infrastructure
Why Join
- Opportunity to solve real healthcare problems with cutting-edge technology
- Well-funded startup with a strong market presence
- Work with advanced AI technology in a healthcare context
- Collaborate with a talented team in a fast-paced environment
- Competitive salary with equity options
- Performance and quarterly bonuses
- Professional development opportunities
Compensation and Logistics
- Remote, full-time
- Reports to: Head of Engineering
- Competitive based on experience
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
Key Responsibilities
- Automate application deployments from Bitbucket to servers using CI/CD pipelines.
- Design and manage scalable, highly available AWS infrastructure.
- Implement Auto Scaling, ELB, and Route 53 for traffic management and high availability.
- Work with AWS services including IAM, RDS, DynamoDB, EC2, and other cloud services.
- Build and manage Docker containers and server images.
- Deploy and manage applications using Kubernetes.
- Implement Infrastructure as Code using Terraform, CloudFormation, or Ansible.
- Develop automation scripts using Python and Bash.
- Implement monitoring and logging using tools such as Prometheus, Grafana, and ELK.
- Integrate security and compliance practices into CI/CD pipelines.
- Optimize infrastructure for security, scalability, performance, and cost.
Required Skills
- 3+ years of experience in DevOps or a similar role.
- Strong knowledge of AWS beyond EC2.
- Hands-on experience with Jenkins or similar CI/CD tools.
- Experience with Docker and Kubernetes.
- Good understanding of Terraform/IaC and automation.
- Proficiency in Python and/or Bash scripting.
- Knowledge of DevSecOps, security, and compliance best practices.
- Strong troubleshooting and problem-solving skills.










