DevOps Engineer at Zocket · Chennai · 1 - 3 years · ₹5L - ₹12L / yr · Raised funding · Posted 21 May 2026

Why this role exists
Our infrastructure footprint is growing faster than our headcount, and we believe most of that
gap should be closed by automation and AI agents — not by hiring more humans to do toil. We
need someone early in their career who treats manual work as a bug, ships scripts and agents
instead of tickets, and wants to grow into deeper ownership over the next two years.
You will not be the most senior person on the team. You will be the one who multiplies the team.
What you'll own
In your first 1 months
• Take ownership of one slice of our CI/CD pipeline and make it measurably
faster, more reliable, or cheaper. We expect a number on a dashboard to move.
• Build at least three internal automations that replace manual ops toil —
using AI agents (Claude Code, agentic CLIs, scripted LLM workflows) as your force
multiplier.
• Be the first responder for a defined set of alerts. Write the runbooks. Drive
the alert volume down.
• Support senior engineers on AI/ML infrastructure (GPU nodes, inference
services, model deployment) — observe, document, and gradually take on contained
changes under review.
By 3 months you should be
• The go-to person for at least two production systems.
• Shipping routine infrastructure changes without needing senior review.
• Treating "manual" as a code smell.
Required (we will reject without these)
• 0–3 years hands-on experience with one major cloud (AWS, GCP, or
Azure — one is fine, depth beats breadth).
• Fluent in Linux command line, bash, and at least one scripting language
(Python or Go preferred).
• Have shipped something to production that real users hit. A side project
counts; a graded coursework lab does not.
• Comfortable with Docker — you can explain what an image vs. a
container is and why it matters.
• Working knowledge of networking fundamentals: DNS, HTTP/HTTPS,
TLS, ports, basic subnets — enough to debug "it works on my machine."
• Git fluency: branches, merges, rebases, conflict resolution.
• CI/CD pipelines — you have authored or substantially modified pipelines
in GitHub Actions, GitLab CI, ArgoCD, Jenkins, or similar. Not just "I clicked Re-run."
• Kubernetes basics — kubectl for real work, can read pod logs,
understand deployments and services, can debug a CrashLoopBackOff without
panicking. You do not need to have run a cluster; you do need to have lived inside one.
• Active user of AI coding agents (Claude Code, Cursor, Copilot, agentic
CLIs, etc.). You should be able to walk us through specific tasks where they made you
faster, and specific tasks where they failed you and how you noticed. "I have tried it" is
not enough.
Bonus (real plus, not required)
• Infrastructure as Code: Terraform, Pulumi, or Ansible.
• Observability: Prometheus/Grafana, Datadog, OpenTelemetry, any APM.
• Have built or extended an LLM-based agent — a custom MCP server, a
scripted multi-step workflow, an internal tool that calls models in a loop. Anything beyond
chat-with-Claude.
• Exposure to GPU workloads, model serving (vLLM, Triton, TGI, etc.), or
ML pipelines.
What we don't care about
• Whether your degree is in CS — or whether you have a degree at all.
• Brand-name companies on your resume.
• Certifications. They are fine. They do not substitute for having shipped.
How we work
• We default to automation. If you do something manually twice, the third
time you script it or hand it to an agent.
• AI agents are part of the workflow, not a novelty. Expect interview
questions about exactly how you use them — and where you have caught them being
wrong.
• Small, reversible changes beat big-bang rollouts.
• Postmortems are blameless and written down.
• We push back on each other. If you only execute, you will be unhappy
here.
How to apply
Send:
• Your resume.
• A short note (≤200 words) describing one infra or automation problem you
solved, and how AI agents factored in — or did not, and why. We read these. Generic
notes get rejected.
Internal note — delete before posting externally
• Comp band, location policy, team name, and reporting line marked
[CONFIRM] need to be filled in before this goes external.
• The Required list is intentionally tight: CI/CD and Kubernetes basics
promoted from bonus. Expect this to filter ~80% of typical junior DevOps applicants. The
remaining pool will skew toward people who have actually shipped infra at a startup, not
bootcamp grads or pure cloud-cert holders.
• IaC, observability, agent-building, and GPU/ML serving stay as bonus.
Promoting any of these to required at 0–3 yrs collapses the pool to near-zero or forces
hiring senior people at junior comp. If you want IaC required, re-level this to mid (3–5
yrs) and raise the band.
• Screening implication: the resume screen should explicitly check for
CI/CD pipeline authorship and any K8s-touching production work. If neither is on the
resume, reject at screen. Do not waste interview slots.
• Pipeline watch: if fewer than ~15 qualified resumes after 2 weeks of
active sourcing, the first thing to relax is the AI-agent-fluency bar (move to bonus and
screen for it in interview instead). Do not relax the "shipped to production" requirement
— that is the load-bearing filter.

About Zocket
About
Connect with the team
Similar jobs (10)
Company Overview:
Planview is hiring a DevOps Engineer in Bengaluru, India to support Planview SaaS applications across the product line. You will work in a global, collaborative team — owning CI/CD pipelines, cloud infrastructure, and automation to keep deployments fast and systems reliable.
Responsibilities
- Build and maintain CI/CD pipelines in Jenkins for reliable, fast delivery.
- Manage containerized workloads on Docker and ECS — task definitions, services, and clusters.
- Provision and manage AWS infrastructure using Terraform (CloudFormation a plus).
- Automate configuration and deployment tasks using Python (Ansible a plus).
- Set up and maintain monitoring and alerting via New Relic (CloudWatch, Datadog, or Prometheus/Grafana a plus).
- Write Shell and Python scripts to automate operations and reduce manual work.
- Manage Git workflows — branching, merge strategies, and pull request reviews.
- Administer and support MSSQL databases underpinning the product line — backups, restores, and basic performance troubleshooting.
- Troubleshoot deployment, performance, and infrastructure issues with development teams.
- Participate in on-call rotations and drive incident response.
- Continuously improve infrastructure resilience and deployment speed.
- Apply AI-assisted engineering tools (e.g., GitHub Copilot, Claude Code) to speed up IaC authoring, pipeline debugging, and day-to-day scripting.
Qualifications
Must-Have Skills
- Experience: 4–6 years of experience in DevOps, SRE, or Infrastructure Engineering.
- OS: Windows administration (systemd, package management, log analysis).
- Cloud: AWS (Active Directory, ECS, EC2, CloudFront, S3, VPC, IAM, RDS, Lambda basics).
- Source Control: Git — branching, merge/rebase, PR reviews.
- CI/CD (Jenkins): Pipeline creation and basic Groovy scripting.
- Containerization (Docker + ECS): Task definitions, services, and clusters.
- IaC: Terraform.
- Monitoring: New Relic.
- Scripting: Bash and Python scripting for automation.
- Networking Basics: DNS, load balancers, security groups, VPNs.
- Logging: ELK stack / CloudWatch Logs.
- Database: MSSQL administration — backups, restores, basic performance troubleshooting.
- Infrastructure Automation: Hands-on experience automating infrastructure provisioning, configuration, and deployment end-to-end.
- AI-Assisted Engineering: Comfortable working with AI coding/DevOps assistants (e.g., GitHub Copilot, Claude Code) for IaC generation, scripting, and troubleshooting — verified via a mandatory AI proficiency assessment during interviews.
Nice-to-Have Skills
• Configuration Management: Ansible.
• Additional IaC: CloudFormation.
• Architecture: Knowledge of microservices architecture.
• Cloudflare: DNS, CDN, WAF.
• Artifact Repositories: Nexus, JFrog Artifactory, ECR.
• Other CI/CD Tools: GitHub Actions.
• AWS cost optimization / FinOps awareness.
• Datadog, CloudWatch, or Prometheus/Grafana.
• AIOps: Exposure to AI-driven anomaly detection, root-cause analysis, or incident triage (e.g., Dynatrace Davis AI, Datadog Bits AI, Harness AIDA).
• Database Basics: RDS backups, restores, performance tuning.
Job Description:
Pre-requisite skills required for a DevOps Engineer role include:
- 6+yrs exp in DevOps
- Experience working on Linux based infrastructure
- knowledge in AWS, docker, CI/CD tools
- Hands on exp in Python/shell scripting language
- hands on exp in AWS and Azure
- Work exp in Docker, Terraform, Ansible, Kubernetes, LINUX
- Excellent understanding of Ruby, Python, Perl, and Java
- Configuration and managing databases such as MySQL, Mongo
- Excellent troubleshooting
- Working knowledge of various tools, open-source technologies, and cloud services
Job Summary :
We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.
Key Responsibilities :
- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.
- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.
- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management
- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.
- Collaborate with developers to ensure smooth and frequent deployments to production.
- Manage versioning and rollback strategies for critical deployments.
- Containerization & Orchestration using Kubernetes.
- Containerize applications using Docker, and manage them using Kubernetes.
- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.
- Develop utilities and tools to enhance operational efficiency and reliability.
- Monitoring & Incident Management
- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.
- Optimize application and system performance through proactive monitoring and configuration tuning.
Desired Skills and Experience :
- Experience Required - 6+ yrs.
- Hands-on experience on cloud services like AWS, EKS etc.
- Ability to design a good cloud solution.
- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.
- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.
- Use knowledge and research to constantly modernize our applications and infrastructure stacks.
- Be a team player and strong problem-solver to work with a diverse team.
- Having good communication skills.
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 4+ Years
Location : Gurugram, Sector 48, Haryana (On-site)
Employment Type : Full-Time
Working Days : Monday to Saturday (1st & 3rd Saturday Off)
About the Role :
We are looking for a hands-on DevOps Engineer / Site Reliability Engineer (SRE) with strong experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, and production application deployments.
The ideal candidate should have real-world production experience, excellent troubleshooting skills, and the ability to manage both infrastructure and application-level issues.
Mandatory Skills :
Linux, AWS, Docker, Kubernetes, Terraform, Ansible, Jenkins, GitHub Actions, GitLab CI/CD, CI/CD, Infrastructure as Code (IaC), Python, Bash, Git, Grafana, Prometheus, ELK, CloudWatch, New Relic, SRE (SLI/SLO/SLA), Networking (DNS, HTTP/HTTPS, TCP/IP, Load Balancer), Production Application Deployment & Troubleshooting
Key Responsibilities :
- Manage and maintain AWS cloud infrastructure.
- Build and optimize CI/CD pipelines using Jenkins, GitHub Actions, or GitLab CI.
- Deploy, monitor, and troubleshoot applications across production environments.
- Automate infrastructure using Terraform and Ansible.
- Manage Docker containers and Kubernetes clusters.
- Monitor systems using Grafana, Prometheus, ELK, CloudWatch, and New Relic.
- Perform Linux server administration and troubleshooting.
- Handle production incidents, Root Cause Analysis (RCA), and improve system reliability.
- Collaborate with development teams to support application releases and automation.
Required Qualifications :
- Bachelor's degree in Computer Science or related field.
- 4+ years of hands-on experience in DevOps / SRE.
- Strong Linux administration and production troubleshooting skills.
- Experience with AWS and modern DevOps toolchains.
- Hands-on experience with application deployment and production support.
What We're Looking For :
- Strong practical Linux and cloud knowledge.
- Real production experience with application deployments.
- Ability to troubleshoot both infrastructure and application issues.
- Experience handling live production incidents.
- Excellent communication and problem-solving skills.
- Candidates should be comfortable with scenario-based technical discussions and demonstrate genuine hands-on expertise.
Interview Process :
- HR Screening
- Technical Round
- Client Technical Round
- Final Discussion
Note : The interview will focus on practical hands-on experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, application deployment, production troubleshooting, and real-world DevOps scenarios.
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
As a DevOps Engineer at YOYO, you'll own the infrastructure and delivery backbone that keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and observability that let a small, fast-moving team ship confidently and you'll keep our AI and data workloads reliable and affordable at scale. This is a hands-on role with real ownership: you won't be maintaining someone else's setup, you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make deployment boring, incidents rare, and scaling a non-event.
If you are Interested DM me on LinkedIn - Saquib Mundagnur
What You'll Own
CI/CD & developer experience - Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible automated testing, rollbacks, and release controls.Cloud infrastructure & IaC - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
Containers & orchestration - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch processing for the speech pipeline.
Reliability & observability (SRE) - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
Data & pipeline infrastructure - Support the infrastructure behind large-scale, edge-to-cloud data movement and processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
Security & compliance - Bake security into the platform: secrets management, IAM/least-privilege, encryption in transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security requirements) for a product that handles sensitive customer conversations.
Cost & scale - Own cloud cost visibility and optimization; make scaling decisions that balance reliability and spend.
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production systems at meaningful scale.
- Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and Infrastructure-as-Code (**Terraform** or equivalent).
- Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong automation-first mindset.
- Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, OpenTelemetry, or similar) and running incident response / on-call.
- A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team
Key Responsibilities
- Automate application deployments from Bitbucket to servers using CI/CD pipelines.
- Design and manage scalable, highly available AWS infrastructure.
- Implement Auto Scaling, ELB, and Route 53 for traffic management and high availability.
- Work with AWS services including IAM, RDS, DynamoDB, EC2, and other cloud services.
- Build and manage Docker containers and server images.
- Deploy and manage applications using Kubernetes.
- Implement Infrastructure as Code using Terraform, CloudFormation, or Ansible.
- Develop automation scripts using Python and Bash.
- Implement monitoring and logging using tools such as Prometheus, Grafana, and ELK.
- Integrate security and compliance practices into CI/CD pipelines.
- Optimize infrastructure for security, scalability, performance, and cost.
Required Skills
- 3+ years of experience in DevOps or a similar role.
- Strong knowledge of AWS beyond EC2.
- Hands-on experience with Jenkins or similar CI/CD tools.
- Experience with Docker and Kubernetes.
- Good understanding of Terraform/IaC and automation.
- Proficiency in Python and/or Bash scripting.
- Knowledge of DevSecOps, security, and compliance best practices.
- Strong troubleshooting and problem-solving skills.
The Role
As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that
keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and
observability that let a small, fast-moving team ship confidently — and you'll keep our AI and
data workloads reliable and affordable at scale.
This is a hands-on role with real ownership: you won't be maintaining someone else's setup,
you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make
deployment boring, incidents rare, and scaling a non-event. ---
What You'll Own
**CI/CD & developer experience**
- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with
confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible
automated testing, rollbacks, and release controls.
**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or
similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and
resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch
processing for the speech pipeline.
**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting,
on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and
processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in
transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security
requirements) for a product that handles sensitive customer conversations.
**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability
and spend. ---
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production
systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and
Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI,
Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong
automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog,
OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team.
Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch
pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data
pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud
ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. ---
Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real
scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data,
and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
We are seeking a highly skilled and experienced DevOps Engineer to join our development team.
Years of experience needed: 8+ Yrs
Required Skills
Strong experience with GitLab, TeamCity, Terraform, Kubernetes and Docker
Proficiency in various deployment strategies and CI/CD pipeline setups
Solid understanding of Linux administration and Windows IIS.
Advanced scripting skills in bash and PowerShell
Hands-on experience with Google Cloud Platform (GCP) services, particularly GKE, GCS, GCE, Cloud SQL, and load balancers
Familiarity with Autosys for job scheduling and automation
Excellent problem-solving and troubleshooting skills
Ability to work independently and collaboratively in a fast-paced environment
- • Strong communication and interpersonal skills.













