Principal DevOps Engineer at Securin Labs · Chennai, Tamilnadu · 10 - 18 years · ₹30L - ₹45L / yr (ESOP available) · Bootstrapped · Posted 10 Jul 2026

Principal DevOps Engineer
Note - Screening Requirement: Please note that this position requires a minimum of 8+ years of hands-on Architecting DevOps/SRE/Platform Engineering experience specializing in AWS, EKS/Kubernetes, Terraform, Python, Jenkins, and AI workflows (Mandatory skills) from scratch.
We are looking for an absolute builder who has a proven track record of personally architecting, designing and setting up production-grade AWS EKS clusters entirely from scratch. If your experience is primarily limited to managing, maintaining, or optimizing pre-existing environments that were already stood up by another team, this is not the right opportunity for you.
Who are we?
Securin is an AI-driven cybersecurity company focused on proactive, adversarial exposure and vulnerability management. Our mission is to help organizations reduce cyber risk by identifying, prioritising, and remediating the issues that matter most. Powered by a seasoned team of threat researchers and status as a Certified Naming Authority (CNA), Securin combines artificial intelligence / machine learning, threat intelligence, and deep vulnerability research (including the Dark Web) to deliver an adversarial approach to cyber defense. We help enterprises shift from reactive patching to strategic, risk-based exposure and vulnerability management – driving smarter security decisions and faster remediation.
What do we promise?
We are a highly effective tech-enabled cybersecurity solutions provider and promise continual security posture improvement, enhanced attack surface visibility, and proactive prioritized remediation for every one of our client businesses.
What do we provide?
● A chance to be on the leading edge of cybersecurity and AI
● Ability to have direct impact on company growth and revenue strategy
● An opportunity to mentor and be mentored by experts in multiple disciplines
What do we deliver?
Securin helps organizations to identify and remediate the most dangerous exposures, vulnerabilities, and risks in their environment. We deliver predictive and definitive intelligence and facilitate proactive remediation to help organizations stay a step ahead of attackers. By utilising our cybersecurity solutions, our clients can have a proactive and holistic view of their security posture and protect their assets from even the most advanced and dynamic attacks.
Securin has been recognized by national and international organizations for its role in accelerating innovation in offensive and proactive security. Our combination of domain expertise, cutting-edge technology, and advanced tech-enabled cybersecurity solutions has made Securin a leader in the industry.
Core Technology Stack
AWS , EKS / Kubernetes, Jenkins , Python ,Terraform ,CI/CD , AI
Key Responsibilities:
● Architect and manage the end-to-end SaaS platform infrastructure on AWS, including EKS cluster design, VPC networking, IAM, and multi-region availability.
● Build, maintain, and optimize Jenkins-based CI/CD pipelines and develop Python automation scripts for provisioning, deployments, and runbook automation.
● Define and enforce platform SLOs/SLAs; own the observability strategy across logging, metrics, and tracing.
● Manage and participate in the on-call rotation; act as escalation point for P1/P2 incidents and drive post-incident reviews.
● Drive Infrastructure-as-Code (IaC) practices with Terraform/CloudFormation and champions a culture of automation and operational excellence.
● Collaborate cross-functionally with product, security, and engineering teams to align infrastructure roadmap with business goals.
● Identify opportunities to leverage AI to automate operational and DevOps workflows.
● Design and implement AI-assisted solutions for incident triaging, root cause analysis, log analysis, and performance optimization.
● Drive the adoption of AI-powered tools for infrastructure management, deployment automation, monitoring, and troubleshooting.
● Build intelligent workflows that reduce manual effort in release management, capacity planning, and operational support.
● Integrate AI capabilities into CI/CD pipelines to improve code quality, deployment reliability, and operational efficiency.
● Collaborate with engineering teams to automate repetitive tasks and improve developer productivity.
● Define best practices and governance for the safe and effective use of AI across DevOps processes.
● Measure and report on productivity gains, operational improvements, and cost savings achieved through AI adoption.
Requirements:
● 8+ years of experience in DevOps, SRE, or cloud infrastructure engineering roles.
● Deep hands-on expertise with AWS services (EC2, EKS, RDS, S3, IAM, VPC, CloudFront, Route53, Lambda, etc).
● Strong Kubernetes experience: cluster management, Helm, autoscaling (HPA/KEDA).
● Proficiency with Jenkins for complex CI/CD pipeline design and maintenance.
● Solid Python scripting skills for automation, tooling, and infrastructure management tasks.
● Experience with Infrastructure-as-Code using Terraform and/or AWS Cloud Formation.
● Proven track record of architecting and managing end-to-end SaaS products in a cloud-native environment.
● Strong understanding of networking fundamentals, security best practices, and compliance frameworks (SOC 2, ISO 27001 a plus).
● Hands-on experience with on-call processes and incident management frameworks.
Preferred Qualifications:
● AWS certifications: Solutions Architect Professional, DevOps Engineer Professional, or equivalent.
● Familiarity with service mesh, secrets management (Vault, AWS Secrets Manager), and zero-trust security models.
● Experience with multi-tenant SaaS architectures and tenant isolation strategies.
● Knowledge of FinOps principles and AWS cost management tooling.
● Experience with database DevOps: RDS, Aurora schema migrations, and backup strategies.
Core Competencies:
● Strategic Thinking – ability to translate business goals into scalable technical architecture.
● Operational Excellence – strong bias for reliability, automation, and continuous improvement.
● Communication – ability to clearly articulate complex technical topics to non-technical stakeholders.
● Ownership Mindset – proactively identifies and resolves risks without waiting to be asked.
● Resilience Under Pressure – calm and decisive during incidents; leads by example in high-stress situations.
Why should we connect?
We are a bunch of passionate cybersecurity professionals who are building a culture of security. Today, cybersecurity is no more a luxury but a necessity with a global market value of $150 billion.
At Securin, we live by a people-first approach. We firmly believe that our employees should enjoy what they do. For our employees, we provide a hybrid work environment with competitive best-in-industry pay, while providing them with an environment to learn, thrive, and grow. Our hybrid working environment allows employees to work from the comfort of their homes or the office if they choose to. For the right candidate, this will feel like your second home.
If you are passionate about cybersecurity just as we are, we would love to connect and share ideas

About Securin Labs
About
At Securin, we're building an AI-native cybersecurity platform that helps organizations stay ahead of evolving threats. We enable security teams to discover, validate, prioritize, and remediate exploitable risks before attackers can take advantage of them. By combining AI-driven intelligence with real-world threat context, we help businesses focus on the vulnerabilities that matter most and strengthen their overall security posture.
We believe cybersecurity should be proactive, intelligent, and actionable. Our platform empowers enterprises to reduce risk through continuous exposure management, automated prioritization, and seamless collaboration across security teams. If you're passionate about solving complex security challenges with cutting-edge technology and making a meaningful impact, you'll find an exciting opportunity to grow with us.
Similar jobs (10)
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 4+ Years
Location : Gurugram, Sector 48, Haryana (On-site)
Employment Type : Full-Time
Working Days : Monday to Saturday (1st & 3rd Saturday Off)
About the Role :
We are looking for a hands-on DevOps Engineer / Site Reliability Engineer (SRE) with strong experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, and production application deployments.
The ideal candidate should have real-world production experience, excellent troubleshooting skills, and the ability to manage both infrastructure and application-level issues.
Mandatory Skills :
Linux, AWS, Docker, Kubernetes, Terraform, Ansible, Jenkins, GitHub Actions, GitLab CI/CD, CI/CD, Infrastructure as Code (IaC), Python, Bash, Git, Grafana, Prometheus, ELK, CloudWatch, New Relic, SRE (SLI/SLO/SLA), Networking (DNS, HTTP/HTTPS, TCP/IP, Load Balancer), Production Application Deployment & Troubleshooting
Key Responsibilities :
- Manage and maintain AWS cloud infrastructure.
- Build and optimize CI/CD pipelines using Jenkins, GitHub Actions, or GitLab CI.
- Deploy, monitor, and troubleshoot applications across production environments.
- Automate infrastructure using Terraform and Ansible.
- Manage Docker containers and Kubernetes clusters.
- Monitor systems using Grafana, Prometheus, ELK, CloudWatch, and New Relic.
- Perform Linux server administration and troubleshooting.
- Handle production incidents, Root Cause Analysis (RCA), and improve system reliability.
- Collaborate with development teams to support application releases and automation.
Required Qualifications :
- Bachelor's degree in Computer Science or related field.
- 4+ years of hands-on experience in DevOps / SRE.
- Strong Linux administration and production troubleshooting skills.
- Experience with AWS and modern DevOps toolchains.
- Hands-on experience with application deployment and production support.
What We're Looking For :
- Strong practical Linux and cloud knowledge.
- Real production experience with application deployments.
- Ability to troubleshoot both infrastructure and application issues.
- Experience handling live production incidents.
- Excellent communication and problem-solving skills.
- Candidates should be comfortable with scenario-based technical discussions and demonstrate genuine hands-on expertise.
Interview Process :
- HR Screening
- Technical Round
- Client Technical Round
- Final Discussion
Note : The interview will focus on practical hands-on experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, application deployment, production troubleshooting, and real-world DevOps scenarios.
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
Key Responsibilities
- Automate application deployments from Bitbucket to servers using CI/CD pipelines.
- Design and manage scalable, highly available AWS infrastructure.
- Implement Auto Scaling, ELB, and Route 53 for traffic management and high availability.
- Work with AWS services including IAM, RDS, DynamoDB, EC2, and other cloud services.
- Build and manage Docker containers and server images.
- Deploy and manage applications using Kubernetes.
- Implement Infrastructure as Code using Terraform, CloudFormation, or Ansible.
- Develop automation scripts using Python and Bash.
- Implement monitoring and logging using tools such as Prometheus, Grafana, and ELK.
- Integrate security and compliance practices into CI/CD pipelines.
- Optimize infrastructure for security, scalability, performance, and cost.
Required Skills
- 3+ years of experience in DevOps or a similar role.
- Strong knowledge of AWS beyond EC2.
- Hands-on experience with Jenkins or similar CI/CD tools.
- Experience with Docker and Kubernetes.
- Good understanding of Terraform/IaC and automation.
- Proficiency in Python and/or Bash scripting.
- Knowledge of DevSecOps, security, and compliance best practices.
- Strong troubleshooting and problem-solving skills.
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 5+ Years
Location : Gurugram, Haryana
Work Mode : On-site (Full-time)
About the Role :
We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.
Mandatory Skills :
AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux
Key Responsibilities :
- Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
- Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
- Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
- Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
- Automate operational tasks using Python, Bash, and Groovy scripting.
- Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.
Required Skills & Qualifications :
- Bachelor's degree in Computer Science, IT, Electronics, or a related field.
- 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
- Strong expertise in AWS, with exposure to Azure/GCP.
- Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
- Strong scripting skills in Python and Bash.
- Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Good understanding of Linux, networking, SQL, and cloud security best practices.
Preferred Skills :
- Experience with multi-cloud environments and DevSecOps practices.
- Knowledge of disaster recovery, automation, and microservices architecture.
- Strong troubleshooting, communication, and problem-solving skills.
About the Role
The non-negotiable is deep, hands-on infrastructure expertise spanning hybrid cloud and self-hosted systems. You will architect and maintain our unique infrastructure combining AWS CDN, bare metal servers, and Kubernetes clusters. This is not just maintenance work: you will establish company-wide DevOps policy, build security-hardened environments, and create the documentation and processes that scale with us. You will design our CI/CD pipelines, implement zero-trust networking, and guide the technical team on how to keep our production systems robustly online. This role is hands-on, autonomous, and sets the standard for how we approach infrastructure as the company grows.
What You'll Build
- Hybrid Infrastructure Management: Architect and maintain our unique infrastructure spanning AWS CDN, bare metal servers, and Kubernetes clusters for computationally intensive facial analysis workloads.
- CI/CD Pipeline Architecture: Design and implement CircleCI or Jenkins pipelines with comprehensive build testing, versioning, and change logging.
- Zero-Trust Networking: Build and maintain mesh topology networks using Tailscale or Wireguard to securely connect our hybrid infrastructure.
- Security-First Culture: Establish and enforce security policies including key rotation, access controls, compliance frameworks, and employee security management.
- Infrastructure as Code: Document and codify all infrastructure decisions, creating repeatable, auditable deployments.
- Containerization Strategy: Implement and optimize Docker/K8s deployments for our AI/ML workloads.
- Cost Optimization: Continue our approach of strategic compute placement using owned, rented or borrowed infrastructure where it makes financial sense without sacrificing security or reliability.
- Observability & Monitoring: Implement comprehensive logging, monitoring, and alerting across our distributed systems.
What We're Looking For
- 5+ years of DevOps or infrastructure engineering experience, with a track record of building from scratch
- Hybrid infrastructure expertise: experience managing both cloud (AWS) and self-hosted infrastructure, understanding the tradeoffs and security risks of each
- Kubernetes production experience: deep knowledge of cluster design, operations, and scaling
- Networking mastery: strong understanding of VPCs, mesh networks, VPNs, and zero-trust architectures
- Security-first mindset: experience with security compliance, key management, IAM policies, and hardening production systems
- CI/CD expertise: hands-on experience building robust pipelines for build testing before deployment
- Infrastructure as Code: proficiency with Terraform, Ansible, or similar tools
- Scripting and automation: strong Python, Bash, or Go skills for tooling and automation
- Policy and documentation: ability to establish best practices and document them clearly for team adoption
- Leadership mentality: comfortable setting standards and directing technical decisions, not just executing them
Nice to Have
- Experience architecting infrastructure for AI/ML workloads
- Background in a fast-moving startup or scale-up environment
- H ands-on experience with cost optimization across cloud and on-premises infrastructure
Why Join
- Opportunity to solve real healthcare problems with cutting-edge technology
- Well-funded startup with a strong market presence
- Work with advanced AI technology in a healthcare context
- Collaborate with a talented team in a fast-paced environment
- Competitive salary with equity options
- Performance and quarterly bonuses
- Professional development opportunities
Compensation and Logistics
- Remote, full-time
- Reports to: Head of Engineering
- Competitive based on experience
Job Summary :
We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.
Key Responsibilities :
- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.
- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.
- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management
- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.
- Collaborate with developers to ensure smooth and frequent deployments to production.
- Manage versioning and rollback strategies for critical deployments.
- Containerization & Orchestration using Kubernetes.
- Containerize applications using Docker, and manage them using Kubernetes.
- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.
- Develop utilities and tools to enhance operational efficiency and reliability.
- Monitoring & Incident Management
- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.
- Optimize application and system performance through proactive monitoring and configuration tuning.
Desired Skills and Experience :
- Experience Required - 6+ yrs.
- Hands-on experience on cloud services like AWS, EKS etc.
- Ability to design a good cloud solution.
- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.
- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.
- Use knowledge and research to constantly modernize our applications and infrastructure stacks.
- Be a team player and strong problem-solver to work with a diverse team.
- Having good communication skills.
The Role
As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that
keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and
observability that let a small, fast-moving team ship confidently — and you'll keep our AI and
data workloads reliable and affordable at scale.
This is a hands-on role with real ownership: you won't be maintaining someone else's setup,
you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make
deployment boring, incidents rare, and scaling a non-event. ---
What You'll Own
**CI/CD & developer experience**
- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with
confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible
automated testing, rollbacks, and release controls.
**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or
similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and
resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch
processing for the speech pipeline.
**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting,
on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and
processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in
transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security
requirements) for a product that handles sensitive customer conversations.
**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability
and spend. ---
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production
systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and
Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI,
Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong
automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog,
OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team.
Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch
pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data
pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud
ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. ---
Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real
scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data,
and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an
🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience Level : 4+ Years
Location : Gurugram Sector 48, Haryana (On-site)
Employment Type : Full Time Opportunity
About the Role :
We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.
In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).
Mandatory Skills :
AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux
Key Responsibilities :
1. Cloud Infrastructure & Infrastructure as Code (IaC) :
- Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
- Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
- Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.
2. CI/CD, Automation & Development :
- Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
- Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.
3. Containerization & Orchestration :
- Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
- Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.
4. Observability, SRE & Incident Management :
- Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
- Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
- Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
- Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).
Required Qualifications & Skills :
- Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
- Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
- Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
- CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
- Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
- Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
- Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
- Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
- Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Strong Hands-On DevOps Engineer Profile with deep Kubernetes and multi-cloud experience
2
Mandatory (Experience 1): Must have 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.
3
Mandatory (Experience 2): Must be able to own end-to-end customer deployments across cloud and enterprise environments (the role covers cloud, BYOC, air-gapped, and data-center deployments).
4
Mandatory (Tech skill 1): Must have strong hands-on Kubernetes experience — troubleshooting, networking, workloads, storage, RBAC, and security
5
Mandatory (Tech skill 2): Must have hands-on Helm experience — deployment, templating, and troubleshooting.
6
Mandatory (Tech skill 3): Must have hands-on Terraform experience — infrastructure provisioning and automation
7
Mandatory (Tech skill 4): Must have strong Linux and cloud infrastructure troubleshooting skills
8
Mandatory (Tech skill 5): Must have hands-on experience with observability — metrics, logs, traces, and alerting
9
Mandatory (Tech skill 6): Must have hands-on experience with one or more of AWS / GCP / Azure / OCI
10
Mandatory (Tech skill 7): Must have a strong incident-management and production-troubleshooting mindset, with a "fix first, optimize later" approach to customer escalations.
11
Mandatory (Company): B2B SaaS product companies
12
Preferred (Enterprise deployments): Prior experience with BYOC, air-gapped, or restricted-network environments.
13
Preferred (Observability tools): Prometheus / Grafana / OpenTelemetry / ELK / OpenSearch / Datadog
14
Preferred (Other): GitOps / CI-CD; Python or Golang for automation; multi-tenant SaaS infrastructure scaling; exposure to AI/ML pipeline deployments or iPaaS / reverse ETL connectors; SRE concepts (SLIs/SLOs, DR, high availability
Company Overview:
Planview is hiring a DevOps Engineer in Bengaluru, India to support Planview SaaS applications across the product line. You will work in a global, collaborative team — owning CI/CD pipelines, cloud infrastructure, and automation to keep deployments fast and systems reliable.
Responsibilities
- Build and maintain CI/CD pipelines in Jenkins for reliable, fast delivery.
- Manage containerized workloads on Docker and ECS — task definitions, services, and clusters.
- Provision and manage AWS infrastructure using Terraform (CloudFormation a plus).
- Automate configuration and deployment tasks using Python (Ansible a plus).
- Set up and maintain monitoring and alerting via New Relic (CloudWatch, Datadog, or Prometheus/Grafana a plus).
- Write Shell and Python scripts to automate operations and reduce manual work.
- Manage Git workflows — branching, merge strategies, and pull request reviews.
- Administer and support MSSQL databases underpinning the product line — backups, restores, and basic performance troubleshooting.
- Troubleshoot deployment, performance, and infrastructure issues with development teams.
- Participate in on-call rotations and drive incident response.
- Continuously improve infrastructure resilience and deployment speed.
- Apply AI-assisted engineering tools (e.g., GitHub Copilot, Claude Code) to speed up IaC authoring, pipeline debugging, and day-to-day scripting.
Qualifications
Must-Have Skills
- Experience: 4–6 years of experience in DevOps, SRE, or Infrastructure Engineering.
- OS: Linux & Windows administration (systemd, package management, log analysis).
- Cloud: AWS (Active Directory, ECS, EC2, CloudFront, S3, VPC, IAM, RDS, Lambda basics).
- Source Control: Git — branching, merge/rebase, PR reviews.
- CI/CD (Jenkins): Pipeline creation and basic Groovy scripting.
- Containerization (Docker + ECS): Task definitions, services, and clusters.
- IaC: Terraform.
- Monitoring: New Relic.
- Scripting: Bash and Python scripting for automation.
- Networking Basics: DNS, load balancers, security groups, VPNs.
- Logging: ELK stack / CloudWatch Logs.
- Database: MSSQL administration — backups, restores, basic performance troubleshooting.
- Infrastructure Automation: Hands-on experience automating infrastructure provisioning, configuration, and deployment end-to-end.
- AI-Assisted Engineering: Comfortable working with AI coding/DevOps assistants (e.g., GitHub Copilot, Claude Code) for IaC generation, scripting, and troubleshooting — verified via a mandatory AI proficiency assessment during interviews.
Nice-to-Have Skills
• Configuration Management: Ansible.
• Additional IaC: CloudFormation.
• Architecture: Knowledge of microservices architecture.
• Cloudflare: DNS, CDN, WAF.
• Artifact Repositories: Nexus, JFrog Artifactory, ECR.
• Other CI/CD Tools: GitHub Actions.
• AWS cost optimization / FinOps awareness.
• Datadog, CloudWatch, or Prometheus/Grafana.
• AIOps: Exposure to AI-driven anomaly detection, root-cause analysis, or incident triage (e.g., Dynatrace Davis AI, Datadog Bits AI, Harness AIDA).
• Database Basics: RDS backups, restores, performance tuning.













