Associate Principal Engineer, Linux Administrator at Digital Product Engineering company · Bengaluru (Bangalore) · 9 - 11 years · ₹32L - ₹35L / yr · Posted 1 Jun 2026

Associate Principal Engineer, Linux Administrator
at Digital Product Engineering company
Associate Principal Engineer, Linux Administrator
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience:9-11 years
Job description
REQUIREMENTS:
- Strong experience in DevOps, Platform Engineering, and Infrastructure Automation
- Deep hands-on expertise in Linux Administration (RHEL, CentOS, Ubuntu) – OS hardening, security, patching, and performance management (Must Have)
- Strong experience with Cloud Technologies – Public & Private Cloud environments (Must Have)
- Hands-on experience with Infrastructure as Code (IaC) using Terraform (Must Have)
- Strong automation expertise using Ansible for configuration management and infrastructure provisioning (Must Have)
- Experience building and managing CI/CD pipelines and end-to-end deployment automation
- Strong experience with Kubernetes administration, orchestration, and cluster management (Must Have)
- Hands-on experience with Docker containerization and Helm package management
- Experience managing large-scale development and infrastructure environments
- Strong understanding of Networking concepts, connectivity, design, troubleshooting, and network automation
- Experience with Observability & Monitoring tools and best practices
- Experience with Proxmox virtualization platform administration and management
- Knowledge of Edge Technologies and distributed infrastructure environments
- Basic understanding and administration of Active Directory (AD)
- Experience implementing AI-driven Automation solutions and operational efficiencies
- Strong understanding of infrastructure security, compliance, and governance
- Experience working in Agile/Scrum environments
- Strong troubleshooting, analytical, and problem-solving skills
- Excellent communication and stakeholder management skills
RESPONSIBILITIES:
- Design, build, and manage scalable infrastructure platforms across cloud and on-premise environments
- Administer and maintain Linux servers including security hardening, patching, performance tuning, and troubleshooting
- Develop and manage Infrastructure as Code (IaC) solutions using Terraform
- Automate infrastructure provisioning, configuration management, and operational tasks using Ansible
- Design, implement, and maintain CI/CD pipelines for application and infrastructure deployments
- Deploy, manage, and optimize Kubernetes clusters and containerized workloads
- Manage Docker environments and Helm-based application deployments
- Design and implement network solutions ensuring security, reliability, and scalability
- Monitor infrastructure health, performance, and availability using observability and monitoring tools
- Manage and support Proxmox virtualization environments
- Implement AI-driven automation initiatives to improve operational efficiency and reduce manual effort
- Support edge infrastructure deployments and distributed computing environments
- Collaborate with development, security, and operations teams to deliver reliable platform services
- Troubleshoot production incidents and perform root cause analysis
- Define infrastructure standards, automation frameworks, and operational best practices
- Ensure high availability, scalability, security, and reliability of infrastructure platforms
- Mentor junior engineers and provide technical leadership on DevOps and platform engineering initiatives
- Participate in Agile ceremonies and contribute to continuous improvement initiatives
- Work closely with stakeholders to understand infrastructure requirements and deliver optimal solutions
Qualifications
Bachelor’s or master’s degree in computer science, Information Technology, or a related fields

Similar jobs (10)
Job Title : SDE 3 – Infrastructure Platform Engineer
Experience : 5.5 to 8.5 Years
Number of Positions : 2
Employment Type : C2H (Contract to Hire)
Work Mode : Remote during contractual period → 5 Days WFO after conversion
Contract Duration : 3 Months
Post-Conversion Location : Pune
Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred
(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)
Role Overview :
We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.
The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.
Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.
Key Responsibilities :
- Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
- Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
- Manage and optimize Docker and Kubernetes workloads.
- Build internal platform tools to improve developer productivity and engineering efficiency.
- Implement monitoring, logging, alerting, and observability solutions.
- Participate in incident response, RCA, postmortems, and reliability improvements.
- Design and improve CI/CD pipelines and deployment automation.
- Contribute to system design, architecture discussions, scalability, security, and cost optimization.
- Collaborate with application, data, and product engineering teams.
Required Skills :
- 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
- Strong hands-on experience with AWS, GCP, or Azure.
- Strong expertise in Terraform / CloudFormation.
- Experience with CI/CD, Docker, and Kubernetes.
- Strong programming skills in at least one of:
- Python, Go, Java, or Ruby.
- Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
- Experience working with production infrastructure and highly available systems.
- Strong troubleshooting and problem-solving skills.
Nice to Have :
- Experience with SRE practices and production on-call ownership.
- Experience in fintech, payments, banking, or transaction-heavy systems.
- Knowledge of cloud security, compliance, or FinOps/cost optimization.
- Experience building internal developer platforms or productivity tools.
- Previous product company experience.
Interview Process :
Round 1 : Take-Home Coding Assignment – Submit within 48 hours
Round 2 : Coding Assignment Discussion – 1 Hour
Round 3 : Technical Managerial Round – 30 Minutes
Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.
Ideal Candidate :
Strong Platform / Infrastructure Engineer with hands-on experience in :
Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations
Pure DevOps profiles without strong coding and platform engineering experience are not preferred.
Job Summary :
We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.
Key Responsibilities :
- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.
- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.
- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management
- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.
- Collaborate with developers to ensure smooth and frequent deployments to production.
- Manage versioning and rollback strategies for critical deployments.
- Containerization & Orchestration using Kubernetes.
- Containerize applications using Docker, and manage them using Kubernetes.
- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.
- Develop utilities and tools to enhance operational efficiency and reliability.
- Monitoring & Incident Management
- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.
- Optimize application and system performance through proactive monitoring and configuration tuning.
Desired Skills and Experience :
- Experience Required - 6+ yrs.
- Hands-on experience on cloud services like AWS, EKS etc.
- Ability to design a good cloud solution.
- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.
- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.
- Use knowledge and research to constantly modernize our applications and infrastructure stacks.
- Be a team player and strong problem-solver to work with a diverse team.
- Having good communication skills.
Cloud Infrastructure Engineer – BANG | 10+ Years
Location: Bangalore
Experience: 10+ Years
Job Description:
- Design, implement, and manage cloud infrastructure across AWS/Azure/GCP environments.
- Strong experience in cloud architecture, compute, storage, networking, and security.
- Manage VMs, VPC/VNet, load balancers, DNS, DHCP, firewalls, and IAM.
- Hands-on experience with Windows/Linux servers, VMware, virtualization, and infrastructure operations.
- Automate infrastructure provisioning and configuration using Terraform, Ansible, or similar tools.
- Monitor infrastructure performance, availability, and capacity using tools such as Grafana, Prometheus, or CloudWatch/Azure Monitor.
- Handle incident management, troubleshooting, disaster recovery, backup, and high-availability requirements.
- Work with cross-functional teams to support cloud migration, infrastructure upgrades, and production environments.
- Ensure infrastructure follows security, compliance, and operational best practices.
Must-Have Skills:
Cloud Infrastructure | AWS/Azure/GCP | Networking | Linux/Windows | VMware | Terraform | Ansible | DNS/DHCP | IAM | Monitoring | Backup & DR
Job Description:
- Infrastructure Management: Design, implement, and manage scalable, reliable, and secure cloud infrastructure using AWS, GCP, and/or Azure.
- CI/CD Pipelines: Develop and maintain continuous integration and continuous deployment (CI/CD) pipelines to streamline the development lifecycle.
- Automation: Automate infrastructure provisioning, configuration management, and application deployment processes.
- Monitoring and Performance: Implement monitoring, logging, and alerting solutions to ensure system health, performance, and reliability.
- Security: Ensure the security of cloud infrastructure and applications, including identity management and compliance with industry standards.
- Collaboration: Work closely with client and development teams to integrate DevOps practices and deliver high-quality software.
- Documentation: Maintain comprehensive documentation of infrastructure, configurations, and processes.
- Innovation: Stay current with emerging technologies and industry trends, integrating them into the DevOps strategy as appropriate.
Qualifications:
- Education: Bachelor's degree in Computer Science, Information Technology, or a related field.
- Experience: 7 - 10 years of overall experience with relevant experience of at least 7 years in DevOps and served as a lead or senior engineer.

Key Skills:
• Bachelor's or Master's degree in Computer Science or related field.
• Minimum 5 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure Engineering.
• Experience migrating data and systems between AWS IaaS and PaaS.
• Experience operating and supporting applications using AWS VPC, EKS, and related services for multi-account operations.
• Experience developing fast and reliable Continuous Integration/Continuous Deployment (CI/CD) workflows used by hundreds of application teams.
• Experience administering and troubleshooting Operating Systems such as Linux, Windows, and MacOS.
• Professional Certifications in AWS Networks, CNCF Technologies, or Kubernetes.
• Experience using and configuring observability tools such as ELK, Prometheus/Grafana, AWS CloudWatch, and Jaeger.
• Experience of applied GitOps principles using ArgoCD or Flux.
• Public examples of code you've worked on with other people using any of these technologies:
o Configuration management/Infrastructure as Code (IAC) tools, such as AWS CDK, AWS CloudFormation, Terraform, Ansible, or Puppet.
o Systems solutions in one or more programming languages, such as Golang, Python, Java.
o Build, Release, Deploy or Ops Workflows using Bamboo, Argo Project, or GitHub Actions.
About the Team
SecurITe’s mission is to build an Agentic-AI driven security platform that protects critical infrastructure from modern cyber threats. Our focus is on delivering highly performant, resilient, and intelligent network security systems that help defenders stay ahead of adversaries.
About the Role
We’re looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.
As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premise environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.
You’ll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security.
This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.
Why This Role Matters
Cybersecurity is undergoing a fundamental shift. AI is no longer an enhancement—it’s becoming the core engine of how detection, investigation, and response are executed. As our Platform Engineer, you will architect and build the infrastructure, automation, deployment, and operational systems that make this transformation real.
Your work will directly influence the scalability, reliability, and security of our AI-driven cybersecurity platform across both cloud and enterprise on-premise deployments. You’ll help establish the operational backbone that enables rapid innovation, secure product delivery, and resilient large-scale deployments in mission-critical environments.
This is a chance to solve novel technical challenges involving distributed systems, hybrid infrastructure, observability, automation, and secure software delivery while shaping how defenders outpace modern attackers.
What You’ll Do
● Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, security groups)
● Administer and harden AlmaLinux VMs across production, staging, and dev environments
● Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling)
● Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch
● Lead incident response — diagnosis, RCA, and post-incident documentation — with no dedicated ops team to escalate to
● Make and document build-vs-buy and architecture decisions as the product and team scale
● Work directly with founders/engineering to translate ambiguous asks into scoped technical plans
Impact You’ll Have
● Accelerate engineering velocity through scalable developer platforms and automation
● Improve deployment reliability, platform uptime, and operational efficiency
● Enable secure and scalable AI-driven cybersecurity workloads
● Reduce operational overhead through infrastructure automation and self-service systems
● Help establish enterprise-grade cloud and on-premise deployment capabilities
● Enhance product resiliency, observability, and operational excellence
● Shape the long-term platform architecture powering next-generation cybersecurity products
● Enable rapid and secure delivery of critical security innovations to customers
Required Experience
● 4+ years hands-on Linux administration (RHEL-family strongly preferred — AlmaLinux, CentOS, RHEL)
● Deep Linux internals: systemd, networking, storage/LVM, process/resource management, kernel-level troubleshooting
● Real AWS architecture experience — not just operating existing infra, but designing it (VPC, EC2, IAM, security groups, networking)
● Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to
● Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts — built tooling that runs unattended
● Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one)
● Clear, proactive communicator — documents decisions and explains reasoning without being asked
Required Skills & Qualifications
● Strong Linux system administration and troubleshooting skills
● Redhat certifications
● Strong understanding of networking fundamentals, security, and distributed systems
● Proficiency with Docker, and container orchestration
● Experience with Terraform, Ansible, or similar infrastructure automation tools
● Strong scripting or programming skills in Python, Bash, or Go
● Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry
● Understanding of platform security best practices and secure infrastructure design
● Familiarity with virtualization technologies and hybrid infrastructure environments
● Strong problem-solving and debugging abilities
● Excellent communication and collaboration skills
● Ability to thrive in fast-paced startup environments
Nice to Have
● Configuration management/automation at scale (Ansible, AWX, Terraform, or similar)
● Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog)
● Container experience (Docker; Kubernetes a plus but not core to this VM-based stack)
● Experience in a startup or small-team environment where infra was built from scratch
● Security/compliance exposure (vulnerability remediation, hardening, SSO/access control)
The Mindset
Problem Solver
You thrive on complex, ambiguous challenges and engineer elegant solutions.
Ownership-Driven
You take initiative, move fast, and deliver outcomes without hand-holding.
Continuous Learner
You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.
Startup DNA
You excel in fast-moving environments where priorities evolve and impact is immediate.
This is a senior role in our Application & Database Modernization pillar, on a specific mission: leading a large-scale VMware-to-AWS migration and landing it in production. VMware estates are exactly where modernization programs go to stall — sprawling dependency graphs, undocumented workloads, and a hypervisor bill that grows while the migration deck gathers dust. Our customers need that estate moved, cut over, and retired, on a date in the contract.
You'll lead the design and implementation of the target infrastructure: Landing Zones built to AWS best practices, EKS-based container platforms, Infrastructure as Code across the stack with Terraform and CloudFormation, and HA, DR, security, and backup strategies that hold up in production. Aedeon absorbs the discovery, dependency-mapping, and validation grind that would otherwise consume the project's first two quarters. You own the judgment calls, architecture tradeoffs, cutover sequencing, risk decisions and you carry them through to production. You'll also mentor engineers and shape long-term infrastructure strategy with customer stakeholders. If you want your migration experience to end in retired VMware clusters rather than revised project plans, this is the role.
What you will do?
- Lead the design and implementation of infrastructure for a large-scale VMware-to-AWS migration — from discovery through production cutover.
- Architect and build secure, scalable, and highly available AWS environments, including Landing Zones that follow AWS best practices.
- Design and implement containerized application platforms on Amazon EKS; exposure to ECS is a plus.
- Implement Infrastructure as Code using Terraform and AWS CloudFormation.
- Define and enforce best practices across security, backups, high availability (HA), disaster recovery (DR), monitoring, and operations.
- Enable configuration management using tools such as Ansible, Chef, or similar.
- Build CI/CD pipelines and manage the complete build and release lifecycle for customer applications.
- Drive automation across provisioning, deployment, and operational workflows.
- Retire legacy infrastructure and land cloud-native architectures in its place.
- Improve and maintain DevOps platforms: Jenkins, Git repositories, monitoring, and observability stacks.
- Provide L3-level support for complex infrastructure and platform issues.
- Work with stakeholders on technical strategy and long-term architecture.
- Mentor engineers, contribute to talent evaluation, and support team development.
What are we looking for?
- 6+ years of hands-on DevOps experience, with strong expertise in designing and managing cloud infrastructure.
- Strong experience in VMware-to-AWS migration projects or large-scale infrastructure migrations.
- 4+ years of Terraform and CloudFormation for Infrastructure as Code (IaC).
- 4+ years in configuration management, systems engineering, and managing production-grade infrastructure.
- Solid Linux and/or Windows administration background.
- Deep understanding of AWS services, including VPC, EC2, IAM, EKS, ECS, RDS, S3, Backup, CloudWatch, etc.
- Experience with Kubernetes (EKS), ECS, and Docker in production environments.
- Hands-on experience designing HA, DR, security, and backup strategies.
- Experience with Landing Zone setup, multi-account strategy, and AWS governance frameworks.
- Proficiency in Git with a strong understanding of branching and merging strategies.
- Experience with CI/CD pipelines, automation, and operational tooling.
- A problem-solving mindset that's proactive and oriented toward reliability, performance, and scalability.
You will be preferred if
- Experience across multiple cloud platforms (AWS, Azure, GCP).
- AWS certifications such as Solutions Architect Associate/Professional or DevOps Engineer Professional.
- Exposure to data platforms such as Amazon EMR, Redshift, Lake Formation, and SageMaker.
- Experience designing cloud architectures at an L3/Architect level.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
We are seeking a skilled and experienced Senior DevOps Engineer to join our team. As a Senior DevOps Engineer, you will play a crucial role in designing, implementing, and managing our cloud infrastructure on AWS, while utilizing Docker, Kubernetes, Terraform and Argo CD for containerization and deployment. You will collaborate closely with development teams to ensure smooth software releases, scalability, and reliability of our systems. Also, you will be working on Airflow and Spark configurations for Data pipelines.
Responsibilities:
- Design, implement, and maintain scalable and secure cloud infrastructure on AWS, utilizing services such as EC2, S3, VPC, IAM, and Lambda.
- Build, deploy, and manage containerized applications using Docker and Kubernetes.
- Develop and maintain CI/CD pipelines to enable continuous integration, automated testing, and deployment using tools like Jenkins, GitLab CI/CD, or AWS CodePipeline.
- Implement and manage infrastructure as code (IaC) using tools like Terraform or CloudFormation to automate the provisioning and configuration of AWS resources.
- Monitor and troubleshoot infrastructure and application issues, ensuring high availability, performance, and scalability of the system.
- Collaborate with development teams to ensure smooth and efficient software releases, including version control, branching strategies, and release management.
- Implement and maintain robust security practices, including access control, network security, and data encryption.
- Automate manual tasks and processes using scripting languages (e.g., Python, Bash) and configuration management tools (e.g., Ansible).
- Collaborate with cross-functional teams to gather requirements, provide technical guidance, and implement best practices for infrastructure and deployment processes.
- Stay up-to-date with the latest DevOps tools, technologies, and best practices, and identify opportunities for improvement and optimization in our infrastructure and processes.
Requirements:
- Bachelor’s or master’s degree in computer science, Software Engineering, or a related field.
- Proven experience as a DevOps Engineer, with a focus on AWS, Docker, Kubernetes, and Argo CD.
- Strong experience in designing, implementing, and managing cloud infrastructure on AWS, including services like EC2, S3, VPC, IAM, and Lambda.
- Hands-on experience with containerization technologies, specifically Docker and Kubernetes.
- Proficiency in CI/CD pipeline setup and management using tools like Jenkins, GitLab CI/CD, or AWS CodePipeline.
- Solid understanding of infrastructure as code (IaC) principles and experience with tools like Terraform or CloudFormation.
- Experience in monitoring and troubleshooting complex infrastructure and application issues.
- Strong scripting and automation skills using languages such as Python, Bash, or PowerShell.
- Excellent problem-solving and analytical skills, with the ability to handle and prioritize multiple tasks in a fast-paced environment.
- Strong communication and collaboration skills, with the ability to work effectively with cross-functional teams.
Preferred Qualifications:
- AWS certifications, such as AWS Certified DevOps Engineer or AWS Certified Solutions Architect.
- Experience with infrastructure monitoring and logging tools like CloudWatch, ELK stack, or Prometheus/Grafana.
- Knowledge of serverless architectures and experience with AWS Lambda.
- Familiarity with other cloud platforms (e.g., Azure, Google Cloud Platform).
- Understanding of security best practices and experience implementing security measures in cloud infrastructure.
Job Description:
Pre-requisite skills required for a DevOps Engineer role include:
- 6+yrs exp in DevOps
- Experience working on Linux based infrastructure
- knowledge in AWS, docker, CI/CD tools
- Hands on exp in Python/shell scripting language
- hands on exp in AWS and Azure
- Work exp in Docker, Terraform, Ansible, Kubernetes, LINUX
- Excellent understanding of Ruby, Python, Perl, and Java
- Configuration and managing databases such as MySQL, Mongo
- Excellent troubleshooting
- Working knowledge of various tools, open-source technologies, and cloud services
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)





