Senior Linux Infrastructure Engineer (Bare Metal Storage & AI Factory) at Uvation India Pvt Ltd · Remote only · 9 - 18 years · ₹10L - ₹30L / yr · Remote only · Posted 27 Aug 2026

Senior Linux Infrastructure Engineer (Bare Metal Storage & AI Factory)
Job Overview
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms.
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
- Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
- Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
- Strong understanding of server hardware, including:
- BIOS/UEFI
- RAID controllers
- Firmware management
- iLO/iDRAC/IPMI
- NICs and SmartNICs
- HBA cards
- Hardware diagnostics and troubleshooting
- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
- Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
- Understanding of NVIDIA GPU technologies including:
- A100, H100, H200, B200, or equivalent GPU platforms
- NVIDIA DGX and OEM GPU servers
- GPU provisioning and lifecycle management
- GPU monitoring and performance optimization
- Knowledge of AI Factory architecture and infrastructure requirements
- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
- Understanding of:
- GPU resource allocation and scheduling
- Multi-GPU systems
- GPU networking requirements
- High-bandwidth, low-latency infrastructure design
- Familiarity with NVIDIA ecosystem technologies such as:
- CUDA
- NCCL
- GPUDirect Storage
- NVIDIA Fabric Manager
- NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
- Advanced Linux storage administration:
- LVM
- XFS, EXT4
- NFS
- iSCSI
- Fibre Channel SAN
- Multipath I/O
- Strong hands-on experience with Ceph, including:
- Cluster architecture
- MON, OSD, MDS
- RBD, CephFS, RGW
- Capacity planning
- Performance tuning
- Failure recovery
- Experience with high-performance AI storage platforms such as:
- WEKA
- VAST Data
- Dell PowerScale
- Pure Storage FlashBlade
- NetApp
- Understanding of:
- NVMe-over-Fabrics (NVMe-oF)
- RDMA
- GPUDirect Storage
- Parallel file systems
- AI data pipelines
Networking & Infrastructure
- Strong networking knowledge:
- Bonding
- VLANs
- Routing
- MTU optimization
- DNS
- DHCP
- Experience with high-performance data center networking:
- 100G/200G/400G Ethernet
- RoCE
- RDMA
- Spine-Leaf architectures
- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
- Experience with high availability, clustering, and disaster recovery
- Strong troubleshooting skills across:
- Linux operating systems
- Hardware platforms
- GPU infrastructure
- Networking
- Enterprise storage
- Experience supporting mission-critical production environments
- Bash and Python scripting for automation and operational efficiency
- Experience creating operational documentation, runbooks, and infrastructure standards
Nice to Have
- Kubernetes infrastructure (especially AI/ML and GPU integration)
- KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
- Ansible automation
- NVIDIA Base Command Manager
- Slurm or HPC workload schedulers
- Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
- Data Center Infrastructure Management (DCIM) tools
- IPAM solutions
- AWS, Azure, or hybrid cloud exposure

Similar jobs (10)
Cloud Infrastructure Engineer – BANG | 10+ Years
Location: Bangalore
Experience: 10+ Years
Job Description:
- Design, implement, and manage cloud infrastructure across AWS/Azure/GCP environments.
- Strong experience in cloud architecture, compute, storage, networking, and security.
- Manage VMs, VPC/VNet, load balancers, DNS, DHCP, firewalls, and IAM.
- Hands-on experience with Windows/Linux servers, VMware, virtualization, and infrastructure operations.
- Automate infrastructure provisioning and configuration using Terraform, Ansible, or similar tools.
- Monitor infrastructure performance, availability, and capacity using tools such as Grafana, Prometheus, or CloudWatch/Azure Monitor.
- Handle incident management, troubleshooting, disaster recovery, backup, and high-availability requirements.
- Work with cross-functional teams to support cloud migration, infrastructure upgrades, and production environments.
- Ensure infrastructure follows security, compliance, and operational best practices.
Must-Have Skills:
Cloud Infrastructure | AWS/Azure/GCP | Networking | Linux/Windows | VMware | Terraform | Ansible | DNS/DHCP | IAM | Monitoring | Backup & DR
Job Title : DevOps / Infrastructure Engineer
Experience : 3+ Years
Location : Gurugram Sector 48
Work Mode : 6 Days WFO (Monday to Saturday) / 01st & 03rd Saturdays are off
Employment Type : Full-time
Role Overview :
We are looking for a DevOps / Infrastructure Engineer with strong hands-on experience in Linux administration, Jenkins, Docker, networking, bare-metal servers, AWS, Redis, and MongoDB. The candidate should be capable of independently troubleshooting infrastructure, deployment, networking, and application-related issues in production environments.
Mandatory / Non-Negotiable Skills :
- Strong hands-on experience with Linux Administration & Troubleshooting.Strong experience with Jenkins and CI/CD pipelines.
- Hands-on experience with Docker and containerized environments.
- Strong understanding of Networking concepts – TCP/IP, DNS, HTTP/HTTPS, ports, routing, firewalls, load balancing, etc.
- Hands-on experience with Bare Metal Servers / Server Administration.
- Strong hands-on experience with AWS (EC2, VPC, IAM, Security Groups, Load Balancers, S3 & CloudWatch)
- Working knowledge of Redis
- Working knowledge of MongoDB
- Strong production troubleshooting and incident-resolution skills
Key Responsibilities :
- Manage, configure, monitor, and troubleshoot Linux and bare-metal servers
- Build, maintain, and troubleshoot Jenkins CI/CD pipelines
- Deploy, manage, and troubleshoot applications using Docker
- Manage and troubleshoot AWS infrastructure and services
- Configure and maintain networking, security groups, firewalls, ports, and connectivity
- Support and maintain Redis and MongoDB environments
- Perform server health checks, log analysis, performance troubleshooting, and incident resolution
- Work closely with development teams to support application deployments
- Identify root causes of infrastructure and production issues and implement preventive solutions
- Maintain infrastructure security, availability, and reliability
- Automate repetitive operational tasks wherever possible
Required Candidate Profile :
- 3+ years of relevant experience in DevOps, Infrastructure, System Administration, or related roles.
- Strong hands-on / production experience with all mandatory technologies.
- Good understanding of Linux systems and infrastructure.
- Strong troubleshooting and problem-solving abilities.
- Ability to take ownership of production infrastructure and deployment issues.
- Good communication and collaboration skills.
About the Team
SecurITe’s mission is to build an Agentic-AI driven security platform that protects critical infrastructure from modern cyber threats. Our focus is on delivering highly performant, resilient, and intelligent network security systems that help defenders stay ahead of adversaries.
About the Role
We’re looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.
As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premise environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.
You’ll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security.
This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.
Why This Role Matters
Cybersecurity is undergoing a fundamental shift. AI is no longer an enhancement—it’s becoming the core engine of how detection, investigation, and response are executed. As our Platform Engineer, you will architect and build the infrastructure, automation, deployment, and operational systems that make this transformation real.
Your work will directly influence the scalability, reliability, and security of our AI-driven cybersecurity platform across both cloud and enterprise on-premise deployments. You’ll help establish the operational backbone that enables rapid innovation, secure product delivery, and resilient large-scale deployments in mission-critical environments.
This is a chance to solve novel technical challenges involving distributed systems, hybrid infrastructure, observability, automation, and secure software delivery while shaping how defenders outpace modern attackers.
What You’ll Do
● Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, security groups)
● Administer and harden AlmaLinux VMs across production, staging, and dev environments
● Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling)
● Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch
● Lead incident response — diagnosis, RCA, and post-incident documentation — with no dedicated ops team to escalate to
● Make and document build-vs-buy and architecture decisions as the product and team scale
● Work directly with founders/engineering to translate ambiguous asks into scoped technical plans
Impact You’ll Have
● Accelerate engineering velocity through scalable developer platforms and automation
● Improve deployment reliability, platform uptime, and operational efficiency
● Enable secure and scalable AI-driven cybersecurity workloads
● Reduce operational overhead through infrastructure automation and self-service systems
● Help establish enterprise-grade cloud and on-premise deployment capabilities
● Enhance product resiliency, observability, and operational excellence
● Shape the long-term platform architecture powering next-generation cybersecurity products
● Enable rapid and secure delivery of critical security innovations to customers
Required Experience
● 4+ years hands-on Linux administration (RHEL-family strongly preferred — AlmaLinux, CentOS, RHEL)
● Deep Linux internals: systemd, networking, storage/LVM, process/resource management, kernel-level troubleshooting
● Real AWS architecture experience — not just operating existing infra, but designing it (VPC, EC2, IAM, security groups, networking)
● Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to
● Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts — built tooling that runs unattended
● Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one)
● Clear, proactive communicator — documents decisions and explains reasoning without being asked
Required Skills & Qualifications
● Strong Linux system administration and troubleshooting skills
● Redhat certifications
● Strong understanding of networking fundamentals, security, and distributed systems
● Proficiency with Docker, and container orchestration
● Experience with Terraform, Ansible, or similar infrastructure automation tools
● Strong scripting or programming skills in Python, Bash, or Go
● Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry
● Understanding of platform security best practices and secure infrastructure design
● Familiarity with virtualization technologies and hybrid infrastructure environments
● Strong problem-solving and debugging abilities
● Excellent communication and collaboration skills
● Ability to thrive in fast-paced startup environments
Nice to Have
● Configuration management/automation at scale (Ansible, AWX, Terraform, or similar)
● Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog)
● Container experience (Docker; Kubernetes a plus but not core to this VM-based stack)
● Experience in a startup or small-team environment where infra was built from scratch
● Security/compliance exposure (vulnerability remediation, hardening, SSO/access control)
The Mindset
Problem Solver
You thrive on complex, ambiguous challenges and engineer elegant solutions.
Ownership-Driven
You take initiative, move fast, and deliver outcomes without hand-holding.
Continuous Learner
You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.
Startup DNA
You excel in fast-moving environments where priorities evolve and impact is immediate.
We are looking for a Linux Administrator to keep our servers secure, patched and reliable.
Responsibilities
- Administer RHEL and Ubuntu servers
- Automate tasks with Shell scripts and Ansible
- Patch, monitor and troubleshoot systems
- Manage users, storage and networking
Requirements
- 2+ years of Linux administration
- Strong shell scripting skills
- RHCSA or RHCE is a plus
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
DevOps / Infrastructure Engineer
Location: Chennai
Experience: 5+ Years
Role: DevOps / Infrastructure Engineer
Job Description
We are looking for an experienced DevOps / Infrastructure Engineer with strong hands-on experience in Linux administration, containerization, Kubernetes, automation, monitoring, and troubleshooting.
Mandatory Skills
- 5+ years of experience in DevOps / Infrastructure Administration
- Strong hands-on experience with Linux Administration
- Experience with Docker and Kubernetes
- Monitoring tools: AppDynamics, Prometheus, Grafana
- Strong Shell Scripting / Python Scripting skills
- Hands-on experience with Ansible 4.1
- Strong troubleshooting and problem-solving skills
- Experience in infrastructure/application monitoring and production support
- Good understanding of DevOps practices and automation
Key Responsibilities
- Manage and support Linux-based infrastructure and production environments.
- Deploy, manage, and troubleshoot applications using Docker and Kubernetes.
- Develop and maintain automation scripts using Shell/Python.
- Automate infrastructure and configuration management using Ansible.
- Monitor applications and infrastructure using AppDynamics, Prometheus, and Grafana.
- Perform root-cause analysis and resolve infrastructure/application issues.
- Handle incidents, troubleshoot performance issues, and ensure system availability.
- Collaborate with development and operations teams to improve deployment and operational processes.
🚀 Hiring – AWS / Kubernetes / OpenShift Engineer
📍 Location: Bangalore
💼 Experience: 7–10 Years
⚡ Joining: Immediate Joiners Only
🔑 Required Skills
- Strong hands-on experience in AWS
- Expertise in Kubernetes & OpenShift
- Strong Linux Administration & Troubleshooting
- Experience in containerized environments and platform operations
- Production support, monitoring and incident troubleshooting
- Good understanding of cloud and infrastructure technologies
📌 Interview Process
2nd Round – Face-to-Face Interview in Bangalore
👉 Please share profiles of candidates who are available for a F2F interview in Bangalore.
#Hiring #AWS #Kubernetes #OpenShift #Linux #CloudEngineer #PlatformEngineer #DevOps #BangaloreJobs #ImmediateJoiners #WFO #ITJobs
Job Title : SDE 3 – Infrastructure Platform Engineer
Experience : 5.5 to 8.5 Years
Number of Positions : 2
Employment Type : C2H (Contract to Hire)
Work Mode : Remote during contractual period → 5 Days WFO after conversion
Contract Duration : 3 Months
Post-Conversion Location : Pune
Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred
(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)
Role Overview :
We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.
The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.
Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.
Key Responsibilities :
- Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
- Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
- Manage and optimize Docker and Kubernetes workloads.
- Build internal platform tools to improve developer productivity and engineering efficiency.
- Implement monitoring, logging, alerting, and observability solutions.
- Participate in incident response, RCA, postmortems, and reliability improvements.
- Design and improve CI/CD pipelines and deployment automation.
- Contribute to system design, architecture discussions, scalability, security, and cost optimization.
- Collaborate with application, data, and product engineering teams.
Required Skills :
- 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
- Strong hands-on experience with AWS, GCP, or Azure.
- Strong expertise in Terraform / CloudFormation.
- Experience with CI/CD, Docker, and Kubernetes.
- Strong programming skills in at least one of:
- Python, Go, Java, or Ruby.
- Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
- Experience working with production infrastructure and highly available systems.
- Strong troubleshooting and problem-solving skills.
Nice to Have :
- Experience with SRE practices and production on-call ownership.
- Experience in fintech, payments, banking, or transaction-heavy systems.
- Knowledge of cloud security, compliance, or FinOps/cost optimization.
- Experience building internal developer platforms or productivity tools.
- Previous product company experience.
Interview Process :
Round 1 : Take-Home Coding Assignment – Submit within 48 hours
Round 2 : Coding Assignment Discussion – 1 Hour
Round 3 : Technical Managerial Round – 30 Minutes
Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.
Ideal Candidate :
Strong Platform / Infrastructure Engineer with hands-on experience in :
Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations
Pure DevOps profiles without strong coding and platform engineering experience are not preferred.
Job Description: Lead - Cloud Engineering (AWS / Azure)
Role Title: Lead - Cloud Engineering
Experience Level: 10+ Years
Domain Focus: Healthcare AI & Cloud Infrastructure
Location: Remote
Job Overview
We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.
The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.
Key Responsibilities
Cloud Architecture & Migration
- Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
- Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
- Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.
Security & Healthcare Compliance
- Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
- Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.
Leadership & Team Management
- Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
- Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
- Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.
Operations & FinOps
- Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
- Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.
Key Requirements
- Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
- Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
- DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
- Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
- AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
- Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).
We are looking for an experienced DevOps Engineer to take ownership of production infrastructure, cloud environments, Kubernetes platforms, and infrastructure automation. This is a hands-on role for someone who enjoys solving complex infrastructure challenges and is comfortable being responsible for systems in production.
Key Responsibilities
- Own and operate production infrastructure, including participating in an on-call rotation and responding to production incidents.
- Design, operate, and continuously improve Kubernetes clusters in production.
- Manage and automate infrastructure using Infrastructure as Code, primarily with Terraform.
- Build, maintain, and optimise cloud infrastructure across AWS, GCP, or Azure.
- Work extensively with Linux, including system administration, networking, troubleshooting, and system-level configuration.
- Manage production deployment and GitOps workflows using ArgoCD.
- Improve infrastructure reliability, scalability, security, monitoring, and operational efficiency.
- Troubleshoot complex production issues and drive problems through to resolution.
- Develop automation and processes that reduce manual operational work.
Essential Requirements
- 4+ years of hands-on experience operating production infrastructure, with personal ownership and responsibility for live systems, including on-call experience.
- Deep, hands-on Kubernetes experience — you must have operated and managed Kubernetes clusters, rather than simply deploying applications onto clusters managed by another team.
- Strong experience with Infrastructure as Code, with Terraform strongly preferred. Experience with Pulumi or CloudFormation is also considered.
- Strong experience with at least one major cloud platform, ideally AWS. Strong GCP or Azure experience is also welcome, provided you are willing to work with AWS.
- Strong Linux skills and confidence working from the command line, including networking, troubleshooting, system configuration, and performance issues.
- Production experience with ArgoCD and GitOps-based deployment workflows.
- Strong troubleshooting and problem-solving skills, with the ability to take ownership of production incidents and infrastructure issues.
Nice to Have
Experience with email infrastructure would be a strong advantage, particularly:
- Exim
- IMAP / SMTP
- Postfix
- Dovecot
- General mail server administration and maintenance






