Cutshort logo
For Employers
LodgIQ logo
Senior Devops Engineer
Senior Devops Engineer

Senior Devops Engineer at LodgIQ · Remote only · 5 - 10 years · ₹14L - ₹35L / yr · Raised funding · Remote only · Posted 28 Jul 2026

Hashone Careers's logo

Senior Devops Engineer

Agency job
5 - 10 yrs
₹14L - ₹35L / yr
Remote only
Skills
DevOps
AWS
skill iconKubernetes
Amazon CloudWatch
skill icongrafana
prometheus

About LodgIQ

Headquartered in New York, LodgIQ delivers a revolutionary B2B SaaS platform to the travel industry. By leveraging machine learning and artificial intelligence, we enable precise forecasting and optimized pricing for hotel revenue management. Backed by Highgate Ventures and Trilantic Capital Partners, LodgIQ is a well-funded, high-growth startup with a global presence.


Role Summary:

We are seeking a Senior DevOps Engineer with 5+ years of strong hands-on experience in AWS, Kubernetes, CI/CD, infrastructure as code, and cloud-native technologies. This role involves designing and implementing scalable infrastructure, improving system reliability, and driving automation across our cloud ecosystem.


Key Responsibilities:

• Architect, implement, and manage scalable, secure, and resilient cloud

infrastructure on AWS

• Lead DevOps initiatives including CI/CD pipelines, infrastructure automation, and monitoring

• Deploy and manage Kubernetes clusters and containerized microservices

• Define and implement infrastructure as code using Terraform/CloudFormation

• Monitor production and staging environments using tools like CloudWatch, Prometheus, and Grafana

• Support MongoDB and MySQL database administration and optimization

• Ensure high availability, performance tuning, and cost optimization

• Guide and mentor junior engineers, and enforce DevOps best practices

• Drive system security, compliance, and audit readiness in cloud environments

• Collaborate with engineering, product, and QA teams to streamline release processes


Required Qualifications:

• 5+ years of DevOps/Infrastructure experience in production-grade environments

• Strong expertise in AWS services: EC2, EKS, IAM, S3, RDS, Lambda, VPC, etc.

• Proven experience with Kubernetes and Docker in production

• Proficient with Terraform, CloudFormation, or similar IaC tools

• Hands-on experience with CI/CD pipelines using Jenkins, GitHub Actions, or similar

• Advanced scripting in Python, Bash, or Go

• Solid understanding of networking, firewalls, DNS, and security protocols

• Exposure to monitoring and logging stacks (e.g., ELK, Prometheus, Grafana)

• Experience with MongoDB and MySQL in cloud environments


Preferred Qualifications:

• AWS Certified DevOps Engineer or Solutions Architect

• Experience with service mesh (Istio, Linkerd), Helm, or ArgoCD

• Familiarity with Zero Downtime Deployments, Canary Releases, and Blue/Green Deployments

• Background in high-availability systems and incident response

• Prior experience in a SaaS, ML, or hospitality-tech environment


Tools and Technologies You’ll Use:

• Cloud: AWS

• Containers: Docker, Kubernetes, Helm

• CI/CD: Jenkins, GitHub Actions

• IaC: Terraform, CloudFormation

• Monitoring: Prometheus, Grafana, CloudWatch

• Databases: MongoDB, MySQL

• Scripting: Bash, Python

• Collaboration: Git, Jira, Confluence, Slack


Why Join Us?

• Competitive salary and performance bonuses.

• Remote-friendly work culture.

• Opportunity to work on cutting-edge tech in AI and ML.

• Collaborative, high-growth startup environment.

• For more information, visit http://www.lodgiq.com

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About LodgIQ

Founded :
2015
Type :
Product
Size :
20-100
Stage :
Raised funding

About

LodgIQ is re-imagining revenue management with predictive and prescriptive analytics methods using state of the art BigData Analytics and AI / Machine Learning algorithms to forecast demand and price hotel rooms.

Headquartered in New York, LodgIQ delivers a revolutionary B2B SaaS platform to the travel industry by incorporating machine learning and artificial intelligence to advance precise forecasting and optimized pricing. For more information, visit http://www.lodgiq.com" target="_blank">http://www.lodgiq.com

Backed by Highgate Ventures and Trilantic Capital Partners, LodgIQ is a well-funded young company, seeking a motivated and entrepreneurial Data Engineer to join its Data Science team in Silicon Valley. Qualified A+ candidates will be offered an excellent compensation and benefit package.
Read more

Company video

LodgIQ's video section
LodgIQ's video section

Connect with the team

Profile picture
Sougata Chatterjee

Company social profiles

linkedin

Similar jobs (10)

company logo
Sakshi Mittal
Posted by Sakshi Mittal
Bengaluru (Bangalore)
3 - 5 yrs
₹6L - ₹12L / yr
skill iconAmazon Web Services (AWS)
DevOps
skill iconKubernetes
Terraform
CI/CD
+1 more

Job Summary :

We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.

Key Responsibilities :

- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.

- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.

- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management

- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.

- Collaborate with developers to ensure smooth and frequent deployments to production.

- Manage versioning and rollback strategies for critical deployments.

- Containerization & Orchestration using Kubernetes.

- Containerize applications using Docker, and manage them using Kubernetes.

- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.

- Develop utilities and tools to enhance operational efficiency and reliability.

- Monitoring & Incident Management

- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.

- Optimize application and system performance through proactive monitoring and configuration tuning.

Desired Skills and Experience :

- Experience Required - 6+ yrs.

- Hands-on experience on cloud services like AWS, EKS etc.

- Ability to design a good cloud solution.

- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.

- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.

- Use knowledge and research to constantly modernize our applications and infrastructure stacks.

- Be a team player and strong problem-solver to work with a diverse team.

- Having good communication skills.

Read more
company logo
Megha Shetty
Posted by Megha Shetty
Bengaluru (Bangalore)
6 - 8 yrs
₹10L - ₹20L / yr
DevOps
skill iconAmazon Web Services (AWS)
skill iconPython
Terraform
skill iconDocker

Job Description:

Pre-requisite skills required for a DevOps Engineer role include:


  • 6+yrs exp in DevOps
  • Experience working on Linux based infrastructure
  • knowledge in AWS, docker, CI/CD tools
  • Hands on exp in Python/shell scripting language
  • hands on exp in AWS and Azure
  • Work exp in Docker, Terraform, Ansible, Kubernetes, LINUX
  • Excellent understanding of Ruby, Python, Perl, and Java
  • Configuration and managing databases such as MySQL, Mongo
  • Excellent troubleshooting
  • Working knowledge of various tools, open-source technologies, and cloud services
Read more
company logo
Pradeep Chandkiran
Posted by Pradeep Chandkiran
Bengaluru (Bangalore)
3 - 5 yrs
₹8L - ₹14L / yr
skill iconKubernetes
Terraform
skill iconAmazon Web Services (AWS)

Location: Bangalore preferred / Hybrid as applicable

Experience: 3+ years

Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline

Salary: Above market standards, flexible for the right candidate

Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations


About FrontM

FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.

The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.


Role Summary

As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.

This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.


Key Responsibilities

Cloud Infrastructure & DevOps Architecture (≈45%)

· Own, maintain and improve AWS cloud infrastructure for FrontM platforms

· Create and maintain Terraform scripts for infrastructure deployment and management

· Manage Kubernetes workloads deployed within AWS EKS

· Support multi-zone AWS infrastructure design for availability, resilience and scale

· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap

CI/CD, Operations & Platform Reliability (≈35%)

· Build, maintain and improve CI/CD pipelines for backend and platform services

· Oversee technical operations with hands-on administration, monitoring and release support

· Ensure continuous server uptime, stability, performance and maintainability

· Debug, respond to and restore system outages in production and staging environments

· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io

· Support backend stability, scale and performance across Node.js, Java and related services

Security, Networking & Production Support (≈20%)

· Maintain AWS security configurations, access controls and monitoring practices

· Support complex networking requirements across multi-domain SaaS implementations

· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users

· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements

· Document operational procedures, incident findings and technical support steps clearly


Required Technical Skills

Cloud Infrastructure & AWS

· Strong hands-on experience with AWS infrastructure and cloud operations

· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Experience with AWS security setup, monitoring and multi-zone infrastructure

· Ability to manage infrastructure using Terraform

Kubernetes, CI/CD & Observability

· Strong experience with Kubernetes, preferably AWS EKS

· Extensive CI/CD and DevOps experience

· Experience with infrastructure observability and application monitoring tools

· Ability to diagnose production bottlenecks, server failures and performance issues

Backend, Networking & SaaS Operations

· Experience supporting Node.js, Java and backend system procedures for stability and scale

· Good understanding of APIs, integrations and backend service dependencies

· Experience with complex networking and multi-domain SaaS implementations

· Ability to troubleshoot technical issues with non-technical end users

Nice to Have

· Experience with MongoDB clusters in MongoDB Atlas

Personal Attributes

· Strong ownership mindset for uptime, reliability and production stability

· Practical problem-solving approach with the ability to act quickly during incidents

· Clear written and spoken communication in English

· Ability to work independently and coordinate with senior management when required

· Comfortable working in fast-moving engineering teams

· Attention to detail in security, monitoring, documentation and operational processes


Why join FrontM?

Long-Term Career Growth

Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.

Engineering Challenges That Matter

Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.

Broad Technical Ownership

Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.


Apply now

Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.

Read more
company logo
Ashish Singh
Posted by Ashish Singh
Remote only
0 - 1 yrs
₹12000 - ₹18000 / mo
CI/CD
DevOps

Build production-grade cloud infrastructure that powers enterprise applications with cutting-edge DevOps practices.


What you'll do:

  • Design CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI)
  • Containerize apps with Docker, deploy on Kubernetes clusters
  • Manage infrastructure as code (Terraform, CloudFormation)
  • Set up monitoring (Prometheus, Grafana, ELK Stack)
  • Cloud migrations (AWS EC2, EKS, RDS → GCP equivalent)
  • Optimize costs and performance for live production systems

What we need:

  • Basic Python/Bash scripting
  • Docker basics, Git workflows
  • Cloud exposure (AWS/GCP/Azure free tier projects)
  • Problem-solving mindset, eagerness to learn

Real impact:

  • Deploy apps used by 1000+ daily users
  • Work with senior DevOps engineers on client deliverables
  • Build portfolio for FAANG-level interviews


Read more
MNC
MNC
Agency job
via by aafia parveen
Bengaluru (Bangalore)
7 - 11 yrs
₹15L - ₹18L / yr
skill iconAmazon Web Services (AWS)
skill iconKubernetes
Linux/Unix
openshift

AWS / Kubernetes / OpenShift / Linux –

Location: Bangalore

Experience: 7–10 Years

Mandatory Skills:

  • Strong hands-on experience with AWS
  • Experience in Kubernetes & OpenShift
  • Strong knowledge of Linux administration
  • Experience with Docker & containerization
  • Knowledge of CI/CD pipelines and DevOps practices
  • Troubleshooting, monitoring, and deployment experience

Role: Cloud/DevOps Engineer – AWS, Kubernetes & OpenShift

Read more
company logo
Silfa Rodrigues
Posted by Silfa Rodrigues
HSR Bangalore
6 - 8 yrs
₹18L - ₹24L / yr
CI/CD
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Google Cloud Platform (GCP)
Terraform
+3 more

 The Role 

As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that 

keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and 

observability that let a small, fast-moving team ship confidently — and you'll keep our AI and 

data workloads reliable and affordable at scale. 

This is a hands-on role with real ownership: you won't be maintaining someone else's setup, 

you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make 

deployment boring, incidents rare, and scaling a non-event. --- 


What You'll Own 

**CI/CD & developer experience** 

- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with 

confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible 

automated testing, rollbacks, and release controls. 

**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or 

similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow. 

**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and 

resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch 

processing for the speech pipeline. 

**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, 

on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients. 

**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and 

processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware. 

**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in 

transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security 

requirements) for a product that handles sensitive customer conversations. 

**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability 

and spend. --- 


What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production 

systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and 

Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, 

Jenkins, Argo, or similar). 

- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong 

automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, 

OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team. 


Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch 

pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data 

pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud 

ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. --- 


Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real 

scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data, 

and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an

Read more
Gurugram
4 - 10 yrs
₹4L - ₹10L / yr
DevOps
Site Reliability Engineer (SRE)
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
+14 more

🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience Level : 4+ Years

Location : Gurugram Sector 48, Haryana (On-site)

Employment Type : Full Time Opportunity


About the Role :

We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.

In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).


Mandatory Skills :

AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux


Key Responsibilities :

1. Cloud Infrastructure & Infrastructure as Code (IaC) :

  • Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
  • Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
  • Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.

2. CI/CD, Automation & Development :

  • Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
  • Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.

3. Containerization & Orchestration :

  • Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
  • Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.

4. Observability, SRE & Incident Management :

  • Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
  • Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
  • Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
  • Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).


Required Qualifications & Skills :

  • Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
  • Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
  • Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
  • CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
  • Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
  • Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
  • Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
  • Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
  • Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Read more
company logo
Swathi S
Posted by Swathi S
Chennai
7 - 12 yrs
₹30L - ₹55L / yr
skill iconAmazon Web Services (AWS)
skill iconPython
CI/CD
DevOps
Platform as a Service (PaaS)
+7 more

Amura’s Vision 


We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.


Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.


Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.


These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.

We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence. 


Role Overview 


We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.


This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.


You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability. 


Key Responsibilities 


Cloud Infrastructure & Platform Engineering (AWS) 

  • Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
  • Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
  • Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
  • Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
  • Build reusable platform templates and shared infrastructure modules. 


AI/ML Infrastructure & MLOps 

  • Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
  • Support GPU-based workloads and optimize compute/storage usage.
  • Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
  • Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines. 


CI/CD, Automation & Developer Productivity 

  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
  • Automate deployments, environment provisioning, and release workflows.
  • Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
  • Implement automated patching, scaling, backups, cleanup workflows, and drift detection. 


Containers, Kubernetes & Platform Reliability

  • Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
  • Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
  • Optimize infrastructure for performance, resilience, and cost-efficiency.
  • Implement progressive deployment strategies including blue/green, canary, and rolling deployments. 


Observability, Incident Response & SRE Practices

  • Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
  • Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
  • Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.

FinOps, Cost Governance & Security

  • Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
  • Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
  • Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.

Collaboration, Leadership & Platform Culture

  • Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
  • Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.

Skills & Qualifications


Must-Have:

  • 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
  • Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Here are answers to some questions you may have

Where is your office?

Chennai (Velachery)

Work Model

Work from Office – because great stories are built in person!

Do you have an online presence?

https://amura.ai (we are @AmuraHealth on all social media)


Read more
company logo
Aditya Singh
Posted by Aditya Singh
Kolkata
8 - 20 yrs
₹15L - ₹25L / yr
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
ArgoCD
Amazon EC2
+6 more

We are seeking a skilled and experienced Senior DevOps Engineer to join our team. As a Senior DevOps Engineer, you will play a crucial role in designing, implementing, and managing our cloud infrastructure on AWS, while utilizing Docker, Kubernetes, Terraform and Argo CD for containerization and deployment. You will collaborate closely with development teams to ensure smooth software releases, scalability, and reliability of our systems. Also, you will be working on Airflow and Spark configurations for Data pipelines.

Responsibilities:

  • Design, implement, and maintain scalable and secure cloud infrastructure on AWS, utilizing services such as EC2, S3, VPC, IAM, and Lambda.
  • Build, deploy, and manage containerized applications using Docker and Kubernetes.
  • Develop and maintain CI/CD pipelines to enable continuous integration, automated testing, and deployment using tools like Jenkins, GitLab CI/CD, or AWS CodePipeline.
  • Implement and manage infrastructure as code (IaC) using tools like Terraform or CloudFormation to automate the provisioning and configuration of AWS resources.
  • Monitor and troubleshoot infrastructure and application issues, ensuring high availability, performance, and scalability of the system.
  • Collaborate with development teams to ensure smooth and efficient software releases, including version control, branching strategies, and release management.
  • Implement and maintain robust security practices, including access control, network security, and data encryption.
  • Automate manual tasks and processes using scripting languages (e.g., Python, Bash) and configuration management tools (e.g., Ansible).
  • Collaborate with cross-functional teams to gather requirements, provide technical guidance, and implement best practices for infrastructure and deployment processes.
  • Stay up-to-date with the latest DevOps tools, technologies, and best practices, and identify opportunities for improvement and optimization in our infrastructure and processes.

Requirements:

  • Bachelor’s or master’s degree in computer science, Software Engineering, or a related field.
  • Proven experience as a DevOps Engineer, with a focus on AWS, Docker, Kubernetes, and Argo CD.
  • Strong experience in designing, implementing, and managing cloud infrastructure on AWS, including services like EC2, S3, VPC, IAM, and Lambda.
  • Hands-on experience with containerization technologies, specifically Docker and Kubernetes.
  • Proficiency in CI/CD pipeline setup and management using tools like Jenkins, GitLab CI/CD, or AWS CodePipeline.
  • Solid understanding of infrastructure as code (IaC) principles and experience with tools like Terraform or CloudFormation.
  • Experience in monitoring and troubleshooting complex infrastructure and application issues.
  • Strong scripting and automation skills using languages such as Python, Bash, or PowerShell.
  • Excellent problem-solving and analytical skills, with the ability to handle and prioritize multiple tasks in a fast-paced environment.
  • Strong communication and collaboration skills, with the ability to work effectively with cross-functional teams.

Preferred Qualifications:

  • AWS certifications, such as AWS Certified DevOps Engineer or AWS Certified Solutions Architect.
  • Experience with infrastructure monitoring and logging tools like CloudWatch, ELK stack, or Prometheus/Grafana.
  • Knowledge of serverless architectures and experience with AWS Lambda.
  • Familiarity with other cloud platforms (e.g., Azure, Google Cloud Platform).
  • Understanding of security best practices and experience implementing security measures in cloud infrastructure.


Read more
company logo
Daniel Castellanos
Posted by Daniel Castellanos
Remote only
1 - 10 yrs
₹10L - ₹50L / yr
skill iconKubernetes
ArgoCD
Exim
Postfix

We are looking for an experienced DevOps Engineer to take ownership of production infrastructure, cloud environments, Kubernetes platforms, and infrastructure automation. This is a hands-on role for someone who enjoys solving complex infrastructure challenges and is comfortable being responsible for systems in production.


Key Responsibilities

  • Own and operate production infrastructure, including participating in an on-call rotation and responding to production incidents.
  • Design, operate, and continuously improve Kubernetes clusters in production.
  • Manage and automate infrastructure using Infrastructure as Code, primarily with Terraform.
  • Build, maintain, and optimise cloud infrastructure across AWS, GCP, or Azure.
  • Work extensively with Linux, including system administration, networking, troubleshooting, and system-level configuration.
  • Manage production deployment and GitOps workflows using ArgoCD.
  • Improve infrastructure reliability, scalability, security, monitoring, and operational efficiency.
  • Troubleshoot complex production issues and drive problems through to resolution.
  • Develop automation and processes that reduce manual operational work.

Essential Requirements

  • 4+ years of hands-on experience operating production infrastructure, with personal ownership and responsibility for live systems, including on-call experience.
  • Deep, hands-on Kubernetes experience — you must have operated and managed Kubernetes clusters, rather than simply deploying applications onto clusters managed by another team.
  • Strong experience with Infrastructure as Code, with Terraform strongly preferred. Experience with Pulumi or CloudFormation is also considered.
  • Strong experience with at least one major cloud platform, ideally AWS. Strong GCP or Azure experience is also welcome, provided you are willing to work with AWS.
  • Strong Linux skills and confidence working from the command line, including networking, troubleshooting, system configuration, and performance issues.
  • Production experience with ArgoCD and GitOps-based deployment workflows.
  • Strong troubleshooting and problem-solving skills, with the ability to take ownership of production incidents and infrastructure issues.


Nice to Have

Experience with email infrastructure would be a strong advantage, particularly:

  • Exim
  • IMAP / SMTP
  • Postfix
  • Dovecot
  • General mail server administration and maintenance


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos