Cutshort logo
For Employers
Thinqor logo
Cloud Engineer
Cloud Engineer

Cloud Engineer at Thinqor · Bengaluru (Bangalore) · 5 - 8 years · ₹15L - ₹20L / yr · Profitable · Posted 20 Mar 2026

Thinqor's logo

Cloud Engineer

sai patel's profile picture
Posted by sai patel
5 - 8 yrs
₹15L - ₹20L / yr
Bengaluru (Bangalore)
Skills
MLOps
Windows Azure
skill iconKubernetes
aks
aro
Terraform
databricks
skill iconJenkins

 Hiring: Cloud Engineer – MLOps Platform 🚨

📍 Location: Bangalore

🧠 Experience: 5–8 Years

We are looking for an experienced Cloud Engineer to support ML teams and drive end-to-end automation for model deployment across modern cloud platforms.

🔹 Tech Stack:

Azure | Databricks | AKS | ARO | Terraform | MLflow | CI/CD

🔹 Key Responsibilities:

• Build and maintain CI/CD and Continuous Training (CT) pipelines using Azure DevOps, GitHub Actions, or Jenkins.

• Deploy Databricks jobs, MLflow models, and microservices on AKS / ARO environments.

• Automate infrastructure using Terraform and GitOps practices.

• Manage Databricks workspaces, AKS clusters, and networking configurations.

• Implement monitoring, logging, and alerting systems for ML workloads.

• Ensure cloud security, governance, and cost optimization best practices.

🔹 Required Skills:

✔ Strong hands-on experience with Azure, AKS, ARO, and Databricks

✔ Experience with MLflow and Kubernetes-based deployments

✔ Proficiency in Python and Bash / PowerShell scripting

✔ Strong understanding of cloud security, infrastructure automation, and distributed systems

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Thinqor

Founded :
2016
Type :
Products & Services
Size :
100-1000
Stage :
Profitable

About

N/A

Company social profiles

bloginstagramtwitterfacebook

Similar jobs (10)

company logo
Ganesh Ram
Posted by Ganesh Ram
Bengaluru (Bangalore), Mumbai, Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Hyderabad, Pune
7 - 10 yrs
₹15L - ₹20L / yr
CI/CD
skill iconKubernetes
helm
Terraform
yaml

Cloud Expertise(Azure):

• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.

• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.

• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).

• Understanding on policies and security aspects of cloud.


Kubernetes & Helm:

• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.

• Understand of Kubernetes templates and its deployment.

• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.

• Implement Kubernetes best practices, including security, networking, and scaling.

• Concepts of Docker and Containers


CI/CD & Programming:

• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).

• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.

• Proven skill in python programming and concepts.


Monitoring and Observability:

Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.

• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.

 

Read more
company logo
Shakthi M
Posted by Shakthi M
Bengaluru (Bangalore)
5 - 14 yrs
Best in industry
skill iconPython
Azure
Terraform
DevOps
  • Strong hands-on experience in Microsoft Azure Cloud.
  • Good understanding of Azure services such as Compute, Storage, Event Hub, Event Subscription, Storage Queue, and PaaS services.
  • Basic understanding of Azure AI Foundry and AI-related Azure service setup.
  • Good Azure networking basics: VNet, subnet, routing, and basic troubleshooting.
  • Strong knowledge of Terraform, especially:
  • Terraform state
  • plan / apply
  • troubleshooting failures
  • migration risks
  • Terraform Enterprise concepts
  • Strong Python coding capability, not just basic scripting.
  • Experience using Python for API integration, automation, JSON/YAML handling, and internal tooling.
  • Good understanding of CI/CD pipelines.
  • Ability to troubleshoot pipeline failures.
  • Comfortable with YAML and JSON.
  • Ability to troubleshoot Azure infrastructure/platform issues.
  • Ability to collect logs/evidence and coordinate with network/app/Microsoft support teams.
  • Basic awareness of agentic AI / LLM concepts.
  • Awareness of security and cost best practices.

Good to Have Skills

  • Hands-on experience with Harness.
  • Hands-on experience with Terraform Enterprise.
  • Exposure to LangGraph / LangChain.
  • Exposure to agentic AI workflows or skill creation.
  • Exposure to Claude or enterprise LLM integrations.
  • Knowledge of Azure ML Workspace, model registry, and managed endpoints.
  • MLOps / LLMOps knowledge.
  • FinOps / Azure cost optimization experience.
  • Azure certifications: AZ-104, AZ-305, AZ-400, AZ-500.

 

Screening Priority:

Azure Cloud + Terraform + Python Coding + CI/CD Troubleshooting + YAML/JSON + Basic Agentic AI Awareness

 

Read more
company logo
Mamta K
Posted by Mamta K
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Hyderabad, Bengaluru (Bangalore)
4 - 7 yrs
₹10L - ₹15L / yr
MLOps
skill iconPython
DevOps
skill iconDocker
skill iconAmazon Web Services (AWS)
+7 more

Job Title : MLOps Engineer

Mode: Hybrid

Experience : 4 to 7 Years

Location : Hyderabad (Priority)/Bengaluru locations only


Notice Period : Immediate Joiner


Job Summary:

 

We are looking for a skilled and proactive ML Engineer with strong expertise in Python, Databricks, and Machine Learning model development. The ideal candidate should be proficient in building scalable data pipelines and deploying ML models, with a working knowledge of MLOps principles and tooling. This role offers an opportunity to work on impactful AI/ML initiatives in a collaborative environment.

 

Key Responsibilities:

 

• Develop and maintain machine learning pipelines for training, testing, and deploying models

• Design and implement infrastructure for managing and monitoring machine learning models

• Work with data scientists to build scalable, efficient, and automated model training and testing processes

• Collaborate with software engineers to integrate machine learning models into production systems

• Automate and optimize the deployment and scaling of machine learning models in a distributed computing environment

• Monitor and troubleshoot machine learning systems and infrastructure to ensure high availability and performance

• Develop and maintain documentation and best practices for MLOps processes and procedures.

 

Experience:

Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field

• 3+ years of experience in MLOps or related field, including building and deploying machine learning models at scale

•Proficiency in programming languages such as Python, Java, and C++

•Experience with machine learning frameworks such as TensorFlow, PyTorch, and Keras

• Experience with containerization technologies such as Docker and Kubernetes

• Strong understanding of DevOps principles and practices

• Experience with cloud computing platforms such as AWS, Azure, or Google Cloud

Read more
company logo
Sandeep Selvan
Posted by Sandeep Selvan
Bengaluru (Bangalore)
4 - 12 yrs
Best in industry
MLOps
databricks
skill iconMachine Learning (ML)
MLFlow
LangGraph
+4 more

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.


Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.


You Will:

  • Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
  • Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
  • CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
  • Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
  • Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
  • Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
  • Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
  • Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
  • Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
  • Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
  • Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
  • Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
  • Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
  • Perform other duties as assigned


You Have:

  • Enterprise SaaS software solutions with high availability and scalability
  • Solution handling large scale structured and unstructured data from varied data sources
  • Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
  • Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
  • In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
  • AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
  • Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
  • Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
  • Programming languages like Python and SQL
  • Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
  • Solution Cost Optimisations and design to cost
  • Legally eligible to work in India on an ongoing basis

 

Get to Know Us:

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.


Equal Opportunity Employer:

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information. 

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.



Job application link : https://grnh.se/z7qx2ehx1us

Read more
company logo
Swathi S
Posted by Swathi S
Chennai
7 - 12 yrs
₹30L - ₹55L / yr
skill iconAmazon Web Services (AWS)
skill iconPython
CI/CD
DevOps
Platform as a Service (PaaS)
+7 more

Amura’s Vision 


We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.


Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.


Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.


These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.

We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence. 


Role Overview 


We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.


This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.


You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability. 


Key Responsibilities 


Cloud Infrastructure & Platform Engineering (AWS) 

  • Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
  • Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
  • Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
  • Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
  • Build reusable platform templates and shared infrastructure modules. 


AI/ML Infrastructure & MLOps 

  • Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
  • Support GPU-based workloads and optimize compute/storage usage.
  • Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
  • Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines. 


CI/CD, Automation & Developer Productivity 

  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
  • Automate deployments, environment provisioning, and release workflows.
  • Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
  • Implement automated patching, scaling, backups, cleanup workflows, and drift detection. 


Containers, Kubernetes & Platform Reliability

  • Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
  • Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
  • Optimize infrastructure for performance, resilience, and cost-efficiency.
  • Implement progressive deployment strategies including blue/green, canary, and rolling deployments. 


Observability, Incident Response & SRE Practices

  • Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
  • Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
  • Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.

FinOps, Cost Governance & Security

  • Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
  • Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
  • Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.

Collaboration, Leadership & Platform Culture

  • Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
  • Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.

Skills & Qualifications


Must-Have:

  • 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
  • Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Here are answers to some questions you may have

Where is your office?

Chennai (Velachery)

Work Model

Work from Office – because great stories are built in person!

Do you have an online presence?

https://amura.ai (we are @AmuraHealth on all social media)


Read more
MNC
MNC
Agency job
via by Jaya Mishra
Remote only
5.5 - 16 yrs
Best in industry
skill iconKubernetes
Platform engineering
Azure networking
Terraform
DevOps

Azure DevOps Engineer

Experience

5-10 years of hands-on experience in Platform Engineering, Cloud Engineering, SRE, DevOps or Infrastructure Engineering roles.

Priority 1 – Must Have (Hands-On)

Azure Cloud Platform

·        Strong hands-on experience supporting Azure workloads in production environments.

·        Experience designing, building and supporting Azure infrastructure using Terraform.

·        Good understanding of Azure networking and connectivity patterns.

·        Experience supporting:

Ø AKS

Ø Application Gateway

Ø Azure Traffic Manager

Ø Key Vault

Ø Azure Monitor / Log Analytics

Ø Managed Identities

Ø Service Principals

Ø Private Endpoints

Ø VNets, NSGs and Route Tables

 

Kubernetes / AKS

  • Strong practical experience operating and supporting AKS.
  • Ability to troubleshoot:

Ø Pod failures

Ø Ingress issues

Ø DNS issues

Ø SSL/TLS certificate issues

Ø Network routing issues

Ø Performance and availability incidents

  • Experience with:

Ø Helm

Ø Ingress Controllers

Ø Cluster upgrades

Ø Scaling

Ø Monitoring

 

Terraform

·        Strong hands-on experience writing and maintaining Terraform.

·        Experience creating reusable modules.

·        Experience managing:

·        State files

·        Remote backends

·        Environment promotion

·        Infrastructure lifecycle

Linux & Scripting

·        Strong Linux administration fundamentals.

·        Practical experience troubleshooting production issues.

·        Bash scripting mandatory.

·        Python desirable.

Application Support / Troubleshooting

Must be comfortable supporting business applications end-to-end.

 

 Priority 2 – Highly Desirable

GitHub & DevOps Platform

Hands-on experience with:

·        GitHub Enterprise

·        GitHub Actions

·        Shared workflows

·        Reusable pipelines

·        Repository onboarding

·        Branch protections

·        GitHub security features

Experience supporting:

·        Runner issues

·        Disk space issues

·        Network connectivity issues

·        Dependency failures

·        Self-hosted runners lifecycle management

API Management

Pipeline failures

GitOps

Experience with:

·        ArgoCD

·        GitOps deployment models

·        Kubernetes deployment automation

 

Monitoring & Observability

Experience working with:

·        Prometheus

·        Grafana

·        Azure Monitor

·        Log Analytics

·        Application Insights

Read more
company logo
Arpita Pathak
Posted by Arpita Pathak
Indore, Pune, Ahmedabad
4 - 6 yrs
₹7L - ₹10L / yr
skill iconPython
skill iconMachine Learning (ML)
Artificial Intelligence (AI)
Generative AI
Large Language Models (LLM) tuning
+5 more

Experience - 4 to 6 year

Location – Ahmedabad/Pune/Indore

  • Additional Job Description

Additional Job Description

Required Skills and Experience: 

  • Strong proficiency in Python and experience with ML/AI libraries (scikit-learn, TensorFlow, PyTorch, Hugging Face ecosystem).
  • Hands-on experience with LLMs, RAG, vector databases, and retrieval pipelines.
  • Practical experience deploying agentic workflows and building multi-step, tool-enabled agents.
  • Experience using Garak (or similar LLM red-teaming/vulnerability scanners) to identify model weaknesses and harden deployments.
  • Demonstrated experience implementing content filtering / moderation systems.
  • Solid skills working with structured and unstructured data and advanced feature engineering.
  • Familiarity with cloud GenAI platforms and services (Azure AI Services preferred; AWS/GCP acceptable).
  • Experience building APIs/microservices; containerization (Docker), orchestration (Kubernetes).
  • Strong understanding of model evaluation, performance profiling, inference cost optimization, and observability.
  • Good knowledge of security, data governance, and privacy best practices for AI systems.


Read more
company logo
Silfa Rodrigues
Posted by Silfa Rodrigues
HSR Bangalore
6 - 8 yrs
₹18L - ₹24L / yr
CI/CD
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Google Cloud Platform (GCP)
Terraform
+3 more

 The Role 

As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that 

keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and 

observability that let a small, fast-moving team ship confidently — and you'll keep our AI and 

data workloads reliable and affordable at scale. 

This is a hands-on role with real ownership: you won't be maintaining someone else's setup, 

you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make 

deployment boring, incidents rare, and scaling a non-event. --- 


What You'll Own 

**CI/CD & developer experience** 

- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with 

confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible 

automated testing, rollbacks, and release controls. 

**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or 

similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow. 

**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and 

resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch 

processing for the speech pipeline. 

**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, 

on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients. 

**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and 

processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware. 

**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in 

transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security 

requirements) for a product that handles sensitive customer conversations. 

**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability 

and spend. --- 


What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production 

systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and 

Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, 

Jenkins, Argo, or similar). 

- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong 

automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, 

OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team. 


Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch 

pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data 

pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud 

ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. --- 


Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real 

scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data, 

and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an

Read more
company logo
Remote only
10 - 15 yrs
₹30L - ₹35L / yr
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Generative AI
Implementation
System deployment
+9 more

Job Description: Lead - Cloud Engineering (AWS / Azure)

Role Title: Lead - Cloud Engineering

Experience Level: 10+ Years

Domain Focus: Healthcare AI & Cloud Infrastructure

Location: Remote

Job Overview

We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.

The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.

Key Responsibilities

Cloud Architecture & Migration

  • Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
  • Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
  • Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.

Security & Healthcare Compliance

  • Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
  • Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.

Leadership & Team Management

  • Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
  • Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
  • Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.

Operations & FinOps

  • Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
  • Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.

Key Requirements

  • Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
  • Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
  • DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
  • Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
  • AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
  • Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).


Read more
company logo
Banu S
Posted by Banu S
Chennai, Hyderabad
9 - 12 yrs
₹7L - ₹28L / yr
Azure Devops
Platform as a Service (PaaS)
AKS
Azure kubernetes service

Job Summary

We are looking for an experienced Azure Cloud / DevOps Engineer with strong hands-on expertise in Azure PaaS services, Azure Kubernetes Service (AKS), and Azure DevOps. The candidate will be responsible for designing, implementing, deploying, and supporting highly available and scalable cloud applications and infrastructure on Microsoft Azure.

The ideal candidate should have strong experience in CI/CD, containerization, Kubernetes, Infrastructure as Code, Azure networking, security, monitoring, and automation.

Key Responsibilities

  • Design, deploy, and manage Azure PaaS services including App Services, Azure Functions, Azure Storage, Azure SQL, Key Vault, Service Bus, Event Grid, and related services.
  • Design, configure, and administer Azure Kubernetes Service (AKS) clusters.
  • Manage Kubernetes workloads, deployments, services, ingress, namespaces, secrets, config maps, and autoscaling.
  • Implement and maintain CI/CD pipelines using Azure DevOps.
  • Develop and maintain build and release pipelines for application and infrastructure deployments.
  • Implement Infrastructure as Code (IaC) using Terraform and/or ARM/Bicep templates.
  • Build and manage Docker containers and container registries using Azure Container Registry (ACR).
  • Implement deployment strategies such as rolling, blue-green, and canary deployments where required.
  • Configure Azure networking components such as VNets, subnets, NSGs, private endpoints, load balancers, Application Gateway, and Azure DNS.
  • Implement Azure security best practices including Managed Identity, RBAC, Key Vault, secrets management, and network security.
  • Configure monitoring, logging, alerting, and troubleshooting using Azure Monitor, Log Analytics, Application Insights, and Container Insights.
  • Automate infrastructure and operational activities using PowerShell, Azure CLI, Bash, or Python.
  • Troubleshoot application, container, Kubernetes, pipeline, networking, and Azure infrastructure issues.
  • Work closely with development, architecture, security, and operations teams to deliver reliable cloud solutions.
  • Participate in production deployments, incident management, root-cause analysis, and performance optimization.

Required Skills

Azure

  • Strong hands-on experience with Microsoft Azure.
  • Azure PaaS services such as:
  • Azure App Service
  • Azure Functions
  • Azure Storage
  • Azure SQL
  • Azure Key Vault
  • Azure Service Bus
  • Azure Event Grid
  • Azure API Management
  • Good understanding of Azure networking, IAM/RBAC, security, and governance.

AKS / Kubernetes

  • Strong hands-on experience with AKS.
  • Kubernetes architecture and administration.
  • Deployments, Services, Ingress, ConfigMaps, Secrets, Namespaces.
  • Helm and Kubernetes manifests.
  • Horizontal Pod Autoscaler and cluster autoscaling.
  • Container troubleshooting and performance
Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos