Cutshort logo
For Employers
IT Industry - Night Shifts logo
Senior Cloud & ML Infrastructure Engineer
IT Industry - Night Shifts
Senior Cloud & ML Infrastructure Engineer

Senior Cloud & ML Infrastructure Engineer at IT Industry - Night Shifts Ā· Bengaluru (Bangalore), Hyderabad, Mumbai, Navi Mumbai, Pune, Mohali, Delhi Ā· 5 - 10 years Ā· ₹20L - ₹30L / yr Ā· Posted 16 Sep 2025

Hunarstreet Technologies Pvt Ltd's logo

Senior Cloud & ML Infrastructure Engineer

at IT Industry - Night Shifts

5 - 10 yrs
₹20L - ₹30L / yr
Bengaluru (Bangalore), Hyderabad, Mumbai, Navi Mumbai, Pune, Mohali, Delhi
Skills
skill iconAmazon Web Services (AWS)
IT infrastructure
skill iconMachine Learning (ML)
DevOps
Automation
skill iconPython

šŸš€ We’re Hiring: Senior Cloud & ML Infrastructure Engineer šŸš€


We’re looking for an experienced engineer to lead the design, scaling, and optimization of cloud-native ML infrastructure on AWS.

If you’re passionate about platform engineering, automation, and running ML systems at scale, this role is for you.


What you’ll do:

šŸ”¹ Architect and manage ML infrastructure with AWS (SageMaker, Step Functions, Lambda, ECR)

šŸ”¹ Build highly available, multi-region solutions for real-time & batch inference

šŸ”¹ Automate with IaC (AWS CDK, Terraform) and CI/CD pipelines

šŸ”¹ Ensure security, compliance, and cost efficiency

šŸ”¹ Collaborate across DevOps, ML, and backend teams


What we’re looking for:

āœ”ļø 6+ years AWS cloud infrastructure experience

āœ”ļø Strong ML pipeline experience (SageMaker, ECS/EKS, Docker)

āœ”ļø Proficiency in Python/Go/Bash scripting

āœ”ļø Knowledge of networking, IAM, and security best practices

āœ”ļø Experience with observability tools (CloudWatch, Prometheus, Grafana)


✨ Nice to have: Robotics/IoT background (ROS2, Greengrass, Edge Inference)


šŸ“ Location: Bengaluru, Hyderabad, Mumbai, Pune, Mohali, Delhi

5 days working, Work from Office

Night shifts: 9pm to 6am IST

šŸ‘‰ If this sounds like you (or someone you know), let’s connect!


Apply here:

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

Amura Health
at Amura Health
3 candid answers
1 video
Swathi S
Posted by Swathi S
Chennai
7 - 12 yrs
₹30L - ₹55L / yr
skill iconAmazon Web Services (AWS)
skill iconPython
CI/CD
DevOps
Platform as a Service (PaaS)
+7 more

Amura’s VisionĀ 


We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.


Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.


Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.


These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.

We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.Ā 


Role OverviewĀ 


We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.


This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.


You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.Ā 


Key ResponsibilitiesĀ 


Cloud Infrastructure & Platform Engineering (AWS)Ā 

  • Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
  • Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
  • Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
  • Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
  • Build reusable platform templates and shared infrastructure modules.Ā 


AI/ML Infrastructure & MLOpsĀ 

  • Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
  • Support GPU-based workloads and optimize compute/storage usage.
  • Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
  • Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.Ā 


CI/CD, Automation & Developer ProductivityĀ 

  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
  • Automate deployments, environment provisioning, and release workflows.
  • Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
  • Implement automated patching, scaling, backups, cleanup workflows, and drift detection.Ā 


Containers, Kubernetes & Platform Reliability

  • Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
  • Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
  • Optimize infrastructure for performance, resilience, and cost-efficiency.
  • Implement progressive deployment strategies including blue/green, canary, and rolling deployments.Ā 


Observability, Incident Response & SRE Practices

  • Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
  • Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
  • Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.

FinOps, Cost Governance & Security

  • Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
  • Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
  • Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.

Collaboration, Leadership & Platform Culture

  • Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
  • Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.

Skills & Qualifications


Must-Have:

  • 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
  • Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.Ā 


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.Ā 


Here are answers to some questions you may have

Where is your office?

Chennai (Velachery)

Work Model

Work from Office – because great stories are built in person!

Do you have an online presence?

https://amura.ai (we are @AmuraHealth on all social media)


Read more
Mango Sciences
Remote only
10 - 15 yrs
₹30L - ₹35L / yr
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Generative AI
Implementation
System deployment
+9 more

Job Description: Lead - Cloud Engineering (AWS / Azure)

Role Title: Lead - Cloud Engineering

Experience Level: 10+ Years

Domain Focus: Healthcare AI & Cloud Infrastructure

Location: Remote

Job Overview

We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.

The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.

Key Responsibilities

Cloud Architecture & Migration

  • Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
  • Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
  • Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.

Security & Healthcare Compliance

  • Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
  • Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.

Leadership & Team Management

  • Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
  • Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
  • Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.

Operations & FinOps

  • Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
  • Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.

Key Requirements

  • Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
  • Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
  • DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
  • Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
  • AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
  • Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).


Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen
Bengaluru (Bangalore)
10 - 17 yrs
₹5L - ₹22L / yr
skill iconAmazon Web Services (AWS)
AWS Lambda
AWS RDS

We are looking for a hands-on Senior AWS Cloud Engineer to lead the infrastructure build, optimization, automation, and production deployment of a Multi-Agent AI Chatbot Platform hosted on AWS. The development environment is already in place, and the successful candidate will drive the solution through testing, integrations, and production go-live.

Key Responsibilities

  • Review, validate, and optimize existing Terraform code and AWS infrastructure.
  • Establish and manage integrations with enterprise platforms such as ServiceNow, Workday, and other third-party systems.
  • Design, build, and support secure, scalable, and highly available AWS environments.
  • Implement and automate CI/CD pipelines and Infrastructure-as-Code practices.
  • Lead infrastructure testing, performance tuning, and production readiness activities.
  • Drive deployment and operationalization of the platform in the Production environment.
  • Implement cloud governance, security, monitoring, and FinOps best practices.
  • Troubleshoot and resolve complex cloud infrastructure issues.

Required Skills & Experience

  • 10+ years of IT experience with strong expertise in AWS Cloud Engineering.
  • Proven experience designing, deploying, and managing AWS production environments.
  • Strong hands-on experience with Terraform and Infrastructure-as-Code.
  • Experience with CI/CD pipeline automation and DevOps practices.
  • Expertise in AWS services including VPC, IAM, EC2, S3, Lambda, CloudWatch, and networking.
  • Experience in performance optimization, reliability, and cloud cost management (FinOps).
  • Strong scripting and automation skills.
  • Experience integrating enterprise applications through APIs and secure connectivity patterns.


Read more
MindBridge
Silfa Rodrigues
Posted by Silfa Rodrigues
HSR Bangalore
6 - 8 yrs
₹18L - ₹24L / yr
CI/CD
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Google Cloud Platform (GCP)
Terraform
+3 more

Ā The RoleĀ 

As a **DevOps Engineer** you'll own the infrastructure and delivery backbone thatĀ 

keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, andĀ 

observability that let a small, fast-moving team ship confidently — and you'll keep our AI andĀ 

data workloads reliable and affordable at scale.Ā 

This is a hands-on role with real ownership: you won't be maintaining someone else's setup,Ā 

you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to makeĀ 

deployment boring, incidents rare, and scaling a non-event. ---Ā 


What You'll OwnĀ 

**CI/CD & developer experience**Ā 

- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day withĀ 

confidence. - Make the path from commit to production simple, safe, and repeatable, with sensibleĀ 

automated testing, rollbacks, and release controls.Ā 

**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform orĀ 

similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.Ā 

**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, andĀ 

resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batchĀ 

processing for the speech pipeline.Ā 

**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting,Ā 

on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.Ā 

**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement andĀ 

processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.Ā 

**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption inĀ 

transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client securityĀ 

requirements) for a product that handles sensitive customer conversations.Ā 

**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliabilityĀ 

and spend. ---Ā 


What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running productionĀ 

systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) andĀ 

Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI,Ā 

Jenkins, Argo, or similar).Ā 

- Comfort with a scripting/automation language (Python, Go, or Bash) and a strongĀ 

automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog,Ā 

OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team.Ā 


Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batchĀ 

pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), dataĀ 

pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloudĀ 

ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. ---Ā 


Why Join - Own infrastructure that's already live with leading retail brands and growing fast — realĀ 

scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data,Ā 

and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an

Read more
Remote only
6 - 16 yrs
₹10L - ₹37L / yr
VMWare
skill iconAmazon Web Services (AWS)
Migration
Terraform
Amazon EKS
+1 more

This is a senior role in our Application & Database Modernization pillar, on a specific mission: leading a large-scale VMware-to-AWS migration and landing it in production. VMware estates are exactly where modernization programs go to stall — sprawling dependency graphs, undocumented workloads, and a hypervisor bill that grows while the migration deck gathers dust. Our customers need that estate moved, cut over, and retired, on a date in the contract.

You'll lead the design and implementation of the target infrastructure: Landing Zones built to AWS best practices, EKS-based container platforms, Infrastructure as Code across the stack with Terraform and CloudFormation, and HA, DR, security, and backup strategies that hold up in production. Aedeon absorbs the discovery, dependency-mapping, and validation grind that would otherwise consume the project's first two quarters. You own the judgment calls, architecture tradeoffs, cutover sequencing, risk decisions and you carry them through to production. You'll also mentor engineers and shape long-term infrastructure strategy with customer stakeholders. If you want your migration experience to end in retired VMware clusters rather than revised project plans, this is the role.

Ā 

What you will do?

  • Lead the design and implementation of infrastructure for a large-scale VMware-to-AWS migration — from discovery through production cutover.
  • Architect and build secure, scalable, and highly available AWS environments, including Landing Zones that follow AWS best practices.
  • Design and implement containerized application platforms on Amazon EKS; exposure to ECS is a plus.
  • Implement Infrastructure as Code using Terraform and AWS CloudFormation.
  • Define and enforce best practices across security, backups, high availability (HA), disaster recovery (DR), monitoring, and operations.
  • Enable configuration management using tools such as Ansible, Chef, or similar.
  • Build CI/CD pipelines and manage the complete build and release lifecycle for customer applications.
  • Drive automation across provisioning, deployment, and operational workflows.
  • Retire legacy infrastructure and land cloud-native architectures in its place.
  • Improve and maintain DevOps platforms: Jenkins, Git repositories, monitoring, and observability stacks.
  • Provide L3-level support for complex infrastructure and platform issues.
  • Work with stakeholders on technical strategy and long-term architecture.
  • Mentor engineers, contribute to talent evaluation, and support team development.

What are we looking for?

  • 6+ years of hands-on DevOps experience, with strong expertise in designing and managing cloud infrastructure.
  • Strong experience in VMware-to-AWS migration projects or large-scale infrastructure migrations.
  • 4+ years of Terraform and CloudFormation for Infrastructure as Code (IaC).
  • 4+ years in configuration management, systems engineering, and managing production-grade infrastructure.
  • Solid Linux and/or Windows administration background.
  • Deep understanding of AWS services, including VPC, EC2, IAM, EKS, ECS, RDS, S3, Backup, CloudWatch, etc.
  • Experience with Kubernetes (EKS), ECS, and Docker in production environments.
  • Hands-on experience designing HA, DR, security, and backup strategies.
  • Experience with Landing Zone setup, multi-account strategy, and AWS governance frameworks.
  • Proficiency in Git with a strong understanding of branching and merging strategies.
  • Experience with CI/CD pipelines, automation, and operational tooling.
  • A problem-solving mindset that's proactive and oriented toward reliability, performance, and scalability.

You will be preferred if

  • Experience across multiple cloud platforms (AWS, Azure, GCP).
  • AWS certifications such as Solutions Architect Associate/Professional or DevOps Engineer Professional.
  • Exposure to data platforms such as Amazon EMR, Redshift, Lake Formation, and SageMaker.
  • Experience designing cloud architectures at an L3/Architect level.


Read more
NeoGenCode Technologies Pvt Ltd
Remote, Pune
5 - 10 yrs
₹20L - ₹32L / yr
Infrastructure Platform Engineer
skill iconAmazon Web Services (AWS)
Terraform
AWS CloudFormation
skill iconKubernetes
+13 more

Job Title : SDE 3 – Infrastructure Platform Engineer

Experience : 5.5 to 8.5 Years

Number of Positions : 2

Employment Type : C2H (Contract to Hire)

Work Mode : Remote during contractual period → 5 Days WFO after conversion

Contract Duration : 3 Months

Post-Conversion Location : Pune

Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred

(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)


Role Overview :

We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.


The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.


Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.


Key Responsibilities :

  • Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
  • Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
  • Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
  • Manage and optimize Docker and Kubernetes workloads.
  • Build internal platform tools to improve developer productivity and engineering efficiency.
  • Implement monitoring, logging, alerting, and observability solutions.
  • Participate in incident response, RCA, postmortems, and reliability improvements.
  • Design and improve CI/CD pipelines and deployment automation.
  • Contribute to system design, architecture discussions, scalability, security, and cost optimization.
  • Collaborate with application, data, and product engineering teams.


Required Skills :

  • 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
  • Strong hands-on experience with AWS, GCP, or Azure.
  • Strong expertise in Terraform / CloudFormation.
  • Experience with CI/CD, Docker, and Kubernetes.
  • Strong programming skills in at least one of:
  • Python, Go, Java, or Ruby.
  • Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
  • Experience working with production infrastructure and highly available systems.
  • Strong troubleshooting and problem-solving skills.


Nice to Have :

  • Experience with SRE practices and production on-call ownership.
  • Experience in fintech, payments, banking, or transaction-heavy systems.
  • Knowledge of cloud security, compliance, or FinOps/cost optimization.
  • Experience building internal developer platforms or productivity tools.
  • Previous product company experience.


Interview Process :

Round 1 : Take-Home Coding Assignment – Submit within 48 hours

Round 2 : Coding Assignment Discussion – 1 Hour

Round 3 : Technical Managerial Round – 30 Minutes


Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.


Ideal Candidate :

Strong Platform / Infrastructure Engineer with hands-on experience in :

Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations


Pure DevOps profiles without strong coding and platform engineering experience are not preferred.

Read more
MNC
Bengaluru (Bangalore)
10 - 18 yrs
₹5L - ₹24L / yr
Infrastructure
Windows Azure
skill iconAmazon Web Services (AWS)

Cloud Infrastructure Engineer – BANG | 10+ Years

Location: Bangalore

Experience: 10+ Years

Job Description:

  • Design, implement, and manage cloud infrastructure across AWS/Azure/GCP environments.
  • Strong experience in cloud architecture, compute, storage, networking, and security.
  • Manage VMs, VPC/VNet, load balancers, DNS, DHCP, firewalls, and IAM.
  • Hands-on experience with Windows/Linux servers, VMware, virtualization, and infrastructure operations.
  • Automate infrastructure provisioning and configuration using Terraform, Ansible, or similar tools.
  • Monitor infrastructure performance, availability, and capacity using tools such as Grafana, Prometheus, or CloudWatch/Azure Monitor.
  • Handle incident management, troubleshooting, disaster recovery, backup, and high-availability requirements.
  • Work with cross-functional teams to support cloud migration, infrastructure upgrades, and production environments.
  • Ensure infrastructure follows security, compliance, and operational best practices.

Must-Have Skills:

Cloud Infrastructure | AWS/Azure/GCP | Networking | Linux/Windows | VMware | Terraform | Ansible | DNS/DHCP | IAM | Monitoring | Backup & DR

Read more
Smartsheet
Sandeep Selvan
Posted by Sandeep Selvan
Bengaluru (Bangalore)
4 - 12 yrs
Best in industry
MLOps
databricks
skill iconMachine Learning (ML)
MLFlow
LangGraph
+4 more

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.


Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.


You Will:

  • Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
  • Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
  • CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
  • Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
  • Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
  • Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
  • Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
  • Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
  • Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
  • Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
  • Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
  • Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
  • Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
  • Perform other duties as assigned


You Have:

  • Enterprise SaaS software solutions with high availability and scalability
  • Solution handling large scale structured and unstructured data from varied data sources
  • Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
  • Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
  • In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
  • AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
  • Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
  • Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
  • Programming languages like Python and SQL
  • Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
  • Solution Cost Optimisations and design to cost
  • Legally eligible to work in India on an ongoing basis

Ā 

Get to Know Us:

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.


Equal Opportunity Employer:

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.Ā 

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.



Job application link : https://grnh.se/z7qx2ehx1us

Read more
Remote only
10 - 15 yrs
₹28L - ₹34L / yr
GenAI,
DevOps
skill iconPython
Infrastructure architecture
Infrastructure-as-Code
+4 more


Job Description: AI Engineer – GenAI Platform Automation

Experience: 10+ Years

Location: Remote – Pan India

Employment Type: Haparz Payroll

Work Mode: Remote

Notice Period: Immediate / Short Notice Preferred


About the Role


We are looking for a senior AI Engineer – GenAI Platform Automation to lead automation initiatives across enterprise Generative AI, Data Science, Data Engineering, and Analytics platforms.

The role focuses on building scalable, secure, and self-service automation capabilities across infrastructure provisioning, CI/CD, cloud environments, AI workload deployment, governance, observability, and operational excellence. The ideal candidate will have strong hands-on experience in platform engineering, cloud automation, DevOps, Infrastructure-as-Code, Python, and enterprise GenAI ecosystems.


Key Responsibilities

  • Lead end-to-end automation initiatives for enterprise GenAI, Data Science, Data Engineering, Metadata, Data Quality, Event Streaming, and Analytics platforms.
  • Design self-service automation for infrastructure provisioning, environment onboarding, deployment, governance, monitoring, and operational workflows.
  • Build automation capabilities supporting the AI lifecycle, including experimentation, model training, deployment, inference, observability, and lifecycle management.
  • Develop scalable Infrastructure-as-Code solutions using Terraform and cloud-native automation frameworks.
  • Design and maintain enterprise CI/CD pipelines, automated testing, deployment, and release processes using modern DevOps toolchains.
  • Automate Kubernetes, containers, serverless, and distributed computing environments in collaboration with cloud and platform engineering teams.
  • Develop automation solutions for GenAI and Agentic AI applications, including MCP-enabled services, API integrations, workflow automation, and event-driven architectures.
  • Implement observability, monitoring, logging, tracing, alerting, automated remediation, and reliability engineering practices.
  • Work with architecture, security, governance, engineering, and business teams to ensure enterprise standards and compliance requirements are met.
  • Conduct technical design reviews, automation assessments, code reviews, and establish engineering best practices.
  • Provide technical leadership and mentorship to engineering teams adopting automation-first and platform engineering practices.

What We’re Looking For

  • 10+ years of hands-on experience in platform engineering, automation engineering, cloud engineering, DevOps, or distributed systems.
  • Strong experience building enterprise self-service platforms supporting AI/ML, Data Science, Data Engineering, or Advanced Analytics workloads.
  • Strong expertise in automation frameworks, CI/CD, DevOps, Infrastructure-as-Code, and software delivery lifecycle automation.
  • Hands-on experience with Terraform and cloud-native infrastructure automation.
  • Strong experience with Python for automation, orchestration, scripting, tooling, and platform engineering.
  • Experience with Bitbucket, Bamboo, Jira, Confluence, or similar enterprise DevOps toolchains.
  • Experience working with Kubernetes, containers, serverless platforms, YARN, and distributed processing environments.
  • Knowledge of Generative AI and Agentic AI architectures, MCP frameworks, APIs, workflow automation, and enterprise AI platforms.
  • Experience with event-driven architectures and technologies such as Kafka and streaming platforms.
  • Strong understanding of cloud engineering, networking, security, scalability, resilience, and cost optimization.
  • Experience implementing observability solutions covering monitoring, logging, tracing, alerting, and operational dashboards.
  • Understanding of metadata management, data lineage, data governance, and semantic-layer concepts is highly valuable.

Good to Have

  • Experience supporting enterprise GenAI platforms, AI governance, model management, and AI operationalization.
  • Experience with GitOps, DevSecOps, Platform Engineering, and Reliability Engineering practices.
  • Exposure to data governance, data quality, metadata management, and model lifecycle automation.
  • Experience creating reusable internal developer platforms and self-service engineering tools at enterprise scale.
  • Banking, AML, fraud detection, financial crime, or risk analytics domain experience is an advantage.


Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Akilandeswari Panneerselvam
Mumbai
8 - 10 yrs
₹7L - ₹15L / yr
skill iconJava
skill iconAmazon Web Services (AWS)
skill iconKubernetes
DevOps
Artificial Intelligence (AI)
+3 more

Job Title: Java AWS Kubernetes DevOps AI/ML

Experience: 8–10 Years

The candidate should have at least 1 year of experience in AI/ML and hands-on experience with the below technologies:

Java

AWS

Kubernetes

DevOps

AI/ML

MongoDB / PostgreSQL

Key-Value Caching

Vector Databases – ChromaDB / Pgvector

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos