Cutshort logo
For Employers
Searce Inc logo
Lead - Cloud Security and Reliability Engineer
Lead - Cloud Security and Reliability Engineer

Lead - Cloud Security and Reliability Engineer at Searce Inc · Coimbatore · 5 - 10 years · Profitable · Posted 11 Jun 2026

Searce Inc's logo

Lead - Cloud Security and Reliability Engineer

Mohammed Rabidheen's profile picture
Posted by Mohammed Rabidheen
5 - 10 yrs
Best in industry
Coimbatore
Skills
Microsoft Windows Azure
skill iconKubernetes
Terraform
Observability
Reliability engineering

About Searce

Searce (pronounced 'search') is a global, AI-native, and engineering-led modern technology consultancy. Founded in 2004 with a vision to "solve for better," we partner with organizations to "futurify" their businesses by leveraging the full power of Cloud, AI, and Data Engineering.

With a presence across 10+ countries—including the US, India, Singapore, and Australia—Searce has evolved over two decades into a trusted technology partner for over 3,000 clients. We are not just a service provider; we are a group of "solvers-at-heart" who thrive on complex technical challenges.

Why Join the "Solvers" Brigade?

  • Award-Winning Excellence: In 2026, Searce was recognized as the Google Cloud Workplace AI Transformation Partner of the Year (APAC). We are a Premier Google Cloud Partner and a top-tier Managed Services Provider (MSP).
  • AI-First Mindset: We specialize in Applied AI (Generative & Conventional), Cloud Modernization, and Location Intelligence, helping industries from FinServ and Healthcare to Retail and Manufacturing reinvent themselves.
  • The "Futurify" DNA: We don't just maintain; we improve. We use our proprietary EVLOS business innovation framework to ensure our clients aren't just moving to the cloud, but are staying ahead of the curve.

Our Culture: The HAPPIER Values

We look for individuals who live and breathe our HAPPIER values:

  • Humble: We learn from everyone.
  • Adaptable: We embrace change as the only constant.
  • Positive: We focus on solutions, not just problems.
  • Passionate: We are obsessed with engineering excellence.
  • Innovative: We challenge the status quo.
  • Excellence: We deliver impactful, futuristic outcomes.
  • Responsible: We take ownership of our work and its impact.


Your Mission: The Role

solving for better.

You are a reliability-owning, hands-on solver. Not just a "break-fix engineer."

As a DRI (directly responsible individual) for our clients' most critical systems, you’ll be the go-to expert within the squad that ensures their environments are secure, reliable, and optimized 24/7. You will deliver measurable impact – improved uptime, faster response times, and real cost savings. Not just closed tickets. Not just alerts. Real outcomes you engineer yourself.

You will lead the charge on technical execution, from complex troubleshooting and root cause analysis to engineering proactive, automated solutions. This role is about building the future of reliable cloud operations and shipping it into today's production environments.


Your Responsibilities

what you will wake up to solve.

This isn’t a “manage tickets” role. You are the architect, the executioner and the DRI for our Cloud Managed Services GTM, deploying solutions that turn operational noise into hardened outcomes. Here’s how you’ll make your mark:

  • Own Service Reliability: You will be the go-to technical expert for 24/7 cloud operations and incident management. You'll ensure strict adherence to SLOs by getting your hands dirty, leading high-stakes troubleshooting to deliver a superior client experience.
  • Engineer the Blueprint: You'll translate client needs into scalable, automated, and secure cloud architectures. You will write and maintain the operational playbooks and Infrastructure as Code (IaC) that your squad uses every day.
  • Automate with Intelligence: You'll lead the charge from the keyboard to futurify our operations. You'll embed AI-driven automation, predictive monitoring, and AIOps into core processes to eliminate toil and preempt incidents.
  • Drive FinOps & Impact: You'll own the technical execution of the FinOps framework. You will continuously analyze, configure, and optimize cloud spend for clients through hands-on engineering.
  • Be the Expert in the Room: You'll share your knowledge through internal demos, documentation, and technical deep dives, representing the deep expertise that turns operational complexity into business resilience.
  • Mentor & Elevate: You will be a technical mentor for your peers. Through code reviews and collaborative problem-solving, you'll help build a high-performing squad that lives the “Always Hardened” mindset.


Experience & Relevance

We are looking for future technology leaders, not just coders. We value raw intelligence, analytical rigor, and an obsessive passion for technology over any prior experience.

  • Cloud Operations Pedigree: 5+ years of experience in Azure cloud infrastructure, with a significant portion in cloud managed services. Hands-on experience in Kubernetes is mandatory.
  • Commercial Acumen: Proven track record of building and scaling a net-new managed services business.
  • Client-Facing Tech Acumen: 2+ years of experience in a client-facing technical role, acting as the trusted advisor for cloud operations, security, and reliability.


Functional Skills:

  • Service Delivery Mindset: A deep understanding of MSP business models, SLAs, and the importance of client satisfaction in an operational context.
  • Client Engagement: Ability to ask appropriate questions to get to the heart of an operational issue and win trust with stakeholders.
  • Cross-Functional Catalyst: Thrive in multi-disciplinary teams, bringing together operations, security, and development teams.
  • Repository builder: Creates reusable frameworks, IaC modules, and operational playbooks for scale.
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Searce Inc

Founded :
2004
Type :
Products & Services
Size :
1000-5000
Stage :
Profitable

About

What is ‘searce’


Searce means ‘a fine sieve’ & indicates ‘to refine, to analyze, to improve’. It signifies our way of working: To improve to the finest degree of excellence, ‘solving for better’ every time. Searcians are passionate improvers & solvers who love to question the status quo.

The primary purpose of all of us, at Searce, is driving intelligent, impactful & futuristic business outcomes using new-age technology. This purpose is driven passionately by HAPPIER people who aim to become better, everyday.

What we do

Searce is a modern tech consulting firm that empowers clients to futurify their businesses, leveraging Cloud, AI & Analytics.


  1. We are a category defining niche’ cloud-native technology consulting company, specializing in modernizing (improve, automate & transform) the full-scope of infra, app, process & work
  2. We partner with clients in their ‘beyond x’ journey to drive intelligent, impactful & futuristic business outcomes
  3. We are the most preferred tech partner of choice when it comes to ‘solving for better’ for the new-age tech startups & digital enterprises, leading disruption in their industries
  4. Our Service Offerings: We offer Advanced Cloud, Data & App Modernization, Cloud Consulting, Management & Improvement (DevOps, SysOps & Cloud Managed Services), Applied AI & Analytics services 
  5. As one of the top 5 niche’ full scope global partners for Google Cloud & a preferred partner for AWS, we are the most preferred ‘engineering-led’ tech company of choice when it comes to solving complex business problems.


Who we are


We are passionate improvers, solvers & futurists. Driven by our engineering excellence mindset, we care most about delivering intelligent, impactful & futuristic business outcomes. Searcians are motivated by continuous improvement & solving for better in everything we do.

At the core, a Searcian is self-driven to become better, everyday. In passionate pursuit of the finest degree of excellence we drive exceptional outcomes in everything we do.

We believe that trust is the most important value. We also believe that we need to ‘earn the trust’, everytime one engages with us. And earning trust for us is far more important than anything else. We aim to be the *most trusted* tech consulting partner for our clients.

We are HAPPIER at heart. Humble, Adaptable, Positive, Passionate, Innovative, Excellence focused, & Responsible. We live the HAPPIER Culture Code.

Being HAPPIER.



How we work


  1. Customers. Partners. Our aim is to build relationships with customers for life. And meaningfully improve the life of every customer.
  2. We do what we say. We say what we do. We are uncomfortably honest and transparent. Being genuine wins trust & makes people happier.
  3. Mistakes are encouraged. We make mistakes. Tons of those. Everyday. And we don’t mind apologizing to our juniors, peers or superiors. We are no ego-doers.
  4. Underpromise. Overdeliver. We work with a deep desire to go above and beyond in everything we do. Everytime.

So, If you are passionate about tech, future & what you read above (we really are!), apply here to experience the ‘Art of Possible’

Read more

Connect with the team

Profile picture
Hardik Parekh
Profile picture
Vishal Jarsania
Profile picture
Vamsi Krishna

Company social profiles

bloglinkedintwitterfacebook

Similar jobs (10)

company logo
Ganesh Ram
Posted by Ganesh Ram
Bengaluru (Bangalore), Mumbai, Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Hyderabad, Pune
7 - 10 yrs
₹15L - ₹20L / yr
CI/CD
skill iconKubernetes
helm
Terraform
yaml

Cloud Expertise(Azure):

• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.

• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.

• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).

• Understanding on policies and security aspects of cloud.


Kubernetes & Helm:

• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.

• Understand of Kubernetes templates and its deployment.

• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.

• Implement Kubernetes best practices, including security, networking, and scaling.

• Concepts of Docker and Containers


CI/CD & Programming:

• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).

• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.

• Proven skill in python programming and concepts.


Monitoring and Observability:

Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.

• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.

 

Read more
company logo
Mohammed Rabidheen
Posted by Mohammed Rabidheen
Coimbatore
3 - 8 yrs
Best in industry
Windows Azure
AKS
DevOps
Microsoft Windows Azure

Senior Cloud Site Reliability Engineer (CSRE) – Azure


About Searce:

Searce is an AI-native, engineering-led modern technology consultancy that empowers

clients to futurify their businesses by delivering real, intelligent business outcomes. As a

trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,

data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"

cultural mindset and our proprietary evlos problem-solving framework, we eliminate

bureaucratic fluff to build working prototypes fast and scale enterprise production

environments intelligently. We don't just fix systems; we leverage multi-cloud technologies

to transform client operations into distinct competitive advantages.

Position Overview:

We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to

architect, secure, and stabilize next-generation hybrid and multi-cloud environments.

Operating at the intersection of infrastructure design, security compliance, and production

operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,

and AWS.

Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a

massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For

the Lead path, you will couple this deep engineering toolkit with stakeholder management

and mentorship to drive an elite operational culture.


Experience & Level Expectation:

Years of Experience: 3 to 10 years of intensive, hands-on production operations

experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.

Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced

triaging of infrastructure failures, and ownership of the CI/CD and deployment

lifecycles.

Intermediate level (5-10 Years): Expected to take architectural ownership, serve as

primary Incident Commander for complex outages, design cross-cloud governance

frameworks, and act as a reliable bridge between technical teams and client

leadership.


Key Responsibilities & Role Expectations:

Multi-Cloud Platforms & Orchestration: Design, configure, and maintain

production-grade Kubernetes clusters across major platforms (AKS).

Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant

isolation.

Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable

infrastructure components using Terraform or Crossplane. Standardize automated

environment provisioning to eliminate configuration drift across multi-branch

environments.

Incident Management & Reliability (SRE): Own and optimize the production on-call

rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing

Mean Time to Recovery (MTTR) through centralized log and metric correlation.

Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to

identify core architectural vulnerabilities and establish long-term fixes preventing

recurrence.

Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,

operating system patching strategies (Linux/Windows), database lifecycle updates,

and multi-region Disaster Recovery (DR) failover drills.

Core Core Operations & Legacy Integration: Manage enterprise-level hybrid

networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP

configurations) while effectively connecting cloud native services to legacy

infrastructures like Active Directory.

Security & Governance: Embed Zero Trust policies, secure secrets management

(Secrets Manager/Key Vault), and continuous vulnerability patching into the

automated SDLC pipeline.


Required Technical Skills:

- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.

- Containers & Orchestration

  • Production-level management of GKE, AKS, and EKS.
  • Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption
Read more
company logo
Shakthi M
Posted by Shakthi M
Bengaluru (Bangalore)
5 - 14 yrs
Best in industry
skill iconPython
Azure
Terraform
DevOps
  • Strong hands-on experience in Microsoft Azure Cloud.
  • Good understanding of Azure services such as Compute, Storage, Event Hub, Event Subscription, Storage Queue, and PaaS services.
  • Basic understanding of Azure AI Foundry and AI-related Azure service setup.
  • Good Azure networking basics: VNet, subnet, routing, and basic troubleshooting.
  • Strong knowledge of Terraform, especially:
  • Terraform state
  • plan / apply
  • troubleshooting failures
  • migration risks
  • Terraform Enterprise concepts
  • Strong Python coding capability, not just basic scripting.
  • Experience using Python for API integration, automation, JSON/YAML handling, and internal tooling.
  • Good understanding of CI/CD pipelines.
  • Ability to troubleshoot pipeline failures.
  • Comfortable with YAML and JSON.
  • Ability to troubleshoot Azure infrastructure/platform issues.
  • Ability to collect logs/evidence and coordinate with network/app/Microsoft support teams.
  • Basic awareness of agentic AI / LLM concepts.
  • Awareness of security and cost best practices.

Good to Have Skills

  • Hands-on experience with Harness.
  • Hands-on experience with Terraform Enterprise.
  • Exposure to LangGraph / LangChain.
  • Exposure to agentic AI workflows or skill creation.
  • Exposure to Claude or enterprise LLM integrations.
  • Knowledge of Azure ML Workspace, model registry, and managed endpoints.
  • MLOps / LLMOps knowledge.
  • FinOps / Azure cost optimization experience.
  • Azure certifications: AZ-104, AZ-305, AZ-400, AZ-500.

 

Screening Priority:

Azure Cloud + Terraform + Python Coding + CI/CD Troubleshooting + YAML/JSON + Basic Agentic AI Awareness

 

Read more
company logo
Remote only
10 - 15 yrs
₹30L - ₹35L / yr
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Generative AI
Implementation
System deployment
+9 more

Job Description: Lead - Cloud Engineering (AWS / Azure)

Role Title: Lead - Cloud Engineering

Experience Level: 10+ Years

Domain Focus: Healthcare AI & Cloud Infrastructure

Location: Remote

Job Overview

We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.

The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.

Key Responsibilities

Cloud Architecture & Migration

  • Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
  • Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
  • Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.

Security & Healthcare Compliance

  • Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
  • Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.

Leadership & Team Management

  • Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
  • Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
  • Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.

Operations & FinOps

  • Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
  • Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.

Key Requirements

  • Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
  • Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
  • DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
  • Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
  • AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
  • Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).


Read more
company logo
Swathi S
Posted by Swathi S
Chennai
7 - 12 yrs
₹30L - ₹55L / yr
skill iconAmazon Web Services (AWS)
skill iconPython
CI/CD
DevOps
Platform as a Service (PaaS)
+7 more

Amura’s Vision 


We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.


Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.


Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.


These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.

We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence. 


Role Overview 


We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.


This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.


You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability. 


Key Responsibilities 


Cloud Infrastructure & Platform Engineering (AWS) 

  • Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
  • Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
  • Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
  • Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
  • Build reusable platform templates and shared infrastructure modules. 


AI/ML Infrastructure & MLOps 

  • Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
  • Support GPU-based workloads and optimize compute/storage usage.
  • Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
  • Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines. 


CI/CD, Automation & Developer Productivity 

  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
  • Automate deployments, environment provisioning, and release workflows.
  • Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
  • Implement automated patching, scaling, backups, cleanup workflows, and drift detection. 


Containers, Kubernetes & Platform Reliability

  • Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
  • Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
  • Optimize infrastructure for performance, resilience, and cost-efficiency.
  • Implement progressive deployment strategies including blue/green, canary, and rolling deployments. 


Observability, Incident Response & SRE Practices

  • Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
  • Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
  • Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.

FinOps, Cost Governance & Security

  • Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
  • Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
  • Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.

Collaboration, Leadership & Platform Culture

  • Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
  • Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.

Skills & Qualifications


Must-Have:

  • 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
  • Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Here are answers to some questions you may have

Where is your office?

Chennai (Velachery)

Work Model

Work from Office – because great stories are built in person!

Do you have an online presence?

https://amura.ai (we are @AmuraHealth on all social media)


Read more
MNC
MNC
Agency job
via by Gauri Naik
Remote only
10 - 18 yrs
Best in industry
Landing page optimization
skill iconKubernetes
Terraform
CI/CD

🚀 Hiring: Senior Azure Platform Engineer | Azure | Terraform | Kubernetes


📍 Location: India

💼 Employment: Full-Time | Long-Term

🎯 Experience: 10+ Years | 8+ Years Hands-on Azure


🔑 What We’re Looking For

▪️ Strong expertise in Azure Cloud Platform & Landing Zone Architecture

▪️ Advanced hands-on experience with Terraform & reusable IaC modules

▪️ Expertise in Kubernetes, Helm & container platforms

▪️ Strong understanding of Azure Networking, Entra ID, RBAC & Azure Policy

▪️ Experience with GitHub, GitHub Actions / Azure Pipelines & CI/CD

▪️ Hands-on GitOps & ArgoCD experience

▪️ Strong observability skills with Prometheus, Grafana & Azure Monitor

▪️ Experience designing and supporting microservices architectures

▪️ Knowledge of security, policy-as-code and IaC security scanning

▪️ Exposure to Azure Arc / Azure Local / Azure Stack HCI is a plus

Read more
company logo
Silfa Rodrigues
Posted by Silfa Rodrigues
HSR Bangalore
6 - 8 yrs
₹18L - ₹24L / yr
CI/CD
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Google Cloud Platform (GCP)
Terraform
+3 more

 The Role 

As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that 

keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and 

observability that let a small, fast-moving team ship confidently — and you'll keep our AI and 

data workloads reliable and affordable at scale. 

This is a hands-on role with real ownership: you won't be maintaining someone else's setup, 

you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make 

deployment boring, incidents rare, and scaling a non-event. --- 


What You'll Own 

**CI/CD & developer experience** 

- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with 

confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible 

automated testing, rollbacks, and release controls. 

**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or 

similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow. 

**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and 

resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch 

processing for the speech pipeline. 

**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, 

on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients. 

**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and 

processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware. 

**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in 

transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security 

requirements) for a product that handles sensitive customer conversations. 

**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability 

and spend. --- 


What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production 

systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and 

Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, 

Jenkins, Argo, or similar). 

- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong 

automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, 

OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team. 


Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch 

pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data 

pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud 

ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. --- 


Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real 

scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data, 

and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an

Read more
company logo
Daniel Castellanos
Posted by Daniel Castellanos
Remote only
1 - 10 yrs
₹10L - ₹50L / yr
skill iconKubernetes
ArgoCD
Exim
Postfix

We are looking for an experienced DevOps Engineer to take ownership of production infrastructure, cloud environments, Kubernetes platforms, and infrastructure automation. This is a hands-on role for someone who enjoys solving complex infrastructure challenges and is comfortable being responsible for systems in production.


Key Responsibilities

  • Own and operate production infrastructure, including participating in an on-call rotation and responding to production incidents.
  • Design, operate, and continuously improve Kubernetes clusters in production.
  • Manage and automate infrastructure using Infrastructure as Code, primarily with Terraform.
  • Build, maintain, and optimise cloud infrastructure across AWS, GCP, or Azure.
  • Work extensively with Linux, including system administration, networking, troubleshooting, and system-level configuration.
  • Manage production deployment and GitOps workflows using ArgoCD.
  • Improve infrastructure reliability, scalability, security, monitoring, and operational efficiency.
  • Troubleshoot complex production issues and drive problems through to resolution.
  • Develop automation and processes that reduce manual operational work.

Essential Requirements

  • 4+ years of hands-on experience operating production infrastructure, with personal ownership and responsibility for live systems, including on-call experience.
  • Deep, hands-on Kubernetes experience — you must have operated and managed Kubernetes clusters, rather than simply deploying applications onto clusters managed by another team.
  • Strong experience with Infrastructure as Code, with Terraform strongly preferred. Experience with Pulumi or CloudFormation is also considered.
  • Strong experience with at least one major cloud platform, ideally AWS. Strong GCP or Azure experience is also welcome, provided you are willing to work with AWS.
  • Strong Linux skills and confidence working from the command line, including networking, troubleshooting, system configuration, and performance issues.
  • Production experience with ArgoCD and GitOps-based deployment workflows.
  • Strong troubleshooting and problem-solving skills, with the ability to take ownership of production incidents and infrastructure issues.


Nice to Have

Experience with email infrastructure would be a strong advantage, particularly:

  • Exim
  • IMAP / SMTP
  • Postfix
  • Dovecot
  • General mail server administration and maintenance


Read more
MNC
MNC
Agency job
via by Sandhiya b
Bengaluru (Bangalore), Hyderabad
7 - 12 yrs
₹2L - ₹15L / yr
Windows Azure
Splunk
skill icongrafana
AppDynamics

Job Description

We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.

Responsibilities

  • Design, deploy, and manage Azure cloud infrastructure and services.
  • Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
  • Configure dashboards, alerts, health rules, and monitoring metrics.
  • Perform log analysis, troubleshooting, and root-cause analysis for production issues.
  • Implement observability solutions for applications, cloud infrastructure, and services.
  • Automate monitoring and operational activities using scripting.
  • Support incident, problem, and change management processes.
  • Collaborate with development, DevOps, and SRE teams to improve system reliability.
  • Maintain monitoring standards, documentation, and operational procedures.

Primary Skills

  • Microsoft Azure
  • Splunk
  • Grafana
  • AppDynamics
  • Cloud Monitoring & Observability
  • Application Performance Monitoring (APM)
  • Log Analysis & Troubleshooting

Secondary Skills

  • Azure Monitor / Log Analytics
  • Azure VMs, Storage, Networking
  • Linux
  • Python / PowerShell / Shell Scripting
  • CI/CD
  • Git
  • ITIL / ServiceNow
Read more
MNC
MNC
Agency job
via by aafia parveen
Bengaluru (Bangalore), Hyderabad
7 - 14 yrs
₹2L - ₹15L / yr
Axure
Windows Azure
skill icongrafana
Splunk
AppDynamics

Job Description

We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.

Responsibilities

  • Design, deploy, and manage Azure cloud infrastructure and services.
  • Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
  • Configure dashboards, alerts, health rules, and monitoring metrics.
  • Perform log analysis, troubleshooting, and root-cause analysis for production issues.
  • Implement observability solutions for applications, cloud infrastructure, and services.
  • Automate monitoring and operational activities using scripting.
  • Support incident, problem, and change management processes.
  • Collaborate with development, DevOps, and SRE teams to improve system reliability.
  • Maintain monitoring standards, documentation, and operational procedures.

Primary Skills

  • Microsoft Azure
  • Splunk
  • Grafana
  • AppDynamics
  • Cloud Monitoring & Observability
  • Application Performance Monitoring (APM)
  • Log Analysis & Troubleshooting

Secondary Skills

  • Azure Monitor / Log Analytics
  • Azure VMs, Storage, Networking
  • Linux
  • Python / PowerShell / Shell Scripting
  • CI/CD
  • Git
  • ITIL / ServiceNow


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos