Cutshort logo
For Employers
Adia Health logo
DevOps Engineer/ Site Reliability Engineer
DevOps Engineer/ Site Reliability Engineer

DevOps Engineer/ Site Reliability Engineer at Adia Health · Bengaluru (Bangalore) · 5 - 12 years · Bootstrapped · Posted 18 Feb 2025

Adia Health's logo

DevOps Engineer/ Site Reliability Engineer

Kavita Singh's profile picture
Posted by Kavita Singh
5 - 12 yrs
Best in industry
Bengaluru (Bangalore)
Skills
skill iconDocker
skill iconKubernetes
DevOps
skill iconAmazon Web Services (AWS)
Google Cloud Platform (GCP)
CI/CD
skill iconGitHub
HIPAA
SOC2
Computer Networking
skill iconPython
Scripting language
Automation
argocd
Terraform

Company Overview 

Adia Health revolutionizes clinical decision support by enhancing diagnostic accuracy and personalizing care. It modernizes the diagnostic process by automating optimal lab test selection and interpretation, utilizing a combination of expert medical insights, real-world data, and artificial intelligence. This approach not only streamlines the diagnostic journey but also ensures precise, individualized patient care by integrating comprehensive medical histories and collective platform knowledge. 

  

Position Overview 

We are seeking a talented and experienced Site Reliability Engineer/DevOps Engineer to join our dynamic team. The ideal candidate will be responsible for ensuring the reliability, scalability, and performance of our infrastructure and applications. You will collaborate closely with development, operations, and product teams to automate processes, implement best practices, and improve system reliability. 

  

Key Responsibilities 

  • Design, implement, and maintain highly available and scalable infrastructure solutions using modern DevOps practices. 
  • Automate deployment, monitoring, and maintenance processes to streamline operations and increase efficiency. 
  • Monitor system performance and troubleshoot issues, ensuring timely resolution to minimize downtime and impact on users. 
  • Implement and manage CI/CD pipelines to automate software delivery and ensure code quality. 
  • Manage and configure cloud-based infrastructure services to optimize performance and cost. 
  • Collaborate with development teams to design and implement scalable, reliable, and secure applications. 
  • Implement and maintain monitoring, logging, and alerting solutions to proactively identify and address potential issues. 
  • Conduct periodic security assessments and implement appropriate measures to ensure the integrity and security of systems and data. 
  • Continuously evaluate and implement new tools and technologies to improve efficiency, reliability, and scalability. 
  • Participate in on-call rotation and respond to incidents promptly to ensure system uptime and availability. 

  

Qualifications 

  • Bachelor's degree in Computer Science, Engineering, or related field 
  • Proven experience (5+ years) as a Site Reliability Engineer, DevOps Engineer, or similar role 
  • Strong understanding of cloud computing principles and experience with AWS 
  • Experience of building and supporting complex CI/CD pipelines using Github 
  • Experience of building and supporting infrastructure as a code using Terraform 
  • Proficiency in scripting and automating tools 
  • Solid understanding of networking concepts and protocols 
  • Understanding of security best practices and experience implementing security controls in cloud environments 
  • Knowing modern security requirements like SOC2, HIPAA, HITRUST will be a solid advantage. 
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Adia Health

Founded :
2024
Type :
Services
Size :
20-100
Stage :
Bootstrapped

About

Adia is transforming healthcare by combining proprietary clinical AI with instant provider payments, solving both the diagnostic accuracy crisis and the financial barriers that plague medical providers. In an era where traditional clinical decision support systems add complexity to clinicians' workflows, our platform orchestrates and automates the entire diagnostic journey while ensuring providers are paid the moment care is delivered.


Adia's breakthrough lies in seamlessly integrating AI decision automation with instant payment solutions. This combination transforms the diagnostic process from a complex maze into a streamlined experience for both providers and patients. Clinicians can finally practice medicine without administrative constraints, improving diagnostic accuracy while boosting provider profitability.


Our platform increases provider revenue through AI-powered diagnostic recommendations. Adia's clinical intelligence identifies over 40% more medically appropriate tests than traditional practice patterns by analyzing patient data and applying clinically validated protocols. This proactive approach ensures comprehensive patient care while substantially increasing revenue per visit. Combined with instant payments, providers realize immediate financial benefits while delivering superior patient outcomes.

Read more

Company social profiles

N/A

Similar jobs (10)

It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
company logo
Mohammed Rabidheen
Posted by Mohammed Rabidheen
Coimbatore
3 - 8 yrs
Best in industry
Windows Azure
AKS
DevOps
Microsoft Windows Azure

Senior Cloud Site Reliability Engineer (CSRE) – Azure


About Searce:

Searce is an AI-native, engineering-led modern technology consultancy that empowers

clients to futurify their businesses by delivering real, intelligent business outcomes. As a

trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,

data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"

cultural mindset and our proprietary evlos problem-solving framework, we eliminate

bureaucratic fluff to build working prototypes fast and scale enterprise production

environments intelligently. We don't just fix systems; we leverage multi-cloud technologies

to transform client operations into distinct competitive advantages.

Position Overview:

We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to

architect, secure, and stabilize next-generation hybrid and multi-cloud environments.

Operating at the intersection of infrastructure design, security compliance, and production

operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,

and AWS.

Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a

massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For

the Lead path, you will couple this deep engineering toolkit with stakeholder management

and mentorship to drive an elite operational culture.


Experience & Level Expectation:

Years of Experience: 3 to 10 years of intensive, hands-on production operations

experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.

Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced

triaging of infrastructure failures, and ownership of the CI/CD and deployment

lifecycles.

Intermediate level (5-10 Years): Expected to take architectural ownership, serve as

primary Incident Commander for complex outages, design cross-cloud governance

frameworks, and act as a reliable bridge between technical teams and client

leadership.


Key Responsibilities & Role Expectations:

Multi-Cloud Platforms & Orchestration: Design, configure, and maintain

production-grade Kubernetes clusters across major platforms (AKS).

Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant

isolation.

Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable

infrastructure components using Terraform or Crossplane. Standardize automated

environment provisioning to eliminate configuration drift across multi-branch

environments.

Incident Management & Reliability (SRE): Own and optimize the production on-call

rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing

Mean Time to Recovery (MTTR) through centralized log and metric correlation.

Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to

identify core architectural vulnerabilities and establish long-term fixes preventing

recurrence.

Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,

operating system patching strategies (Linux/Windows), database lifecycle updates,

and multi-region Disaster Recovery (DR) failover drills.

Core Core Operations & Legacy Integration: Manage enterprise-level hybrid

networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP

configurations) while effectively connecting cloud native services to legacy

infrastructures like Active Directory.

Security & Governance: Embed Zero Trust policies, secure secrets management

(Secrets Manager/Key Vault), and continuous vulnerability patching into the

automated SDLC pipeline.


Required Technical Skills:

- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.

- Containers & Orchestration

  • Production-level management of GKE, AKS, and EKS.
  • Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption
Read more
company logo
Sakshi Mittal
Posted by Sakshi Mittal
Bengaluru (Bangalore)
3 - 5 yrs
₹6L - ₹12L / yr
skill iconAmazon Web Services (AWS)
DevOps
skill iconKubernetes
Terraform
CI/CD
+1 more

Job Summary :

We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.

Key Responsibilities :

- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.

- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.

- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management

- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.

- Collaborate with developers to ensure smooth and frequent deployments to production.

- Manage versioning and rollback strategies for critical deployments.

- Containerization & Orchestration using Kubernetes.

- Containerize applications using Docker, and manage them using Kubernetes.

- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.

- Develop utilities and tools to enhance operational efficiency and reliability.

- Monitoring & Incident Management

- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.

- Optimize application and system performance through proactive monitoring and configuration tuning.

Desired Skills and Experience :

- Experience Required - 6+ yrs.

- Hands-on experience on cloud services like AWS, EKS etc.

- Ability to design a good cloud solution.

- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.

- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.

- Use knowledge and research to constantly modernize our applications and infrastructure stacks.

- Be a team player and strong problem-solver to work with a diverse team.

- Having good communication skills.

Read more
company logo
Megha Shetty
Posted by Megha Shetty
Bengaluru (Bangalore)
6 - 8 yrs
₹10L - ₹20L / yr
DevOps
skill iconAmazon Web Services (AWS)
skill iconPython
Terraform
skill iconDocker

Job Description:

Pre-requisite skills required for a DevOps Engineer role include:


  • 6+yrs exp in DevOps
  • Experience working on Linux based infrastructure
  • knowledge in AWS, docker, CI/CD tools
  • Hands on exp in Python/shell scripting language
  • hands on exp in AWS and Azure
  • Work exp in Docker, Terraform, Ansible, Kubernetes, LINUX
  • Excellent understanding of Ruby, Python, Perl, and Java
  • Configuration and managing databases such as MySQL, Mongo
  • Excellent troubleshooting
  • Working knowledge of various tools, open-source technologies, and cloud services
Read more
Gurugram
5 - 10 yrs
₹12L - ₹18L / yr
DevOps
Reliability engineering
skill iconAmazon Web Services (AWS)
Terraform
Ansible
+18 more

Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience : 5+ Years

Location : Gurugram, Haryana

Work Mode : On-site (Full-time)


About the Role :

We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.


Mandatory Skills :

AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux


Key Responsibilities :

  • Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
  • Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
  • Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
  • Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
  • Automate operational tasks using Python, Bash, and Groovy scripting.
  • Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.


Required Skills & Qualifications :

  • Bachelor's degree in Computer Science, IT, Electronics, or a related field.
  • 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
  • Strong expertise in AWS, with exposure to Azure/GCP.
  • Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
  • Strong scripting skills in Python and Bash.
  • Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Good understanding of Linux, networking, SQL, and cloud security best practices.


Preferred Skills :

  • Experience with multi-cloud environments and DevSecOps practices.
  • Knowledge of disaster recovery, automation, and microservices architecture.
  • Strong troubleshooting, communication, and problem-solving skills.
Read more
Remote only
8 - 12 yrs
Best in industry
Terraform
Artificial Intelligence (AI)
IAC
skill iconAmazon Web Services (AWS)
ECS
+6 more


Senior Platform & Site Reliability Engineer

Location: Remote Employment Type: Contract

The Role

This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.

Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.

This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.

The Scale You Will Operate At

The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.

You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.

What You Will Own

Platform Architecture

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
  • Define, build, and enforce platform standards across portfolio products
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
  • Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
  • Design and maintain batch processing infrastructure alongside live event flows
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products
  • SLOs and error budgets defined per product; reliability tracked consistently
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation
  • Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
  • Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
  • Target: full platform integration within a defined window per acquisition

A Note on Automation

Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.

Platform Stack

Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
  • Strong Terraform depth across multi-environment, multi-account setups
  • CI/CD ownership across a multi-product environment with GitHub Actions
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
  • Hands-on Grafana, Prometheus, and Loki in production
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
  • Acquisition or greenfield platform integration experience strongly preferred

How You Work

  • Comfortable operating across multiple products simultaneously — context-switching without dropping standards
  • Cost-efficiency instinct — you optimise spend as a habit, not as a project
  • You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
  • You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use

Why This Role

The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.

This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.

Read more
Gurugram
4 - 10 yrs
₹4L - ₹10L / yr
DevOps
Site Reliability Engineer (SRE)
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
+14 more

🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience Level : 4+ Years

Location : Gurugram Sector 48, Haryana (On-site)

Employment Type : Full Time Opportunity


About the Role :

We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.

In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).


Mandatory Skills :

AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux


Key Responsibilities :

1. Cloud Infrastructure & Infrastructure as Code (IaC) :

  • Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
  • Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
  • Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.

2. CI/CD, Automation & Development :

  • Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
  • Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.

3. Containerization & Orchestration :

  • Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
  • Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.

4. Observability, SRE & Incident Management :

  • Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
  • Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
  • Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
  • Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).


Required Qualifications & Skills :

  • Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
  • Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
  • Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
  • CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
  • Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
  • Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
  • Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
  • Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
  • Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Read more
company logo
Silfa Rodrigues
Posted by Silfa Rodrigues
HSR Bangalore
6 - 8 yrs
₹18L - ₹24L / yr
CI/CD
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Google Cloud Platform (GCP)
Terraform
+3 more

 The Role 

As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that 

keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and 

observability that let a small, fast-moving team ship confidently — and you'll keep our AI and 

data workloads reliable and affordable at scale. 

This is a hands-on role with real ownership: you won't be maintaining someone else's setup, 

you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make 

deployment boring, incidents rare, and scaling a non-event. --- 


What You'll Own 

**CI/CD & developer experience** 

- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with 

confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible 

automated testing, rollbacks, and release controls. 

**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or 

similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow. 

**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and 

resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch 

processing for the speech pipeline. 

**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, 

on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients. 

**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and 

processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware. 

**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in 

transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security 

requirements) for a product that handles sensitive customer conversations. 

**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability 

and spend. --- 


What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production 

systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and 

Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, 

Jenkins, Argo, or similar). 

- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong 

automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, 

OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team. 


Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch 

pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data 

pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud 

ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. --- 


Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real 

scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data, 

and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an

Read more
company logo
Bhattacharjee Akash
Posted by Bhattacharjee Akash
Bengaluru (Bangalore), Chennai, Mumbai, Hyderabad, Pune, Gurugram
3 - 10 yrs
₹12L - ₹35L / yr
Linux/Unix
skill iconKubernetes
Monitoring
skill iconDocker
skill iconAmazon Web Services (AWS)
+4 more



We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.



Key Responsibilities

  • Own reliability, availability, and performance of production services, including on-call rotation
  • Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
  • Automate deployments, scaling, and operational tasks to reduce manual toil
  • Manage containerized workloads on Kubernetes and cloud infrastructure
  • Design and maintain CI/CD pipelines for safe, frequent releases
  • Lead incident response and drive blameless postmortems with clear follow-ups
  • Perform capacity planning, performance tuning, and cost optimization
  • Define and track SLIs/SLOs and error budgets with product teams


Requirements

  • 3+ years in SRE, DevOps, or production-focused engineering
  • Strong Linux administration and hands-on Kubernetes experience
  • Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
  • Cloud experience with AWS, GCP, or Azure
  • CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
  • Proficient scripting in Python and/or Bash


Nice to have

  • Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
  • Background in high-traffic or distributed systems
Read more
company logo
Pradeep Chandkiran
Posted by Pradeep Chandkiran
Bengaluru (Bangalore)
3 - 5 yrs
₹8L - ₹14L / yr
skill iconKubernetes
Terraform
skill iconAmazon Web Services (AWS)

Location: Bangalore preferred / Hybrid as applicable

Experience: 3+ years

Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline

Salary: Above market standards, flexible for the right candidate

Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations


About FrontM

FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.

The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.


Role Summary

As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.

This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.


Key Responsibilities

Cloud Infrastructure & DevOps Architecture (≈45%)

· Own, maintain and improve AWS cloud infrastructure for FrontM platforms

· Create and maintain Terraform scripts for infrastructure deployment and management

· Manage Kubernetes workloads deployed within AWS EKS

· Support multi-zone AWS infrastructure design for availability, resilience and scale

· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap

CI/CD, Operations & Platform Reliability (≈35%)

· Build, maintain and improve CI/CD pipelines for backend and platform services

· Oversee technical operations with hands-on administration, monitoring and release support

· Ensure continuous server uptime, stability, performance and maintainability

· Debug, respond to and restore system outages in production and staging environments

· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io

· Support backend stability, scale and performance across Node.js, Java and related services

Security, Networking & Production Support (≈20%)

· Maintain AWS security configurations, access controls and monitoring practices

· Support complex networking requirements across multi-domain SaaS implementations

· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users

· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements

· Document operational procedures, incident findings and technical support steps clearly


Required Technical Skills

Cloud Infrastructure & AWS

· Strong hands-on experience with AWS infrastructure and cloud operations

· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Experience with AWS security setup, monitoring and multi-zone infrastructure

· Ability to manage infrastructure using Terraform

Kubernetes, CI/CD & Observability

· Strong experience with Kubernetes, preferably AWS EKS

· Extensive CI/CD and DevOps experience

· Experience with infrastructure observability and application monitoring tools

· Ability to diagnose production bottlenecks, server failures and performance issues

Backend, Networking & SaaS Operations

· Experience supporting Node.js, Java and backend system procedures for stability and scale

· Good understanding of APIs, integrations and backend service dependencies

· Experience with complex networking and multi-domain SaaS implementations

· Ability to troubleshoot technical issues with non-technical end users

Nice to Have

· Experience with MongoDB clusters in MongoDB Atlas

Personal Attributes

· Strong ownership mindset for uptime, reliability and production stability

· Practical problem-solving approach with the ability to act quickly during incidents

· Clear written and spoken communication in English

· Ability to work independently and coordinate with senior management when required

· Comfortable working in fast-moving engineering teams

· Attention to detail in security, monitoring, documentation and operational processes


Why join FrontM?

Long-Term Career Growth

Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.

Engineering Challenges That Matter

Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.

Broad Technical Ownership

Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.


Apply now

Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos