Cutshort logo
For Employers
One2n logo
Virtual Hiring Event on 25th April - Site Reliability Engineer
Virtual Hiring Event on 25th April - Site Reliability Engineer

Virtual Hiring Event on 25th April - Site Reliability Engineer at One2n · Pune · 3 - 7 years · ₹15L - ₹25L / yr · Bootstrapped · Posted 7 May 2026

One2n's logo

Virtual Hiring Event on 25th April - Site Reliability Engineer

Abhishek Ghorpade's profile picture
Posted by Abhishek Ghorpade
3 - 7 yrs
₹15L - ₹25L / yr
Pune
Skills
DevOps
skill iconKubernetes
skill iconDocker
skill icongrafana
prometheus
Terraform
Message Queuing Telemetry Transport (MQTT)
cicd
skill iconAmazon Web Services (AWS)
Data migration

Virtual Hiring Drive Site Reliability Engineer (SRE)


Date: 25th April 2026, Saturday (Single-Day Drive)

Mode: 100% Virtual - All interview rounds on the same day

Experience: 3 to 7 Years


Note : We are looking for quick joiners who can join us within 30 days.


About the Role

We are looking for a Site Reliability Engineer who understands the realities of running production systems at scale. If building reliable, scalable, and observable systems excites you, you'll enjoy working with us.

At One2N, we solve One-to-N problems where proof of concept is already built and the real challenge lies in scalability, maintainability, performance, and reliability.

You will work closely with startups and mid-sized clients, helping them architect production-grade infrastructure and observability systems.


Key Responsibilities

  • Design and build platform engineering solutions with a self-serve model
  • Architect and optimize observability systems (metrics, logs, traces)
  • Implement monitoring, logging, alerting & dashboards
  • Build and optimize CI/CD pipelines
  • Automate repetitive operational and infrastructure tasks (IaC-first approach)
  • Improve Developer Experience (DX)
  • Guide teams on SRE best practices & on-call processes
  • Participate in code reviews and mentor engineers
  • Contribute to cloud-native and platform engineering initiatives


Must-Have Skills

  • 3 - 7 years experience in DevOps / SRE / Platform Engineering
  • Strong hands-on with Kubernetes on AWS
  • Expertise in observability tools like Datadog / Honeycomb / ELK / Grafana / Prometheus
  • Experience with Docker & Microservices architecture
  • Infrastructure as Code using Terraform / Pulumi
  • Strong Linux troubleshooting skill
  • Programming knowledge in Golang / Python / Java
  • Automation & scripting expertise
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About One2n

Founded :
2019
Type :
Services
Size :
20-100
Stage :
Bootstrapped

About

One2N is a boutique technology consulting firm that helps fast-growing companies build and scale high-performance backend systems. As startups evolve from MVP to tens of thousands of users, the engineering challenges change dramatically — and we specialise in solving exactly those problems. Our work ensures that scaling, reliability, and performance never become bottlenecks to growth.


With deep expertise in Site Reliability Engineering, Cloud Infrastructure & DevOps, Data Engineering, and Backend Architecture, we partner with engineering teams to design, build, and operate resilient cloud-native systems. We believe in pragmatic, impact-driven engineering — creating solutions that are robust today and adaptable for tomorrow, so teams can ship faster, stay reliable, and scale confidently.

Read more

Candid answers by the company

What does the company do?
What is the location preference of jobs?
What is the work culture like at One2N?

One2N builds robust, scalable software systems for startups ready to grow from product-market fit to large-scale adoption.

Photos

Company featured pictures
Company featured pictures
Company featured pictures
Company featured pictures
Company featured pictures
Company featured pictures
Company featured pictures
Company featured pictures

Company social profiles

bloginstagramlinkedintwitter

Similar jobs (10)

It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
company logo
Bhattacharjee Akash
Posted by Bhattacharjee Akash
Bengaluru (Bangalore), Chennai, Mumbai, Hyderabad, Pune, Gurugram
3 - 10 yrs
₹12L - ₹35L / yr
Linux/Unix
skill iconKubernetes
Monitoring
skill iconDocker
skill iconAmazon Web Services (AWS)
+4 more



We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.



Key Responsibilities

  • Own reliability, availability, and performance of production services, including on-call rotation
  • Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
  • Automate deployments, scaling, and operational tasks to reduce manual toil
  • Manage containerized workloads on Kubernetes and cloud infrastructure
  • Design and maintain CI/CD pipelines for safe, frequent releases
  • Lead incident response and drive blameless postmortems with clear follow-ups
  • Perform capacity planning, performance tuning, and cost optimization
  • Define and track SLIs/SLOs and error budgets with product teams


Requirements

  • 3+ years in SRE, DevOps, or production-focused engineering
  • Strong Linux administration and hands-on Kubernetes experience
  • Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
  • Cloud experience with AWS, GCP, or Azure
  • CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
  • Proficient scripting in Python and/or Bash


Nice to have

  • Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
  • Background in high-traffic or distributed systems
Read more
Gurugram
4 - 10 yrs
₹4L - ₹10L / yr
DevOps
Site Reliability Engineer (SRE)
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
+14 more

🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience Level : 4+ Years

Location : Gurugram Sector 48, Haryana (On-site)

Employment Type : Full Time Opportunity


About the Role :

We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.

In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).


Mandatory Skills :

AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux


Key Responsibilities :

1. Cloud Infrastructure & Infrastructure as Code (IaC) :

  • Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
  • Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
  • Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.

2. CI/CD, Automation & Development :

  • Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
  • Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.

3. Containerization & Orchestration :

  • Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
  • Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.

4. Observability, SRE & Incident Management :

  • Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
  • Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
  • Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
  • Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).


Required Qualifications & Skills :

  • Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
  • Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
  • Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
  • CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
  • Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
  • Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
  • Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
  • Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
  • Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Read more
Gurugram
3 - 9 yrs
Best in industry
DevOps

Site Reliability Engineer (SRE) / DevOps Engineer (Walk-In Drive)

Location: Gurgaon Experience: 3–6 Years


About the Role

We are looking for a hands-on Site Reliability Engineer (SRE) / DevOps Engineer with strong programming and automation skills.

The role will initially involve development and automation work, helping the engineer build a strong understanding of the applications and platform. Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, observability, production reliability, incident management, and operational automation.

The ideal candidate should be comfortable working with both application code and production systems and should use engineering and automation to improve reliability and reduce manual effort.

Key Responsibilities

· Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.

· Write clean, maintainable, testable, and production-ready code.

· Participate in code reviews, debugging, testing, and technical discussions.

· Build, maintain, and improve CI/CD pipelines and automated deployment processes.

· Work with Docker and Kubernetes for application deployment and operations.

· Support on prem and cloud-based application and infrastructure deployments.

· Maintain reliable, scalable, secure, and highly available production environments.

· Implement and manage monitoring, logging, alerting, and observability solutions.

· Contribute to defining and tracking SLIs, SLOs, and error budgets.

· Troubleshoot application and production issues and perform Root Cause Analysis (RCA).

· Identify recurring operational problems and address them through automation and engineering improvements.

· Support incident response, change management, deployment governance, and disaster recovery practices.

· Maintain runbooks, SOPs, incident documentation, and technical documentation.

· Collaborate with Engineering, Product, Platform, Security, Operations, and external teams.

Technical Skills

· Strong hands-on experience with Python for development and automation.

· Experience developing scripts, APIs, integrations, utilities, or backend services.

· Good understanding of software engineering principles, debugging, logging, testing, and exception handling.

· Experience with REST APIs, JSON, Git, pull requests, and code reviews.

· Strong knowledge of Linux/Unix environments and basic Windows administration.

· Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.

· Experience with at least one cloud platform: AWS, Azure, or GCP.

· Hands-on experience with Docker and Kubernetes.

· Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.

· Familiarity with Infrastructure-as-Code tools such as Terraform is preferred.

· Experience with monitoring and observability tools such as Grafana, Prometheus, Power BI, or equivalent.

· Ability to analyze logs, metrics, alerts, and traces for troubleshooting.

· Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.

· Experience with JIRA, ServiceNow, and Confluence is desirable.

Preferred Experience

· 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, or related roles.

· Strong Python development or automation experience.

· Experience supporting applications across development, deployment, and production environments.

· Exposure to cloud-native, distributed, or production-grade systems.

· Understanding of security and compliance best practices.

· Familiarity with AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools.

Soft Skills

· Strong analytical and troubleshooting skills.

· Engineering and automation mindset.

· Good written and verbal communication skills.

· Effective cross-functional collaboration.

· Ownership-driven approach to problem solving.

· Ability to remain structured during production incidents.

Read more
ride hailing app
ride hailing app
Agency job
via by Kaushik Reddyshetty
Hyderabad
3 - 6 yrs
₹20L - ₹25L / yr
skill iconAmazon Web Services (AWS)
skill iconKubernetes
skill iconDocker
CI/CD
Microservices
+10 more

Platform Engineer

Location: Hyderabad, Telangana — On-site

Experience: 3–6 Years

Employment Type: Full-time


Hyderabad-based mobility-tech startup building the technology infrastructure behind student transportation.

We operate a real-time transportation platform that brings together student tracking, routing, parent notifications, driver applications, operations dashboards, and cloud infrastructure to make student mobility safer, more reliable, and easier to manage.

As we scale, we’re looking for a Platform Engineer who can take ownership of the infrastructure and platform layer that powers these systems.

The Role

As a Platform Engineer, you will own the systems that enable our engineering teams to build, deploy, scale, monitor, and operate reliable production services.

This is an early-stage startup role with significant ownership. You will work closely with engineering and product teams to build infrastructure from the ground up, improve deployment velocity, strengthen reliability, and ensure our platform can scale with the business.

You should be comfortable moving between cloud infrastructure, Kubernetes, CI/CD, observability, security, and distributed systems.

What You'll Do

  • Design, build, and maintain scalable cloud infrastructure on AWS
  • Own production infrastructure across EC2, IAM, RDS, networking, monitoring, and deployments
  • Build and maintain Docker and Kubernetes environments for production workloads
  • Develop and improve CI/CD pipelines for reliable and rapid deployments
  • Manage infrastructure as code using Terraform
  • Establish infrastructure standards for scalability, security, reliability, and cost efficiency
  • Monitor production systems and proactively identify performance and reliability issues
  • Build observability around applications and infrastructure, including metrics, logs, alerts, and incident monitoring
  • Troubleshoot production issues and conduct root-cause analysis (RCA)
  • Work with backend engineers to design infrastructure for Java/Kotlin-based distributed systems
  • Support highly available services involving real-time tracking, routing, notifications, and operational workflows
  • Improve deployment processes, release reliability, rollback strategies, and disaster recovery
  • Identify infrastructure bottlenecks and continuously improve platform performance
  • Help establish engineering practices around reliability, security, and operational excellence
  • Work across the stack when required and take end-to-end ownership of infrastructure problems

What We're Looking For

Must Have

  • 3–6 years of experience in Platform Engineering, DevOps, SRE, Infrastructure Engineering, or a closely related role
  • Strong hands-on experience with AWS
  • Strong understanding of EC2, IAM, RDS, networking, monitoring, and production deployments
  • Experience with Docker and Kubernetes
  • Strong experience building and managing CI/CD pipelines
  • Hands-on experience with Terraform / Infrastructure as Code
  • Strong Linux and networking fundamentals
  • Experience troubleshooting production systems and performing RCA
  • Understanding of distributed systems, scalability, availability, and system design
  • Experience working with backend services built using Java/Kotlin or similar technologies
  • Ability to independently own infrastructure problems from design → implementation → deployment → monitoring

Good to Have

  • Experience with Redis, PostgreSQL, or MongoDB
  • Experience with microservices architecture
  • Experience with AWS security and IAM best practices
  • Experience building observability and alerting systems
  • Experience with multi-cloud environments such as AWS, Azure, or GCP
  • Experience working in an early-stage startup
  • Experience with real-time systems, IoT, location services, or mobility platforms
  • Open-source contributions or meaningful personal engineering projects

What Makes This Role Different

At ZeroMoblt, you won't be working within a large infrastructure team where responsibilities are narrowly defined.

You'll have the opportunity to:

  • Own critical infrastructure decisions
  • Build platform capabilities from 0 → 1
  • Work directly with engineering and product teams
  • Solve real-world scalability and reliability problems
  • Influence architecture and engineering practices
  • See your work directly impact a platform serving 10,000+ students
  • Work in a fast-moving environment with minimal bureaucracy

We're looking for someone who enjoys ownership, ambiguity, and solving problems independently.

This role may not be the right fit if you prefer highly structured processes, narrowly defined responsibilities, or large-company environments with multiple layers of ownership.

Why ZeroMoblt?

  • High ownership and autonomy
  • Direct exposure to product and engineering decisions
  • Opportunity to build infrastructure at an early-stage mobility startup
  • Work on real-time transportation and location-based systems
  • Hyderabad-based, on-site team


Read more
Gurugram
5 - 10 yrs
₹12L - ₹18L / yr
DevOps
Reliability engineering
skill iconAmazon Web Services (AWS)
Terraform
Ansible
+18 more

Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience : 5+ Years

Location : Gurugram, Haryana

Work Mode : On-site (Full-time)


About the Role :

We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.


Mandatory Skills :

AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux


Key Responsibilities :

  • Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
  • Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
  • Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
  • Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
  • Automate operational tasks using Python, Bash, and Groovy scripting.
  • Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.


Required Skills & Qualifications :

  • Bachelor's degree in Computer Science, IT, Electronics, or a related field.
  • 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
  • Strong expertise in AWS, with exposure to Azure/GCP.
  • Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
  • Strong scripting skills in Python and Bash.
  • Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Good understanding of Linux, networking, SQL, and cloud security best practices.


Preferred Skills :

  • Experience with multi-cloud environments and DevSecOps practices.
  • Knowledge of disaster recovery, automation, and microservices architecture.
  • Strong troubleshooting, communication, and problem-solving skills.
Read more
company logo
Priya Rawat
Posted by Priya Rawat
Gurugram
4 - 5 yrs
₹8L - ₹10L / yr
RCA
SLA
skill icongrafana
ELKI
SOP

About the Role


We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.


Key Responsibilities


  • Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
  • Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
  • Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
  • Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
  • Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
  • Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
  • Participate in incident/severity calls, ensuring clear communication and coordination
  • Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting


Required Skills & Experience


  • Strong understanding of Linux/Unix systems for application support
  • Hands-on experience troubleshooting applications in staging and production environments
  • Ability to monitor system performance and identify root causes using logs and metrics
  • Experience working with Kubernetes and microservices-based architectures
  • Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
  • Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
  • Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
  • Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
  • Experience with ticketing and documentation tools such as Jira and Confluence
  • Minimum 4+ years of experience in application support or reliability engineering


Education & Certifications


  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Relevant certifications (Cloud, Kubernetes, Microservices) are a plus


Work Schedule


  • Willingness to work in a 24x7 environment, including weekends and on-call rotations
Read more
Remote only
8 - 12 yrs
Best in industry
Terraform
Artificial Intelligence (AI)
IAC
skill iconAmazon Web Services (AWS)
ECS
+6 more


Senior Platform & Site Reliability Engineer

Location: Remote Employment Type: Contract

The Role

This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.

Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.

This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.

The Scale You Will Operate At

The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.

You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.

What You Will Own

Platform Architecture

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
  • Define, build, and enforce platform standards across portfolio products
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
  • Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
  • Design and maintain batch processing infrastructure alongside live event flows
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products
  • SLOs and error budgets defined per product; reliability tracked consistently
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation
  • Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
  • Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
  • Target: full platform integration within a defined window per acquisition

A Note on Automation

Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.

Platform Stack

Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
  • Strong Terraform depth across multi-environment, multi-account setups
  • CI/CD ownership across a multi-product environment with GitHub Actions
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
  • Hands-on Grafana, Prometheus, and Loki in production
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
  • Acquisition or greenfield platform integration experience strongly preferred

How You Work

  • Comfortable operating across multiple products simultaneously — context-switching without dropping standards
  • Cost-efficiency instinct — you optimise spend as a habit, not as a project
  • You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
  • You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use

Why This Role

The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.

This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.

Read more
company logo
Sakshi Mittal
Posted by Sakshi Mittal
Bengaluru (Bangalore)
3 - 5 yrs
₹6L - ₹12L / yr
skill iconAmazon Web Services (AWS)
DevOps
skill iconKubernetes
Terraform
CI/CD
+1 more

Job Summary :

We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.

Key Responsibilities :

- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.

- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.

- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management

- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.

- Collaborate with developers to ensure smooth and frequent deployments to production.

- Manage versioning and rollback strategies for critical deployments.

- Containerization & Orchestration using Kubernetes.

- Containerize applications using Docker, and manage them using Kubernetes.

- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.

- Develop utilities and tools to enhance operational efficiency and reliability.

- Monitoring & Incident Management

- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.

- Optimize application and system performance through proactive monitoring and configuration tuning.

Desired Skills and Experience :

- Experience Required - 6+ yrs.

- Hands-on experience on cloud services like AWS, EKS etc.

- Ability to design a good cloud solution.

- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.

- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.

- Use knowledge and research to constantly modernize our applications and infrastructure stacks.

- Be a team player and strong problem-solver to work with a diverse team.

- Having good communication skills.

Read more
company logo
Arshiya Shaikh
Posted by Arshiya Shaikh
Mumbai
4 - 7 yrs
₹5L - ₹9L / yr
AWS CloudFormation
Bitbucket
skill iconDocker
skill iconKubernetes

Key Responsibilities

  • Automate application deployments from Bitbucket to servers using CI/CD pipelines.
  • Design and manage scalable, highly available AWS infrastructure.
  • Implement Auto Scaling, ELB, and Route 53 for traffic management and high availability.
  • Work with AWS services including IAM, RDS, DynamoDB, EC2, and other cloud services.
  • Build and manage Docker containers and server images.
  • Deploy and manage applications using Kubernetes.
  • Implement Infrastructure as Code using Terraform, CloudFormation, or Ansible.
  • Develop automation scripts using Python and Bash.
  • Implement monitoring and logging using tools such as Prometheus, Grafana, and ELK.
  • Integrate security and compliance practices into CI/CD pipelines.
  • Optimize infrastructure for security, scalability, performance, and cost.

Required Skills

  • 3+ years of experience in DevOps or a similar role.
  • Strong knowledge of AWS beyond EC2.
  • Hands-on experience with Jenkins or similar CI/CD tools.
  • Experience with Docker and Kubernetes.
  • Good understanding of Terraform/IaC and automation.
  • Proficiency in Python and/or Bash scripting.
  • Knowledge of DevSecOps, security, and compliance best practices.
  • Strong troubleshooting and problem-solving skills.


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos