Cutshort logo
For Employers
Interfaceai logo
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Senior Site Reliability Engineer at Interfaceai · Remote only · 5 - 8 years · ₹20L - ₹35L / yr · Raised funding · Remote only · Posted 8 Feb 2022

Interfaceai's logo

Senior Site Reliability Engineer

Shashikant MS's profile picture
Posted by Shashikant MS
5 - 8 yrs
₹20L - ₹35L / yr
Remote only
Skills
skill iconDocker
skill iconKubernetes
DevOps
skill iconAmazon Web Services (AWS)
Windows Azure
Google Cloud Platform (GCP)
CI/CD
Infrastructure
Reliability

About Us

We have grown over 1400% in revenues in the last year.

Interface.ai provides an Intelligent Virtual Assistant (IVA) to FIs to automate calls and customer inquiries across multiple channels and engage their customers with financial insights and upsell/cross-sell.

Our IVA is transforming financial institutions’ call centers from a cost to a revenue center.

 

Our core technology is built 100% in-house with several breakthroughs in Natural Language Understanding. Our parser is built based on zero-shot learning that helps us to launch industry-specific IVA that can achieve over 90% accuracy on Day-1. 

We are 45 people strong with employees spread across India and US locations. Many of them come from ML teams at Apple, Microsoft, and Salesforce in the US along with enterprise architects with over 20+ years of experience building large-scale systems. Our India team consists of people from ISB, IIMs, and many who have been previously part of early-stage startups. 

 

We are a fully remote team.

 

Founders come from Banking and Enterprise Technology backgrounds with previous experience scaling companies from scratch to $50M+ in revenues.

As a Site Reliability Engineer you will be in charge of:

  • Designing, analyzing and troubleshooting large-scale distributed systems
  • Engaging in cross-functional team discussions on design, deployment, operation, and maintenance,  in a fast-moving, collaborative set up
  • Building automation scripts to validate the stability, scalability, and reliability of interface.ai’s products & services as well as enhance interface.ai’s employees’ productivity
  • Debugging and optimizing code and automating routine tasks
  • Troubleshoot and diagnose issues (hardware or software), propose and implement solutions to ensure they occur with reduced frequency
  • Perform the periodic on-call duty to handle security, availability, and reliability of interface.ai’s products 
  • You will follow and write good code and solid engineering practices

 

Requirements

You can be a great fit if you are :

  1. Extremely self motivated
  2. Ability to learn quickly
  3. Growth Mindset (read this if you don't know what it means - https://www.amazon.com/Mindset-Psychology-Carol-S-Dweck/dp/0345472322" target="_blank">link)
  4. Emotional Maturity (read this if you don't know what it means - https://medium.com/@krisgage/15-signs-of-emotional-maturity-38b1a2ab9766" target="_blank">link)
  5. Passionate about the possibilities at the intersection of AI + Banking
  6. Worked in a startup of 5 to 30 employees
  7. Developer with a strong interest in systems Design. You will be building, maintaining, and scaling our cloud infrastructure through software tooling and automation. 
  8. 4-8 years of industry experience developing and troubleshooting large-scale infrastructure on the cloud
  9. Have a solid understanding of system availability, latency, and performance
  10. Strong programming skills in at least one major programming language and the ability to learn new languages as needed  
  11. Strong System/network debugging skills
  12. Experience with management/automation tools such as Terraform/Puppet/Chef/SALT
  13. Experience with setting up production-level monitoring and telemetry
  14. Expertise in Container management & AWS
  15. Experience with kubernetes is a plus
  16. Experience building CI/CD pipelines
  17. Experience working with Web sockets, Redis, Postgres, Elastic search, Logstash
  18. Experience working in an agile team environment and proficient understanding of code versioning tools, such as Git.
  19. Ability to effectively articulate technical challenges and solutions.
  20. Proactive outlook for ways to make our systems more reliable
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Interfaceai

Founded :
2019
Type :
Product
Size :
20-100
Stage :
Raised funding

About

interface powered Intelligent Virtual Assistants or IVAs make every digital channel of an enterprise intelligent. With our IVAs, an enterprise can leapfrog its customer and employee experiences to voice and text-first based interactions. interface is a platform that can be integrated with your business to provide conversational assistance. With state-of-the-art technology at its core, it can help your business grow manifold.
Read more

Connect with the team

Profile picture
Shashikant MS
Profile picture
Shakun Banthia

Company social profiles

linkedin

Similar jobs (10)

It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
company logo
Harsha Mehrotra
Posted by Harsha Mehrotra
Remote only
3 - 5 yrs
Best in industry
CI/CD
skill iconPostgreSQL
skill iconPython
skill iconAmazon Web Services (AWS)
Artificial Intelligence (AI)
+2 more

SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.


Core responsibilities include:

  • Production monitoring and debugging of live systems.
  • Incident investigation, troubleshooting, and problem resolution.
  • AWS cloud infrastructure support and maintenance.
  • Deployment and operational support activities.
  • Supporting a 24x7 production environment.
  • Working with GitHub-based development workflows.
  • Technical debt remediation and platform improvements.
  • Customer issue investigation and support.
  • Security and compliance-related work, including FedRAMP initiatives.


Preferred Skills:

AWS (especially S3 and EC2)

Strong debugging and troubleshooting skills

Site Reliability Engineering (SRE) experience

GitHub experience

Basic software development skills

TypeScript/JavaScript knowledge

C# preferred

AI experience is a plus.


Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Read more
company logo
Priya Rawat
Posted by Priya Rawat
Gurugram
4 - 5 yrs
₹8L - ₹10L / yr
RCA
SLA
skill icongrafana
ELKI
SOP

About the Role


We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.


Key Responsibilities


  • Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
  • Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
  • Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
  • Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
  • Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
  • Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
  • Participate in incident/severity calls, ensuring clear communication and coordination
  • Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting


Required Skills & Experience


  • Strong understanding of Linux/Unix systems for application support
  • Hands-on experience troubleshooting applications in staging and production environments
  • Ability to monitor system performance and identify root causes using logs and metrics
  • Experience working with Kubernetes and microservices-based architectures
  • Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
  • Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
  • Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
  • Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
  • Experience with ticketing and documentation tools such as Jira and Confluence
  • Minimum 4+ years of experience in application support or reliability engineering


Education & Certifications


  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Relevant certifications (Cloud, Kubernetes, Microservices) are a plus


Work Schedule


  • Willingness to work in a 24x7 environment, including weekends and on-call rotations
Read more
Gurugram
3 - 9 yrs
Best in industry
DevOps

Site Reliability Engineer (SRE) / DevOps Engineer (Walk-In Drive)

Location: Gurgaon Experience: 3–6 Years


About the Role

We are looking for a hands-on Site Reliability Engineer (SRE) / DevOps Engineer with strong programming and automation skills.

The role will initially involve development and automation work, helping the engineer build a strong understanding of the applications and platform. Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, observability, production reliability, incident management, and operational automation.

The ideal candidate should be comfortable working with both application code and production systems and should use engineering and automation to improve reliability and reduce manual effort.

Key Responsibilities

· Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.

· Write clean, maintainable, testable, and production-ready code.

· Participate in code reviews, debugging, testing, and technical discussions.

· Build, maintain, and improve CI/CD pipelines and automated deployment processes.

· Work with Docker and Kubernetes for application deployment and operations.

· Support on prem and cloud-based application and infrastructure deployments.

· Maintain reliable, scalable, secure, and highly available production environments.

· Implement and manage monitoring, logging, alerting, and observability solutions.

· Contribute to defining and tracking SLIs, SLOs, and error budgets.

· Troubleshoot application and production issues and perform Root Cause Analysis (RCA).

· Identify recurring operational problems and address them through automation and engineering improvements.

· Support incident response, change management, deployment governance, and disaster recovery practices.

· Maintain runbooks, SOPs, incident documentation, and technical documentation.

· Collaborate with Engineering, Product, Platform, Security, Operations, and external teams.

Technical Skills

· Strong hands-on experience with Python for development and automation.

· Experience developing scripts, APIs, integrations, utilities, or backend services.

· Good understanding of software engineering principles, debugging, logging, testing, and exception handling.

· Experience with REST APIs, JSON, Git, pull requests, and code reviews.

· Strong knowledge of Linux/Unix environments and basic Windows administration.

· Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.

· Experience with at least one cloud platform: AWS, Azure, or GCP.

· Hands-on experience with Docker and Kubernetes.

· Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.

· Familiarity with Infrastructure-as-Code tools such as Terraform is preferred.

· Experience with monitoring and observability tools such as Grafana, Prometheus, Power BI, or equivalent.

· Ability to analyze logs, metrics, alerts, and traces for troubleshooting.

· Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.

· Experience with JIRA, ServiceNow, and Confluence is desirable.

Preferred Experience

· 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, or related roles.

· Strong Python development or automation experience.

· Experience supporting applications across development, deployment, and production environments.

· Exposure to cloud-native, distributed, or production-grade systems.

· Understanding of security and compliance best practices.

· Familiarity with AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools.

Soft Skills

· Strong analytical and troubleshooting skills.

· Engineering and automation mindset.

· Good written and verbal communication skills.

· Effective cross-functional collaboration.

· Ownership-driven approach to problem solving.

· Ability to remain structured during production incidents.

Read more
company logo
Pune, Mumbai
4 - 8 yrs
₹6L - ₹22L / yr
Reliability engineering
DevOps
Google Cloud Platform (GCP)
Alerting and Monitoring
skill iconKubernetes
+6 more

Lead Cloud Reliability Engineer


Job Responsibilities

● Lead and manage the Cloud Reliability teams to provide strong Managed Services support to end-customers.

● Isolate, troubleshoot and resolve issues reported by CMS clients in their cloud environment

● Drive the communication with the customer providing details about the issue, current steps, next plan of action, ETA

● Gather client's requirements related to use of specic cloud services and provide assistance in seing them up and resolving issues

● Create SOPs and knowledge articles for use by the L1 teams to resolve common issues

● Identify recurring issues, perform root cause analysis and propose/implement preventive actions

● Follow change management procedure to identify, record and implement changes

● Plan and deploy OS, security patches in Windows/Linux environment and upgrade k8s clusters

● Identify the recurring manual activities and contribute to automation

● Provide technical guidance and educate team members on development and operations. Monitor metrics and develop ways to improve.

● System troubleshooting and problem-solving across plaorm and application domains. Ability to use a wide variety of open-source technologies and cloud services.

● Build, maintain, and monitor conguration standards.

● Ensuring critical system security through using best-in-class cloud security solutions.


Qualifications

● 4-7 years experience in Cloud Infrastructure and Operations domains and IT operational experience preferably in a global enterprise environment.

● Specialize in one or two cloud deployment platforms: AWS, GCP

● Hands on experience with AWS/GCP services (EKS, ECS, EC2, VPC, RDS, Lambda, GKE, Compute Engine)

● Understanding of one or more programming languages (Python, JavaScript, Ruby, Java, .Net)

● Logging and Monitoring tools (ELK, Stackdriver, CloudWatch)

● Knowledge on Conguration Management tools such as Ansible, Terraform, Puppet, Chef

● Experience working with deployment and orchestration technologies (such as Docker, Kubernetes, Mesos)

● Good analytical, communication, problem solving, and learning skills.

● Knowledge on programming against cloud plaorms such as Google Cloud Platform and lean development methodologies.

● Strong service aitude and a commitment to quality.

● Willingness to work in shifts.

Read less


Read more
company logo
Pune, Gurugram, Bengaluru (Bangalore), Hyderabad
5 - 12 yrs
₹15L - ₹28L / yr
DevOps
skill iconKubernetes
Incident management
Observability
Reliability engineering
+4 more

Lead Cloud Reliability Engineer


Job Responsibilities

● Lead and manage the Cloud Reliability teams to provide strong Managed Services support to end-customers.

● Isolate, troubleshoot and resolve issues reported by CMS clients in their cloud environment

● Drive the communication with the customer providing details about the issue, current steps, next plan of action, ETA

● Gather client's requirements related to use of specic cloud services and provide assistance in seing them up and resolving issues

● Create SOPs and knowledge articles for use by the L1 teams to resolve common issues

● Identify recurring issues, perform root cause analysis and propose/implement preventive actions

● Follow change management procedure to identify, record and implement changes

● Plan and deploy OS, security patches in Windows/Linux environment and upgrade k8s clusters

● Identify the recurring manual activities and contribute to automation

● Provide technical guidance and educate team members on development and operations. Monitor metrics and develop ways to improve.

● System troubleshooting and problem-solving across plaorm and application domains. Ability to use a wide variety of open-source technologies and cloud services.

● Build, maintain, and monitor conguration standards.

● Ensuring critical system security through using best-in-class cloud security solutions.


Qualifications

● 4-7 years experience in Cloud Infrastructure and Operations domains and IT operational experience preferably in a global enterprise environment.

● Specialize in one or two cloud deployment platforms: AWS, GCP

● Hands on experience with AWS/GCP services (EKS, ECS, EC2, VPC, RDS, Lambda, GKE, Compute Engine)

● Understanding of one or more programming languages (Python, JavaScript, Ruby, Java, .Net)

● Logging and Monitoring tools (ELK, Stackdriver, CloudWatch)

● Knowledge on Conguration Management tools such as Ansible, Terraform, Puppet, Chef

● Experience working with deployment and orchestration technologies (such as Docker, Kubernetes, Mesos)

● Good analytical, communication, problem solving, and learning skills.

● Knowledge on programming against cloud plaorms such as Google Cloud Platform and lean development methodologies.

● Strong service aitude and a commitment to quality.

● Willingness to work in shifts.

Read more
Product Based Co
Product Based Co
Agency job
via by Rishika Teja
Hyderabad
18 - 25 yrs
₹70L - ₹80L / yr
SRE
skill iconAmazon Web Services (AWS)

Hiring SRE - Director


Exp : 18 - 25 yrs

Edu : BE/B.Tech

Work Location : Hyd


Must Have Skills :


Must be from product SaaS based companies.


18+ years in Software Engineering, SRE, or reliability roles; 5+ years in leadership(Director). 


Proven ability to leverage software engineering principles and practices to solve reliability and operational challenges.


Expertise in SLI/SLO and monitoring.


Expertise in CI/CD, observability, and incident response. 


Strong AWS knowledge and experience with container orchestration. 


Proven ability to lead reliability programs across multiple SaaS products. 


Experience architecting applications or infrastructure for high-growth cloud platforms. 


Experience in B2B SaaS environments involving large-scale distributed systems. 



Read more
Remote only
8 - 12 yrs
Best in industry
Terraform
Artificial Intelligence (AI)
IAC
skill iconAmazon Web Services (AWS)
ECS
+6 more


Senior Platform & Site Reliability Engineer

Location: Remote Employment Type: Contract

The Role

This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.

Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.

This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.

The Scale You Will Operate At

The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.

You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.

What You Will Own

Platform Architecture

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
  • Define, build, and enforce platform standards across portfolio products
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
  • Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
  • Design and maintain batch processing infrastructure alongside live event flows
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products
  • SLOs and error budgets defined per product; reliability tracked consistently
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation
  • Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
  • Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
  • Target: full platform integration within a defined window per acquisition

A Note on Automation

Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.

Platform Stack

Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
  • Strong Terraform depth across multi-environment, multi-account setups
  • CI/CD ownership across a multi-product environment with GitHub Actions
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
  • Hands-on Grafana, Prometheus, and Loki in production
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
  • Acquisition or greenfield platform integration experience strongly preferred

How You Work

  • Comfortable operating across multiple products simultaneously — context-switching without dropping standards
  • Cost-efficiency instinct — you optimise spend as a habit, not as a project
  • You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
  • You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use

Why This Role

The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.

This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.

Read more
company logo
Robin Silverster
Posted by Robin Silverster
Bengaluru (Bangalore)
7 - 11 yrs
₹10L - ₹36L / yr
SRE
Reliability engineering
Google Cloud Platform (GCP)

Location - Bangalore Skill/Experience Expectations: 1. Total Experience 7-11 yrs 2. 3-4 years in managing scalable production environment 3. 2-4 yr experience in managing Google cloud infrastructure 4. proficient in terraform and any programming language 5. Expert in designing and managing observability solutions 6. 5 yr experience in DevOps and SRE practices and troubleshooting critical incidents.

Read more
MNC
MNC
Agency job
via by Sandhiya b
Bengaluru (Bangalore), Hyderabad
8 - 12 yrs
₹2L - ₹22L / yr
Production support
Reliability engineering
Linux/Unix
Monitoring
skill icongrafana

Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting

WFO-Immediate

8 to 12 Yrs

Bangalore/Hyderabad

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos