Cutshort logo
For Employers
Full stack fleet software company (Established startup) logo
Senior Platform Engineer
Full stack fleet software company (Established startup)
Senior Platform Engineer

Senior Platform Engineer at Full stack fleet software company (Established startup) · Remote only · 2 - 8 years · ₹15L - ₹35L / yr · Remote only · Posted 9 Dec 2021

Qrata's logo

Senior Platform Engineer

at Full stack fleet software company (Established startup)

Agency job
via Qrata
2 - 8 yrs
₹15L - ₹35L / yr
Remote only
Skills
Platform as a Service (PaaS)
DevOps
skill iconJava
skill iconPython
skill iconAmazon Web Services (AWS)
Windows Azure
Google Cloud Platform (GCP)
Requirements
● Good understanding of how the web works
● Experience with at least one language like Java, Python etc
● Good with Shell scripting
● Experience with *Nix based operating systems
● Experience with k8s, containers
● Fairly good understanding of AWS/GCP/Azure
● Troubleshoot and fix outages and performance issues in infrastructure stack
● Identify gap and design automation tools for all feasible functions in infrastructure
● Good verbal and written communication skills
● Drive SLA/SLO of team
Benefits
This is an opportunity to work on a fairly complex set of systems and improve
them. You will get a chance to learn things like “how to think about code
simplicity”, “how to write for maintainability” and several other things.
● Comprehensive health insurance policy.
● Flexible working hours and a very friendly work environment.
● Flexibility to work either in the office (post Covid) or remotely.
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

Agami Tech
at Agami Tech
3 recruiters
Digish Shah
Posted by Digish Shah
Mumbai
1 - 3 yrs
Best in industry
RHCSA
Linux/Unix
Cloud Computing
VMWare
Firewall
+5 more

Location : Mumbai


Big Picture (The Opportunity) :

Are you looking for an opportunity to advance your Career? & If you are able to maintain a positive attitude even when everything goes wrong, if you are detail oriented and self motivated with a passion to learn and improve your skills and knowledge, we have a perfect job for you !

What do we want from you ? (Our Expectations) :

  • Zero to 2 years experience in Linux Operating System.
  • Flexible working hours - able to support occasional nights, weekends, and call-ins, able to quickly adapt to a constantly changing faced paced environment. Open to travel to sites.
  • An ideal person who is excited and motivated about running and supporting a production - grade critical infrastructure and looks for opportunities to improve processes with automation.


What are you required to do ? (Your Responsibilities) :

  • Proactively maintain and develop all linux infrastructure technology to maintain a 24*7*365 uptime service.
  • Engineering of systems administration-related various solutions for our various SAAS products and projects as well as operational needs.
  • Proactively monitoring system performance and capacity planning.
  • Providing technical support to customers for Applications, Operating systems, and networking.
  • Will also be the first point of contact for our clients, where installations are placed on a permanent basis, for basic troubleshooting and problem solving.
  • Fault finding, analysis and logging information for reporting of performance exceptions.
  • Maintain best practices on managing systems and services across all environments.

 

 

Skills & Qualification Required (Add Value) :

  • Graduate - Preferred to have Bachelor‘s Degree in Engineering, Computer Science or related field.
  • He / She should be familiar with the installation and configuration of Linux operating systems and setup and operation of TCP/IP networking on Linux systems also familiar with Internet concepts including SMTP, IMAP, POP, HTTP, DNS, LDAP and related protocols.
  • You should possess excellent communication skills .
  • Knowledge of Email concepts, Helpdesk Concept , VoIP Concept , Cloud computing will be an added advantage.
Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen
Chennai
5 - 10 yrs
₹2L - ₹15L / yr
DevOps
Linux/Unix
AppDynamics
skill icongrafana
prometheus
+1 more

DevOps / Infrastructure Engineer

Location: Chennai

Experience: 5+ Years

Role: DevOps / Infrastructure Engineer

Job Description

We are looking for an experienced DevOps / Infrastructure Engineer with strong hands-on experience in Linux administration, containerization, Kubernetes, automation, monitoring, and troubleshooting.

Mandatory Skills

  • 5+ years of experience in DevOps / Infrastructure Administration
  • Strong hands-on experience with Linux Administration
  • Experience with Docker and Kubernetes
  • Monitoring tools: AppDynamics, Prometheus, Grafana
  • Strong Shell Scripting / Python Scripting skills
  • Hands-on experience with Ansible 4.1
  • Strong troubleshooting and problem-solving skills
  • Experience in infrastructure/application monitoring and production support
  • Good understanding of DevOps practices and automation

Key Responsibilities

  • Manage and support Linux-based infrastructure and production environments.
  • Deploy, manage, and troubleshoot applications using Docker and Kubernetes.
  • Develop and maintain automation scripts using Shell/Python.
  • Automate infrastructure and configuration management using Ansible.
  • Monitor applications and infrastructure using AppDynamics, Prometheus, and Grafana.
  • Perform root-cause analysis and resolve infrastructure/application issues.
  • Handle incidents, troubleshoot performance issues, and ensure system availability.
  • Collaborate with development and operations teams to improve deployment and operational processes.
Read more
NeoGenCode Technologies Pvt Ltd
Akshay Patil
Posted by Akshay Patil
Gurugram
2 - 5 yrs
₹4L - ₹7.2L / yr
DevOps
skill iconAmazon Web Services (AWS)
Linux/Unix
Linux administration
skill iconJenkins
+5 more

Job Title : DevOps / Infrastructure Engineer

Experience : 3+ Years

Location : Gurugram Sector 48

Work Mode : 6 Days WFO (Monday to Saturday) / 01st & 03rd Saturdays are off

Employment Type : Full-time


Role Overview :

We are looking for a DevOps / Infrastructure Engineer with strong hands-on experience in Linux administration, Jenkins, Docker, networking, bare-metal servers, AWS, Redis, and MongoDB. The candidate should be capable of independently troubleshooting infrastructure, deployment, networking, and application-related issues in production environments.


Mandatory / Non-Negotiable Skills :

  • Strong hands-on experience with Linux Administration & Troubleshooting.Strong experience with Jenkins and CI/CD pipelines.
  • Hands-on experience with Docker and containerized environments.
  • Strong understanding of Networking concepts – TCP/IP, DNS, HTTP/HTTPS, ports, routing, firewalls, load balancing, etc.
  • Hands-on experience with Bare Metal Servers / Server Administration.
  • Strong hands-on experience with AWS (EC2, VPC, IAM, Security Groups, Load Balancers, S3 & CloudWatch)
  • Working knowledge of Redis
  • Working knowledge of MongoDB
  • Strong production troubleshooting and incident-resolution skills


Key Responsibilities :

  • Manage, configure, monitor, and troubleshoot Linux and bare-metal servers
  • Build, maintain, and troubleshoot Jenkins CI/CD pipelines
  • Deploy, manage, and troubleshoot applications using Docker
  • Manage and troubleshoot AWS infrastructure and services
  • Configure and maintain networking, security groups, firewalls, ports, and connectivity
  • Support and maintain Redis and MongoDB environments
  • Perform server health checks, log analysis, performance troubleshooting, and incident resolution
  • Work closely with development teams to support application deployments
  • Identify root causes of infrastructure and production issues and implement preventive solutions
  • Maintain infrastructure security, availability, and reliability
  • Automate repetitive operational tasks wherever possible


Required Candidate Profile :

  • 3+ years of relevant experience in DevOps, Infrastructure, System Administration, or related roles.
  • Strong hands-on / production experience with all mandatory technologies.
  • Good understanding of Linux systems and infrastructure.
  • Strong troubleshooting and problem-solving abilities.
  • Ability to take ownership of production infrastructure and deployment issues.
  • Good communication and collaboration skills.
Read more
NeoGenCode Technologies Pvt Ltd
Gurugram
3 - 6 yrs
₹4L - ₹9L / yr
DevOps
Linux/Unix
Networking
Bare Metal
Server administration
+13 more

Job Title : DevOps Engineer – Linux, AWS & Infrastructure

Experience : 3+ Years

Location : Sector 48, Gurgaon

Work Mode : 6 Days WFO – Monday to Saturday

Week Off : 01st & 03rd Saturday Off

Employment Type : Full-Time


Job Summary :

We are looking for a DevOps Engineer with strong hands-on experience in Linux, Networking, Server Administration, AWS, CI/CD, Containers, Kubernetes, and Infrastructure Automation. The ideal candidate should have strong troubleshooting skills, production ownership, and the ability to manage and automate infrastructure reliably.


Key Responsibilities :

  • Manage and troubleshoot Linux servers, bare-metal infrastructure, and server environments.
  • Perform system administration, networking, performance monitoring, and production troubleshooting.
  • Design, maintain, and optimize Jenkins-based CI/CD pipelines and deployment workflows.
  • Manage Docker containers and Kubernetes environments.
  • Work with AWS Cloud services, infrastructure, security, and deployment environments.
  • Monitor system health, application performance, logs, and infrastructure using appropriate monitoring and logging tools.
  • Implement and maintain Infrastructure as Code (IaC) using tools such as Terraform or CloudFormation.
  • Automate repetitive operational tasks using Python, Bash, Shell scripting, or similar technologies.
  • Implement infrastructure and application security, access controls, patching, and hardening.
  • Investigate production incidents, perform root-cause analysis (RCA), and drive issues to resolution.
  • Take end-to-end ownership of infrastructure reliability, availability, and operational issues.
  • Collaborate with development and other engineering teams to improve deployment, scalability, and system reliability.


Mandatory Skills :

Linux & Networking | Bare Metal / Server Administration | Jenkins / CI-CD | Docker / Containers | AWS Cloud | Kubernetes | Git | Monitoring & Logging | Security | Infrastructure as Code (IaC) | Automation / Scripting | Production Troubleshooting & Ownership


Preferred Skills :

  • Strong understanding of TCP/IP, DNS, HTTP/HTTPS, SSH, load balancing, and networking fundamentals.
  • Hands-on experience with Terraform / CloudFormation.
  • Experience with Prometheus, Grafana, ELK / EFK, CloudWatch, or similar monitoring / logging tools.
  • Good understanding of Linux performance troubleshooting, processes, memory, disk, networking, and file systems.
  • Experience handling production incidents, RCA, deployments, and system reliability.
  • Exposure to cloud security, IAM, secrets management, and server hardening.


What We’re Looking For :

  • 3+ years of hands-on experience in DevOps / SRE / Infrastructure Engineering.
  • Strong practical knowledge rather than certification-based/theoretical understanding.
  • Good troubleshooting and analytical skills.
  • Strong sense of ownership and accountability for production systems.
  • Comfortable working in a 6-day work-from-office environment.
Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via Unique Occupational by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
Gurugram
4 - 10 yrs
₹4L - ₹10L / yr
DevOps
Site Reliability Engineer (SRE)
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
+14 more

🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience Level : 4+ Years

Location : Gurugram Sector 48, Haryana (On-site)

Employment Type : Full Time Opportunity


About the Role :

We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.

In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).


Mandatory Skills :

AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux


Key Responsibilities :

1. Cloud Infrastructure & Infrastructure as Code (IaC) :

  • Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
  • Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
  • Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.

2. CI/CD, Automation & Development :

  • Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
  • Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.

3. Containerization & Orchestration :

  • Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
  • Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.

4. Observability, SRE & Incident Management :

  • Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
  • Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
  • Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
  • Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).


Required Qualifications & Skills :

  • Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
  • Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
  • Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
  • CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
  • Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
  • Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
  • Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
  • Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
  • Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via Unique Occupational by Mantasha Naaz
Bengaluru (Bangalore)
0.6 - 2 yrs
₹8L - ₹10L / yr
skill iconAmazon Web Services (AWS)
Terraform
On call Support
Incident management
Reliability engineering

Jr Platform Engineer

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 0.6-2 years

Shift Timing: 2 PM to 11 PM IST


About Company

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally,It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

At this company, we value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.


Role Overview

It is seeking a  TechOps Engineer! We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. It engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.


What you’ll do:  

  • Ensure system reliability, uptime, and performance of global platform.
  • Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule) 
  • Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization. 
  •  Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
  • End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
  • Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
  •  Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
  • Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
  • Collaborate with development, operations and support  teams to build scalable and resilient systems.
  • Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
  • Participate in capacity planning, performance tuning, and resource optimization.
  • Integrate security and compliance best practices into all infrastructure operations.
  • Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
  • Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
  • Flexible to resolve blocking issues during off hours or weekends if required.  

 

What We’re Looking For: 

Basic Qualifications and skills

  • Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
  • Overall 1+ years of experience as a Site Reliability Engineer, DevOps/ Technical project coordinator role.
  • Proven experience of DevOps, SRE, or Technical Project Coordination with IoT or connected devices based platforms.
  • Hands-on experience with cloud platforms such as AWS and Infrastructure as Code (IaC) tools such as Terraform.  
  • Experience with incident management and on-call best practices. Provide support to on call engineers.
  • Excellent analytical and problem-solving skills with a proactive mindset. 
  • Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
  • Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
  • Fluency in English (spoken and written). 
  • Successfully recommission or decommission chargers following changes in our network.
  • Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.

 Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.  


What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
Gurugram
5 - 10 yrs
₹12L - ₹18L / yr
DevOps
Reliability engineering
skill iconAmazon Web Services (AWS)
Terraform
Ansible
+18 more

Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience : 5+ Years

Location : Gurugram, Haryana

Work Mode : On-site (Full-time)


About the Role :

We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.


Mandatory Skills :

AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux


Key Responsibilities :

  • Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
  • Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
  • Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
  • Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
  • Automate operational tasks using Python, Bash, and Groovy scripting.
  • Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.


Required Skills & Qualifications :

  • Bachelor's degree in Computer Science, IT, Electronics, or a related field.
  • 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
  • Strong expertise in AWS, with exposure to Azure/GCP.
  • Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
  • Strong scripting skills in Python and Bash.
  • Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Good understanding of Linux, networking, SQL, and cloud security best practices.


Preferred Skills :

  • Experience with multi-cloud environments and DevSecOps practices.
  • Knowledge of disaster recovery, automation, and microservices architecture.
  • Strong troubleshooting, communication, and problem-solving skills.
Read more
Bengaluru (Bangalore), Chennai, Mumbai, Hyderabad, Pune, Gurugram
3 - 10 yrs
₹12L - ₹35L / yr
Linux/Unix
skill iconKubernetes
Monitoring
skill iconDocker
skill iconAmazon Web Services (AWS)
+4 more



We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.



Key Responsibilities

  • Own reliability, availability, and performance of production services, including on-call rotation
  • Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
  • Automate deployments, scaling, and operational tasks to reduce manual toil
  • Manage containerized workloads on Kubernetes and cloud infrastructure
  • Design and maintain CI/CD pipelines for safe, frequent releases
  • Lead incident response and drive blameless postmortems with clear follow-ups
  • Perform capacity planning, performance tuning, and cost optimization
  • Define and track SLIs/SLOs and error budgets with product teams


Requirements

  • 3+ years in SRE, DevOps, or production-focused engineering
  • Strong Linux administration and hands-on Kubernetes experience
  • Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
  • Cloud experience with AWS, GCP, or Azure
  • CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
  • Proficient scripting in Python and/or Bash


Nice to have

  • Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
  • Background in high-traffic or distributed systems
Read more
EDM NETWORK
AHLOUCHE AHLOUCHE
Posted by AHLOUCHE AHLOUCHE
Remote only
1 - 10 yrs
$10K - $20K / yr (ESOP available)
User Experience (UX) Design

Key Responsibilities

Platform Monitoring and Reliability

  • Monitor the health and performance of EDM’s advertiser, publisher, broker, and internal platforms.
  • Maintain monitoring, alerting, logging, and system-health dashboards.
  • Investigate platform outages, degraded performance, failed transactions, delayed data, and integration errors.
  • Respond to production incidents and coordinate resolutions with the appropriate engineers and vendors.
  • Perform root-cause analysis and document corrective and preventive actions.
  • Help maintain defined uptime, response-time, recovery-time, and system-reliability targets.
  • Identify recurring problems and recommend permanent solutions.

Cloud Infrastructure and Systems Operations

  • Maintain and support cloud infrastructure, servers, databases, networks, storage, and production environments.
  • Support development, staging, and production environments.
  • Assist with infrastructure scaling, system upgrades, patching, backups, and disaster recovery.
  • Monitor cloud usage and help control infrastructure and technology costs.
  • Maintain access controls, service accounts, certificates, domain configurations, and environment variables.
  • Ensure production systems are properly documented and recoverable.

Deployment and Release Support

  • Support safe and consistent application deployments.
  • Maintain or improve continuous integration and deployment workflows.
  • Coordinate release schedules, deployment validation, rollback procedures, and post-release monitoring.
  • Help engineering teams identify configuration or infrastructure problems before releases reach production.
  • Maintain deployment documentation, technical checklists, and change logs.
  • Reduce manual deployment work through automation.

API and Integration Support

  • Monitor and troubleshoot third-party APIs, webhooks, postbacks, dialer connections, CRM integrations, payment systems, tracking platforms, and compliance services.
  • Investigate failed lead deliveries, missing postbacks, duplicate records, delayed reporting, and authentication problems.
  • Support ping-post, real-time bidding, call-routing, SIP, and data-transfer workflows.
  • Work with advertisers, publishers, and vendors to diagnose technical integration problems.
  • Create clear documentation for common integration methods and troubleshooting procedures.
  • Develop alerts that identify integration failures before clients report them.

Call and Lead Operations

  • Monitor call-routing, tracking, recording, attribution, and disposition systems.
  • Investigate calls that fail to route, connect, record, track, or report correctly.
  • Support number provisioning, routing rules, caps, schedules, geographic restrictions, buyer availability, and failover logic.
  • Validate that leads, calls, and transactions are properly attributed to the correct advertiser, publisher, campaign, and payout.
  • Assist with discrepancies involving call duration, billable events, conversions, payouts, and reporting.
  • Help protect revenue by identifying technical leakage and delivery failures.

Data and Reporting Support

  • Monitor data pipelines, scheduled jobs, reporting processes, and database performance.
  • Investigate discrepancies between platform reporting, billing records, payment records, and third-party systems.
  • Write and maintain database queries for troubleshooting, validation, and operational reporting.
  • Assist with data corrections using controlled and documented procedures.
  • Support dashboards and operational alerts for revenue, margin, consumption, conversion, and platform activity.
  • Maintain appropriate controls around production data access and modification.

Security and Access Management

  • Support role-based access controls, multifactor authentication, audit logging, encryption, and secure system configuration.
  • Provision and remove employee, contractor, client, and vendor access.
  • Monitor suspicious activity and report potential security incidents.
  • Assist with vulnerability remediation, security reviews, access audits, and incident-response procedures.
  • Protect consumer, advertiser, publisher, employee, and company information.
  • Follow company policies for credentials, production access, sensitive data, and change management.

Automation and Process Improvement

  • Automate repetitive operational tasks using scripts, workflows, APIs, and infrastructure tools.
  • Reduce manual work associated with monitoring, deployments, reporting, reconciliation, onboarding, and support.
  • Build internal tools that improve visibility and response times.
  • Identify operational bottlenecks that affect revenue, margin, client satisfaction, or employee productivity.
  • Maintain clear runbooks and standard operating procedures for recurring technical tasks.

Technical Support and Documentation

  • Serve as an escalation point for complex platform and integration issues.
  • Translate technical problems into clear explanations for nontechnical teams.
  • Create and maintain architecture diagrams, system inventories, runbooks, troubleshooting guides, and incident reports.
  • Track incidents and technical requests through completion.
  • Document known issues, temporary workarounds, permanent resolutions, and system dependencies.
  • Participate in an on-call rotation for urgent production incidents.

First 90-Day Priorities

The successful candidate will be expected to:

  • Learn EDM’s platforms, infrastructure, integrations, call-routing systems, reporting processes, and revenue workflows.
  • Document critical systems, dependencies, credentials ownership, vendor contacts, and escalation procedures.
  • Review existing monitoring, alerting, backups, access controls, and deployment procedures.
  • Establish baseline metrics for uptime, incident volume, response time, recovery time, deployment success, and integration failures.
  • Resolve high-priority recurring production and integration issues.
  • Improve alerting for call-routing failures, API errors, delayed data, failed jobs, and reporting discrepancies.
  • Create runbooks for the company’s most common and highest-risk technical incidents.
  • Identify at least three meaningful automation or cost-saving opportunities.
  • Participate in production support and demonstrate ownership of incidents through resolution.

Performance Expectations

Success will be measured by:

  • Platform uptime and reliability
  • Mean time to acknowledge and resolve incidents
  • Reduction in recurring production problems
  • Deployment success and rollback rates
  • API, postback, webhook, and call-routing reliability
  • Reporting and data accuracy
  • Backup and recovery readiness
  • Quality of technical documentation
  • Security and access-control compliance
  • Reduction in manual operational work
  • Infrastructure costs relative to platform volume
  • Responsiveness to internal teams, clients, and technical partners

Required Qualifications

  • Three or more years of experience in operations engineering, DevOps, site reliability engineering, cloud infrastructure, systems administration, or production support.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Experience supporting Linux-based production environments.
  • Working knowledge of networking, DNS, SSL certificates, firewalls, load balancing, and application security.
  • Experience with relational databases and SQL.
  • Experience troubleshooting REST APIs, webhooks, authentication, and third-party integrations.
  • Familiarity with monitoring, logging, alerting, and incident-management tools.
  • Experience with scripting languages such as Python, Bash, JavaScript, or PowerShell.
  • Understanding of source control, deployment pipelines, and release management.
  • Strong troubleshooting, documentation, and communication skills.
  • Ability to prioritize incidents based on business and revenue impact.
  • Availability to participate in an on-call rotation.

Preferred Qualifications

  • Experience in ad-tech, mar-tech, affiliate marketing, lead generation, pay-per-call, telecommunications, or SaaS.
  • Familiarity with SIP, VoIP, dialers, call-tracking platforms, routing systems, and phone-number provisioning.
  • Experience with containers, infrastructure as code, and automated deployment tools.
  • Experience with Docker, Kubernetes, Terraform, GitHub Actions, or similar technologies.
  • Familiarity with payment processing, usage-based billing, reconciliation, and commission systems.
  • Experience working with real-time bidding, ping-post, lead distribution, or high-volume transactional systems.
  • Understanding of TCPA-related controls, consent records, DNC suppression, data privacy, or regulated marketing environments.
  • Experience with security audits, disaster-recovery testing, and compliance documentation.

Ideal Candidate

The ideal candidate:

  • Takes ownership instead of waiting for someone else to fix the problem.
  • Remains calm and methodical during high-impact incidents.
  • Understands the difference between applying a temporary fix and eliminating a root cause.
  • Communicates technical problems clearly and directly.
  • Recognizes that production reliability is a business and revenue responsibility.
  • Automates repetitive work whenever practical.
  • Documents systems so the company is not dependent on one person’s memory.
  • Protects security and stability without creating unnecessary bureaucracy.
  • Is comfortable working in a fast-moving entrepreneurial environment.
  • Can manage competing priorities while maintaining attention to detail.


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos