SRE Engineer at Technology Industry · Hyderabad · 7 - 10 years · ₹20L - ₹40L / yr · Posted 27 Feb 2026

Description
SRE Engineer
Role Overview
As a Site Reliability Engineer, you will play a critical role in ensuring the availability and performance of our customer-facing platform. You will work closely with DevOps, DBA, and Development teams to provision and maintain infrastructure, deploy and monitor our applications, and automate workflows. Your contributions will have a direct impact on customer satisfaction and overall experience.
Responsibilities and Deliverables
• Manage, monitor, and maintain highly available systems (Windows and Linux)
• Analyze metrics and trends to ensure rapid scalability.
• Address routine service requests while identifying ways to automate and simplify.
• Create infrastructure as code using Terraform, ARM Templates, Cloud Formation.
• Maintain data backups and disaster recovery plans.
• Design and deploy CI/CD pipelines using GitHub Actions, Octopus, Ansible, Jenkins, Azure DevOps.
• Adhere to security best practices through all stages of the software development lifecycle
• Follow and champion ITIL best practices and standards.
• Become a resource for emerging and existing cloud technologies with a focus on AWS.
Organizational Alignment
• Reports to the Senior SRE Manager
• This role involves close collaboration with DevOps, DBA, and security teams.
Technical Proficiencies
• Hands-on experience with AWS is a must-have.
• Proficiency analyzing application, IIS, system, security logs and CloudTrail events
• Practical experience with CI/CD tools such as GitHub Actions, Jenkins, Octopus
• Experience with observability tools such as New Relic, Application Insights, AppDynamics, or DataDog.
• Experience maintaining and administering Windows, Linux, and Kubernetes.
• Experience in automation using scripting languages such as Bash, PowerShell, or Python.
• Configuration management experience using Ansible, Terraform, Azure Automation Run book or similar.
• Experience with SQL Server database maintenance and administration is preferred.
• Good Understanding of networking (VNET, subnet, private link, VNET peering).
• Familiarity with cloud concepts including certificates, Oauth, AzureAD, ASE, ASP, AKS, Azure Apps,
Load Balancers, Application Gateway, Firewall, Load Balancer, API Management, SQL Server, Databases on Azure
Experience
• 7+ years of experience in SRE or System Administration role
• Demonstrated ability building and supporting high availability Windows/Linux servers, with emphasis on the WISA stack (Windows/IIS/SQL Server/ASP.net)
• 3+ years of experience working with cloud technologies including AWS, Azure.
• 1+ years of experience working with container technology including Docker and Kubernetes.
• Comfortable using Scrum, Kanban, or Lean methodologies.
Education
• Bachelor’s Degree or College Diploma in Computer Science, Information Systems, or equivalent
experience.
Additional Job Details:
• Working hours: 2:00 PM / 3:00 PM to 11:30 PM IST
• Interview process: 3 technical rounds
• Work model: 3 days’ work from office

Similar jobs (10)
Senior Platform & Site Reliability Engineer
Location: Remote Employment Type: Contract
The Role
This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.
Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.
This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.
The Scale You Will Operate At
The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.
You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.
What You Will Own
Platform Architecture
- Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
- Define, build, and enforce platform standards across portfolio products
- Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
- Self-service developer platform so product teams ship without waiting on platform
Event Streaming & Pipeline Infrastructure
- Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
- Design and maintain batch processing infrastructure alongside live event flows
- Ensure pipeline reliability, throughput, and cost are actively managed at scale
CI/CD & Deployment
- Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
- Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
- Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified
Observability & Reliability
- Own the full observability stack: Grafana, Prometheus, and Loki across all products
- SLOs and error budgets defined per product; reliability tracked consistently
- Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
- Incident response: on-call design, escalation playbooks, post-mortem facilitation
- Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached
Acquisition Onboarding
- Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
- Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
- Target: full platform integration within a defined window per acquisition
A Note on Automation
Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.
Platform Stack
Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For
- 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
- Strong Terraform depth across multi-environment, multi-account setups
- CI/CD ownership across a multi-product environment with GitHub Actions
- Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
- Hands-on Grafana, Prometheus, and Loki in production
- AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
- Acquisition or greenfield platform integration experience strongly preferred
How You Work
- Comfortable operating across multiple products simultaneously — context-switching without dropping standards
- Cost-efficiency instinct — you optimise spend as a habit, not as a project
- You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
- You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use
Why This Role
The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.
This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.
SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.
Core responsibilities include:
- Production monitoring and debugging of live systems.
- Incident investigation, troubleshooting, and problem resolution.
- AWS cloud infrastructure support and maintenance.
- Deployment and operational support activities.
- Supporting a 24x7 production environment.
- Working with GitHub-based development workflows.
- Technical debt remediation and platform improvements.
- Customer issue investigation and support.
- Security and compliance-related work, including FedRAMP initiatives.
Preferred Skills:
AWS (especially S3 and EC2)
Strong debugging and troubleshooting skills
Site Reliability Engineering (SRE) experience
GitHub experience
Basic software development skills
TypeScript/JavaScript knowledge
C# preferred
AI experience is a plus.
Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.
Job Title: Senior Site Reliability Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 6+ years
About Compnay
It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.
Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.
We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.
Role Overview
We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.
Key Responsibilities
- Ensure high availability, performance, and reliability of production systems
- Design, implement, and manage scalable infrastructure solutions
- Build and maintain CI/CD pipelines for efficient software delivery
- Monitor system health using observability tools and respond to incidents proactively
- Automate operational processes using scripting and Infrastructure as Code (IaC)
- Manage containerized environments using Docker and Kubernetes
- Collaborate with cross-functional teams to improve system architecture and resilience
- Participate in on-call rotations and incident management processes
- Continuously optimize cloud infrastructure for cost, performance, and scalability
Required Qualifications & Skills
- Bachelor’s degree in Computer Science, IT, or related field
- 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
- Strong experience with containerization (Docker) and orchestration (Kubernetes)
- Proficiency in Linux administration, networking, and system security
- Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
- Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
- Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
- Proficiency in scripting languages (Python, Bash, or PowerShell)
- Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
- Solid understanding of system architecture, microservices, and SaaS/PaaS models
- Strong analytical and problem-solving skills
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
Job Summary
We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in networking, DNS, load balancing, and hybrid cloud environments. The candidate will be responsible for maintaining the reliability, availability, performance, and scalability of production infrastructure and services.
The ideal candidate should have strong troubleshooting skills and experience working across network, cloud, infrastructure, and application environments.
Key Responsibilities
- Monitor and maintain the availability and reliability of production systems and services.
- Troubleshoot complex network, infrastructure, and application connectivity issues.
- Manage and troubleshoot DNS services, DNS resolution, records, and configuration issues.
- Configure, manage, and troubleshoot Load Balancers and traffic routing.
- Work with Layer 4 and Layer 7 networking and understand TCP/IP, HTTP/HTTPS, routing, and network connectivity.
- Support hybrid cloud environments involving on-premises infrastructure and public cloud platforms.
- Troubleshoot connectivity between on-premises data centers and cloud environments.
- Participate in production incidents, troubleshooting, root cause analysis (RCA), and problem management.
- Develop automation scripts and tools to reduce manual operational activities.
- Configure and maintain monitoring, alerting, and observability solutions.
- Work closely with Network, Cloud, DevOps, Security, and Application teams.
- Participate in on-call support and resolve production issues within defined SLAs.
- Document infrastructure, troubleshooting procedures, incident reports, and operational processes.
- Identify opportunities to improve system reliability, performance, and scalability.
We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.
Key Responsibilities
- Own reliability, availability, and performance of production services, including on-call rotation
- Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
- Automate deployments, scaling, and operational tasks to reduce manual toil
- Manage containerized workloads on Kubernetes and cloud infrastructure
- Design and maintain CI/CD pipelines for safe, frequent releases
- Lead incident response and drive blameless postmortems with clear follow-ups
- Perform capacity planning, performance tuning, and cost optimization
- Define and track SLIs/SLOs and error budgets with product teams
Requirements
- 3+ years in SRE, DevOps, or production-focused engineering
- Strong Linux administration and hands-on Kubernetes experience
- Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
- Cloud experience with AWS, GCP, or Azure
- CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
- Proficient scripting in Python and/or Bash
Nice to have
- Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
- Background in high-traffic or distributed systems
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
As a DevOps Engineer at YOYO, you'll own the infrastructure and delivery backbone that keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and observability that let a small, fast-moving team ship confidently and you'll keep our AI and data workloads reliable and affordable at scale. This is a hands-on role with real ownership: you won't be maintaining someone else's setup, you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make deployment boring, incidents rare, and scaling a non-event.
If you are Interested DM me on LinkedIn - Saquib Mundagnur
What You'll Own
CI/CD & developer experience - Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible automated testing, rollbacks, and release controls.Cloud infrastructure & IaC - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
Containers & orchestration - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch processing for the speech pipeline.
Reliability & observability (SRE) - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
Data & pipeline infrastructure - Support the infrastructure behind large-scale, edge-to-cloud data movement and processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
Security & compliance - Bake security into the platform: secrets management, IAM/least-privilege, encryption in transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security requirements) for a product that handles sensitive customer conversations.
Cost & scale - Own cloud cost visibility and optimization; make scaling decisions that balance reliability and spend.
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production systems at meaningful scale.
- Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and Infrastructure-as-Code (**Terraform** or equivalent).
- Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong automation-first mindset.
- Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, OpenTelemetry, or similar) and running incident response / on-call.
- A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team
We are looking for a hands-on Senior AWS Cloud Engineer to lead the infrastructure build, optimization, automation, and production deployment of a Multi-Agent AI Chatbot Platform hosted on AWS. The development environment is already in place, and the successful candidate will drive the solution through testing, integrations, and production go-live.
Key Responsibilities
- Review, validate, and optimize existing Terraform code and AWS infrastructure.
- Establish and manage integrations with enterprise platforms such as ServiceNow, Workday, and other third-party systems.
- Design, build, and support secure, scalable, and highly available AWS environments.
- Implement and automate CI/CD pipelines and Infrastructure-as-Code practices.
- Lead infrastructure testing, performance tuning, and production readiness activities.
- Drive deployment and operationalization of the platform in the Production environment.
- Implement cloud governance, security, monitoring, and FinOps best practices.
- Troubleshoot and resolve complex cloud infrastructure issues.
Required Skills & Experience
- 10+ years of IT experience with strong expertise in AWS Cloud Engineering.
- Proven experience designing, deploying, and managing AWS production environments.
- Strong hands-on experience with Terraform and Infrastructure-as-Code.
- Experience with CI/CD pipeline automation and DevOps practices.
- Expertise in AWS services including VPC, IAM, EC2, S3, Lambda, CloudWatch, and networking.
- Experience in performance optimization, reliability, and cloud cost management (FinOps).
- Strong scripting and automation skills.
- Experience integrating enterprise applications through APIs and secure connectivity patterns.
Job Summary
We are looking for an experienced AWS Cloud Engineer with strong expertise in AWS infrastructure, deployment, migration, and cloud operations. The candidate should have hands-on experience managing AWS services such as EC2, EBS, S3, EFS, and FSx, along with infrastructure provisioning, migration, troubleshooting, and optimization.
Key Responsibilities
- Design, deploy, configure, and manage AWS infrastructure environments.
- Perform application and infrastructure migration to AWS.
- Provision and manage EC2 instances, including configuration, scaling, patching, and troubleshooting.
- Manage EBS volumes, snapshots, backups, and storage performance.
- Configure and administer S3 buckets, storage policies, lifecycle management, and access controls.
- Manage EFS for scalable shared file storage.
- Implement and manage Amazon FSx file systems based on application requirements.
- Monitor AWS infrastructure performance, availability, and capacity.
- Troubleshoot infrastructure, networking, storage, and deployment-related issues.
- Implement security best practices including IAM, security groups, encryption, and access controls.
- Support backup, disaster recovery, high availability, and business continuity requirements.
- Optimize AWS resources for performance, scalability, reliability, and cost.
Mandatory Skills
- Strong hands-on experience in AWS Cloud Infrastructure.
- Expertise in EC2, EBS, S3, EFS, and FSx.
- Experience in AWS deployment and migration projects.
- Strong knowledge of AWS networking concepts such as VPC, Subnets, Route Tables, Security Groups, and Load Balancers.
- Experience with IAM and AWS security best practices.
- Good knowledge of AWS monitoring and troubleshooting.
- Experience with cloud infrastructure automation using Terraform or CloudFormation is preferred.
- Strong Linux administration and troubleshooting skills.
- Good understanding of backup, disaster recovery, and high-availability concepts

Job Title: TechOps Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 6 Month-2 years (Excluding Internship)
Shift Timing: 2 PM to 11 PM IST
Role Overview
We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. Engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.
What you’ll do:
- Ensure system reliability, uptime, and performance of global platform.
- Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule)
- Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
- Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
- End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
- Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
- Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
- Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
- Collaborate with development, operations and support teams to build scalable and resilient systems.
- Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
- Participate in capacity planning, performance tuning, and resource optimization.
- Integrate security and compliance best practices into all infrastructure operations.
- Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
- Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
- Flexible to resolve blocking issues during off hours or weekends if required.
What We’re Looking For:
Basic Qualifications and skills
- Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
- Overall 1 years of experience as a Site Reliability Engineer, Technical project coordinator role.
- Proven experience of SRE or Technical Project Coordination with IoT or connected devices based platforms.
- Experience with incident management and on-call best practices. Provide support to on call engineers.
- Excellent analytical and problem-solving skills with a proactive mindset.
- Hands-on experience with AWS Cloud and IaC tools such as Terraform or Ansible.
- Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
- Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
- Fluency in English (spoken and written).
- Successfully recommission or decommission chargers following changes in our network.
- Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.
Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.





