Platform Engineer at TestMu AI (Formely LambdaTest) · Delhi, Gurugram, Noida, Ghaziabad, Faridabad · 2 - 4 years · ₹15L - ₹20L / yr · Raised funding · Posted 31 Jul 2026

PLATFORM ENGINEER
Engineering · Platform
📍 Noida 🕐 Full-Time 🧭 2-3 Years
THE MISSION
We aren't automating scripts — we're deprecating the era of manual-heavy testing entirely. TestMu AI is building the world's first AI-native platform where Agentic Intelligence autonomously plans, authors, and self-heals the entire Quality Engineering lifecycle.
Drive reliability and scalability of TestMu AI's core infrastructure and service workflows.
THE PILLARS OF IMPACT
🚀 1. Platform Reliability (50%)
– Ensure robust cloud infrastructure and deployment pipelines.
– Improve observability tools and practices.
– Automate and streamline service workflows.
– Enhance platform systems for developer efficiency.
⚙️ 2. Infrastructure Engineering (30%)
– Build and maintain cloud-based systems.
– Optimize performance and scalability.
– Implement infrastructure improvements based on feedback.
🧠 3. Incident Management (20%)
– Lead incident response with strong SRE principles.
– Analyze and resolve production issues swiftly.
– Develop strategies to prevent future incidents.
MUST-HAVES — DO NOT APPLY UNLESS YOU HAVE THESE
– You must have strong cloud, scripting, and core DevOps fundamentals.
– You must have hands-on experience with observability tools like New Relic, Sumo Logic.
– You must demonstrate real coding ability.
– You must have a strong grasp of incident management and reliability/SRE thinking.
– You must explain service-layer architecture and reasoning effectively.
THE BAR — WHAT YOU MUST PROVE
Cloud Expertise: Designed and optimized cloud infrastructure for scalability and reliability.
Coding Proficiency: Developed scripts and tools to automate deployment processes.
Incident Management: Led a team to resolve critical production incidents effectively.
Architecture Understanding: Explained complex service architectures to non-technical stakeholders.

About TestMu AI (Formely LambdaTest)
About
TestMu AI (formerly LambdaTest) is the world’s first full-stack Agentic AI Quality Engineering Platform.
We built TestMu AI for a reality where software is written by AI and must be shipped at machine speed.
Tech stack
Connect with the team
Similar jobs (10)
Senior Platform & Site Reliability Engineer
Location: Remote Employment Type: Contract
The Role
This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.
Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.
This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.
The Scale You Will Operate At
The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.
You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.
What You Will Own
Platform Architecture
- Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
- Define, build, and enforce platform standards across portfolio products
- Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
- Self-service developer platform so product teams ship without waiting on platform
Event Streaming & Pipeline Infrastructure
- Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
- Design and maintain batch processing infrastructure alongside live event flows
- Ensure pipeline reliability, throughput, and cost are actively managed at scale
CI/CD & Deployment
- Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
- Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
- Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified
Observability & Reliability
- Own the full observability stack: Grafana, Prometheus, and Loki across all products
- SLOs and error budgets defined per product; reliability tracked consistently
- Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
- Incident response: on-call design, escalation playbooks, post-mortem facilitation
- Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached
Acquisition Onboarding
- Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
- Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
- Target: full platform integration within a defined window per acquisition
A Note on Automation
Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.
Platform Stack
Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For
- 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
- Strong Terraform depth across multi-environment, multi-account setups
- CI/CD ownership across a multi-product environment with GitHub Actions
- Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
- Hands-on Grafana, Prometheus, and Loki in production
- AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
- Acquisition or greenfield platform integration experience strongly preferred
How You Work
- Comfortable operating across multiple products simultaneously — context-switching without dropping standards
- Cost-efficiency instinct — you optimise spend as a habit, not as a project
- You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
- You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use
Why This Role
The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.
This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.
About the Team
SecurITe’s mission is to build an Agentic-AI driven security platform that protects critical infrastructure from modern cyber threats. Our focus is on delivering highly performant, resilient, and intelligent network security systems that help defenders stay ahead of adversaries.
About the Role
We’re looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.
As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premise environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.
You’ll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security.
This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.
Why This Role Matters
Cybersecurity is undergoing a fundamental shift. AI is no longer an enhancement—it’s becoming the core engine of how detection, investigation, and response are executed. As our Platform Engineer, you will architect and build the infrastructure, automation, deployment, and operational systems that make this transformation real.
Your work will directly influence the scalability, reliability, and security of our AI-driven cybersecurity platform across both cloud and enterprise on-premise deployments. You’ll help establish the operational backbone that enables rapid innovation, secure product delivery, and resilient large-scale deployments in mission-critical environments.
This is a chance to solve novel technical challenges involving distributed systems, hybrid infrastructure, observability, automation, and secure software delivery while shaping how defenders outpace modern attackers.
What You’ll Do
● Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, security groups)
● Administer and harden AlmaLinux VMs across production, staging, and dev environments
● Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling)
● Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch
● Lead incident response — diagnosis, RCA, and post-incident documentation — with no dedicated ops team to escalate to
● Make and document build-vs-buy and architecture decisions as the product and team scale
● Work directly with founders/engineering to translate ambiguous asks into scoped technical plans
Impact You’ll Have
● Accelerate engineering velocity through scalable developer platforms and automation
● Improve deployment reliability, platform uptime, and operational efficiency
● Enable secure and scalable AI-driven cybersecurity workloads
● Reduce operational overhead through infrastructure automation and self-service systems
● Help establish enterprise-grade cloud and on-premise deployment capabilities
● Enhance product resiliency, observability, and operational excellence
● Shape the long-term platform architecture powering next-generation cybersecurity products
● Enable rapid and secure delivery of critical security innovations to customers
Required Experience
● 4+ years hands-on Linux administration (RHEL-family strongly preferred — AlmaLinux, CentOS, RHEL)
● Deep Linux internals: systemd, networking, storage/LVM, process/resource management, kernel-level troubleshooting
● Real AWS architecture experience — not just operating existing infra, but designing it (VPC, EC2, IAM, security groups, networking)
● Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to
● Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts — built tooling that runs unattended
● Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one)
● Clear, proactive communicator — documents decisions and explains reasoning without being asked
Required Skills & Qualifications
● Strong Linux system administration and troubleshooting skills
● Redhat certifications
● Strong understanding of networking fundamentals, security, and distributed systems
● Proficiency with Docker, and container orchestration
● Experience with Terraform, Ansible, or similar infrastructure automation tools
● Strong scripting or programming skills in Python, Bash, or Go
● Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry
● Understanding of platform security best practices and secure infrastructure design
● Familiarity with virtualization technologies and hybrid infrastructure environments
● Strong problem-solving and debugging abilities
● Excellent communication and collaboration skills
● Ability to thrive in fast-paced startup environments
Nice to Have
● Configuration management/automation at scale (Ansible, AWX, Terraform, or similar)
● Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog)
● Container experience (Docker; Kubernetes a plus but not core to this VM-based stack)
● Experience in a startup or small-team environment where infra was built from scratch
● Security/compliance exposure (vulnerability remediation, hardening, SSO/access control)
The Mindset
Problem Solver
You thrive on complex, ambiguous challenges and engineer elegant solutions.
Ownership-Driven
You take initiative, move fast, and deliver outcomes without hand-holding.
Continuous Learner
You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.
Startup DNA
You excel in fast-moving environments where priorities evolve and impact is immediate.
We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.
Key Responsibilities
- Own reliability, availability, and performance of production services, including on-call rotation
- Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
- Automate deployments, scaling, and operational tasks to reduce manual toil
- Manage containerized workloads on Kubernetes and cloud infrastructure
- Design and maintain CI/CD pipelines for safe, frequent releases
- Lead incident response and drive blameless postmortems with clear follow-ups
- Perform capacity planning, performance tuning, and cost optimization
- Define and track SLIs/SLOs and error budgets with product teams
Requirements
- 3+ years in SRE, DevOps, or production-focused engineering
- Strong Linux administration and hands-on Kubernetes experience
- Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
- Cloud experience with AWS, GCP, or Azure
- CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
- Proficient scripting in Python and/or Bash
Nice to have
- Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
- Background in high-traffic or distributed systems

Platform Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 2-4 years
About Compnay
This is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.
Operating over 70,000 charge points globally, this is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.
Role Overview
What you’ll do:
- Ensure system reliability, uptime, and performance of global platform.
- Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule).
- Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
- Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
- Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
- Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
- Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
- Collaborate with development, operations and support teams to build scalable and resilient systems.
- Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
- Participate in capacity planning, performance tuning, and resource optimization.
- Integrate security and compliance best practices into all infrastructure operations.
- Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
- Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
- Flexible to resolve blocking issues during off hours or weekends if required.
What We’re Looking For:
Basic Qualifications and Skills
- Bachelor’s degree in Engineering, Electrical, ECE, Computer Science, Information Technology, or related field.
- 2–4 years of overall experience with at least 1+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Technical Project Coordinator.
- Proven experience of DevOps, SRE or Technical Project Coordination with IoT or connected devices-based platforms.
- Experience with incident management and on-call best practices. Provide support to on-call engineers.
- Excellent analytical and problem-solving skills with a proactive mindset.
- Expertise with monitoring and observability tools (Dynatrace, Prometheus, Grafana, Zabbix, etc.).
- Solid understanding of cloud platforms (AWS) and AWS native services (EKS, EC2, S3, RDS, Lambda).
- Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
- Fluency in English (spoken and written).
- Successfully recommission or decommission chargers following changes in our network.
- Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.
Additional Information
- This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions.
- Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support on-call engineers.
- This role may involve EU or US time-zone shifts based on business requirements.
- Shift timing: 2 PM IST to 11 PM IST.
What is required to be successful in this role:
- Global platform experience (B2C or B2B).
- AWS native service experience.
- Firmware deployment and cloud cost optimization experience.
- Strong exposure to monitoring and alerts.
- Experience with firmware rollout, IoT devices onboarding and offboarding will be an added advantage.
- Experience as an SRE or DevOps Engineer with some exposure to Project Management or Technical Project Management in IoT-based projects will be helpful.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
Hiring Platform Engineer
Exp: 6 -- 10 yrs
Edu : BE/B.tech/MCA
Work Location : Pune
Skills :
Platform monitoring ,Incident trouble shooting, Incident recovery, openshift ,kubernetes.
2 years of IT operations, infrastructure, cloud or application support experience.
Exp in Linux command-line knowledge.
Exp in networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.
Site Reliability Engineer (SRE) / DevOps Engineer (Walk-In Drive)
Location: Gurgaon Experience: 3–6 Years
About the Role
We are looking for a hands-on Site Reliability Engineer (SRE) / DevOps Engineer with strong programming and automation skills.
The role will initially involve development and automation work, helping the engineer build a strong understanding of the applications and platform. Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, observability, production reliability, incident management, and operational automation.
The ideal candidate should be comfortable working with both application code and production systems and should use engineering and automation to improve reliability and reduce manual effort.
Key Responsibilities
· Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.
· Write clean, maintainable, testable, and production-ready code.
· Participate in code reviews, debugging, testing, and technical discussions.
· Build, maintain, and improve CI/CD pipelines and automated deployment processes.
· Work with Docker and Kubernetes for application deployment and operations.
· Support on prem and cloud-based application and infrastructure deployments.
· Maintain reliable, scalable, secure, and highly available production environments.
· Implement and manage monitoring, logging, alerting, and observability solutions.
· Contribute to defining and tracking SLIs, SLOs, and error budgets.
· Troubleshoot application and production issues and perform Root Cause Analysis (RCA).
· Identify recurring operational problems and address them through automation and engineering improvements.
· Support incident response, change management, deployment governance, and disaster recovery practices.
· Maintain runbooks, SOPs, incident documentation, and technical documentation.
· Collaborate with Engineering, Product, Platform, Security, Operations, and external teams.
Technical Skills
· Strong hands-on experience with Python for development and automation.
· Experience developing scripts, APIs, integrations, utilities, or backend services.
· Good understanding of software engineering principles, debugging, logging, testing, and exception handling.
· Experience with REST APIs, JSON, Git, pull requests, and code reviews.
· Strong knowledge of Linux/Unix environments and basic Windows administration.
· Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.
· Experience with at least one cloud platform: AWS, Azure, or GCP.
· Hands-on experience with Docker and Kubernetes.
· Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.
· Familiarity with Infrastructure-as-Code tools such as Terraform is preferred.
· Experience with monitoring and observability tools such as Grafana, Prometheus, Power BI, or equivalent.
· Ability to analyze logs, metrics, alerts, and traces for troubleshooting.
· Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.
· Experience with JIRA, ServiceNow, and Confluence is desirable.
Preferred Experience
· 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, or related roles.
· Strong Python development or automation experience.
· Experience supporting applications across development, deployment, and production environments.
· Exposure to cloud-native, distributed, or production-grade systems.
· Understanding of security and compliance best practices.
· Familiarity with AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools.
Soft Skills
· Strong analytical and troubleshooting skills.
· Engineering and automation mindset.
· Good written and verbal communication skills.
· Effective cross-functional collaboration.
· Ownership-driven approach to problem solving.
· Ability to remain structured during production incidents.
Job Summary:
We are looking for a Senior SRE/DevOps Engineer with strong experience in site reliability, automation, monitoring, observability, and production support. The candidate will be responsible for ensuring the reliability, availability, security, and performance of enterprise platforms.
Key Responsibilities:
- Own reliability, availability, security, and performance of enterprise browser platforms.
- Manage access and identity controls, Group Policy, and SAML/SSO integrations.
- Handle CI/CD deployments, configuration management, monitoring, and health checks.
- Develop automation and operational workflows using Python and Bash.
- Perform performance monitoring and production troubleshooting.
- Implement and maintain observability frameworks.
- Work with monitoring tools such as Datadog, Splunk, Dynatrace, Prometheus, and Grafana.
- Participate in incident management, RCA, and continuous improvement activities.
- Support highly available production environments in a 24/7 shift model.
- Collaborate with application, infrastructure, security, and operations teams.
Mandatory Skills:
- 10+ years of experience in SRE / DevOps.
- Strong hands-on experience in Python and Bash scripting.
- Experience with CI/CD and configuration management.
- Strong knowledge of monitoring and observability.
- Hands-on experience with Datadog, Splunk, Dynatrace, Prometheus, or Grafana.
- Experience with SAML/SSO, Identity & Access Management, and Group Policy.
- Banking domain experience.
🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience Level : 4+ Years
Location : Gurugram Sector 48, Haryana (On-site)
Employment Type : Full Time Opportunity
About the Role :
We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.
In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).
Mandatory Skills :
AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux
Key Responsibilities :
1. Cloud Infrastructure & Infrastructure as Code (IaC) :
- Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
- Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
- Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.
2. CI/CD, Automation & Development :
- Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
- Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.
3. Containerization & Orchestration :
- Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
- Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.
4. Observability, SRE & Incident Management :
- Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
- Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
- Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
- Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).
Required Qualifications & Skills :
- Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
- Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
- Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
- CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
- Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
- Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
- Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
- Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
- Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.

Key Skills:
• Bachelor's or Master's degree in Computer Science or related field.
• Minimum 5 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure Engineering.
• Experience migrating data and systems between AWS IaaS and PaaS.
• Experience operating and supporting applications using AWS VPC, EKS, and related services for multi-account operations.
• Experience developing fast and reliable Continuous Integration/Continuous Deployment (CI/CD) workflows used by hundreds of application teams.
• Experience administering and troubleshooting Operating Systems such as Linux, Windows, and MacOS.
• Professional Certifications in AWS Networks, CNCF Technologies, or Kubernetes.
• Experience using and configuring observability tools such as ELK, Prometheus/Grafana, AWS CloudWatch, and Jaeger.
• Experience of applied GitOps principles using ArgoCD or Flux.
• Public examples of code you've worked on with other people using any of these technologies:
o Configuration management/Infrastructure as Code (IAC) tools, such as AWS CDK, AWS CloudFormation, Terraform, Ansible, or Puppet.
o Systems solutions in one or more programming languages, such as Golang, Python, Java.
o Build, Release, Deploy or Ops Workflows using Bamboo, Argo Project, or GitHub Actions.

What You Will Own
Platform Architecture
- Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
- Define, build, and enforce platform standards across portfolio products
- Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
- Self-service developer platform so product teams ship without waiting on platform
Event Streaming & Pipeline Infrastructure
- Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
- Design and maintain batch processing infrastructure alongside live event flows
- Ensure pipeline reliability, throughput, and cost are actively managed at scale
CI/CD & Deployment
- Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
- Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
- Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified
Observability & Reliability
- Own the full observability stack: Grafana, Prometheus, and Loki across all products
- SLOs and error budgets defined per product; reliability tracked consistently
- Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
- Incident response: on-call design, escalation playbooks, post-mortem facilitation
- Automated remediation scoped to a defined set of safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns. Novel or ambiguous failures escalate to a human with full context attached
Acquisition Onboarding
- Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
- Migration plan and execution for each portfolio company joining the platform.
- Target: full platform integration within a defined window per acquisition
What We're Looking For
Experience & Background
- 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
- Strong Terraform depth across multi-environment, multi-account setups
- CI/CD ownership across a multi-product environment with GitHub Actions
- Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
- Hands-on Grafana, Prometheus, and Loki in production
- AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
- Acquisition or greenfield platform integration experience strongly preferred






