DevOps Engineer at Leading Internet pioneer in India providing web and email se · Mumbai · 5 - 8 years · ₹10L - ₹14L / yr · Posted 24 Aug 2021
DevOps Engineer
at Leading Internet pioneer in India providing web and email se
- Provides free and subscription-based website and email services hosted and operated at data centres in Mumbai and Hyderabad.
- Serve global audience and customers through sophisticated content delivery networks.
- Operate a service infrastructure using the latest technologies for web services and a very large storage infrastructure.
- Provides virtualized infrastructure, allows seamless migration and the addition of services for scalability.
- Pioneers and earliest adopters of public cloud and NoSQL big data store - since more than a decade.
- Provide innovative internet services with work on multiple technologies like php, java, nodejs, python and c++ to scale our services as per need.
- Has Internet infrastructure peering arrangements with all the major and minor ISPs and telecom service providers.
- Have mail traffic exchange agreements with major Internet services.
Job Details :
- This job position provides competitive professional opportunity both to experienced and aspiring engineers. The company's technology and operations groups are managed by senior professionals with deep subject matter expertise.
- The company believes having an open work environment offering mentoring and learning opportunities with an informal and flexible work culture, which allows professionals to actively participate and contribute to the success of our services and business.
- You will be part of a team that keeps the business running for cloud products and services that are used 24- 7 by the company's consumers and enterprise customers around the world. You will be asked to contribute to operate, maintain and provide escalation support for the company's cloud infrastructure that powers all of cloud offerings.
Job Role :
- As a senior engineer, your role grows as you gain experience in our operations. We facilitate a hands-on learning experience after an induction program, to get you into the role as quickly as possible.
- The systems engineer role also requires candidates to research and recommend innovative and automated approaches for system administration tasks.
- The work culture allows a seamless integration with different product engineering teams. The teams work together and share responsibility to triage in complex operational situations. The candidate is expected to stay updated on best practices and help evolve processes both for resilience of services and compliance.
- You will be required to provide support for both, production and non-production environments to ensure system updates and expected service levels. You will be required to specifically handle 24/7 L2 and L3 oversight for incident responses and have an excellent understanding of the end-to-end support process from client to different support escalation levels.
- The role also requires a discipline to create, update and maintain process documents, based on operation incidents, technologies and tools used in the processes to resolve issues.
QUALIFICATION AND EXPERIENCE :
- A graduate degree or senior diploma in engineering or technology with some or all of the following:
- Knowledge and work experience with KVM, AWS (Glacier, S3, EC2), RabbitMQ, Fluentd, Syslog, Nginx is preferred
- Installation and tuning of Web Servers, PHP, Java servlets, memory-based databases for scalability and performance
- Knowledge of email related protocols such as SMTP, POP3, IMAP along with experience in maintenance and administration of MTAs such as postfix, qmail, etc will be an added advantage
- Must have knowledge on monitoring tools, trend analysis, networking technologies, security tools and troubleshooting aspects.
- Knowledge of analyzing and mitigating security related issues and threats is certainly desirable.
- Knowledge of agile development/SDLC processes and hands-on participation in planning sprints and managing daily scrum is desirable.
- Preferably, programming experience in Shell, Python, Perl or C.

Similar jobs (10)
Location : Mumbai
Big Picture (The Opportunity) :
Are you looking for an opportunity to advance your Career? & If you are able to maintain a positive attitude even when everything goes wrong, if you are detail oriented and self motivated with a passion to learn and improve your skills and knowledge, we have a perfect job for you !
What do we want from you ? (Our Expectations) :
- Zero to 2 years experience in Linux Operating System.
- Flexible working hours - able to support occasional nights, weekends, and call-ins, able to quickly adapt to a constantly changing faced paced environment. Open to travel to sites.
- An ideal person who is excited and motivated about running and supporting a production - grade critical infrastructure and looks for opportunities to improve processes with automation.
What are you required to do ? (Your Responsibilities) :
- Proactively maintain and develop all linux infrastructure technology to maintain a 24*7*365 uptime service.
- Engineering of systems administration-related various solutions for our various SAAS products and projects as well as operational needs.
- Proactively monitoring system performance and capacity planning.
- Providing technical support to customers for Applications, Operating systems, and networking.
- Will also be the first point of contact for our clients, where installations are placed on a permanent basis, for basic troubleshooting and problem solving.
- Fault finding, analysis and logging information for reporting of performance exceptions.
- Maintain best practices on managing systems and services across all environments.
Skills & Qualification Required (Add Value) :
- Graduate - Preferred to have Bachelor‘s Degree in Engineering, Computer Science or related field.
- He / She should be familiar with the installation and configuration of Linux operating systems and setup and operation of TCP/IP networking on Linux systems also familiar with Internet concepts including SMTP, IMAP, POP, HTTP, DNS, LDAP and related protocols.
- You should possess excellent communication skills .
- Knowledge of Email concepts, Helpdesk Concept , VoIP Concept , Cloud computing will be an added advantage.
Key Responsibilities
Platform Monitoring and Reliability
- Monitor the health and performance of EDM’s advertiser, publisher, broker, and internal platforms.
- Maintain monitoring, alerting, logging, and system-health dashboards.
- Investigate platform outages, degraded performance, failed transactions, delayed data, and integration errors.
- Respond to production incidents and coordinate resolutions with the appropriate engineers and vendors.
- Perform root-cause analysis and document corrective and preventive actions.
- Help maintain defined uptime, response-time, recovery-time, and system-reliability targets.
- Identify recurring problems and recommend permanent solutions.
Cloud Infrastructure and Systems Operations
- Maintain and support cloud infrastructure, servers, databases, networks, storage, and production environments.
- Support development, staging, and production environments.
- Assist with infrastructure scaling, system upgrades, patching, backups, and disaster recovery.
- Monitor cloud usage and help control infrastructure and technology costs.
- Maintain access controls, service accounts, certificates, domain configurations, and environment variables.
- Ensure production systems are properly documented and recoverable.
Deployment and Release Support
- Support safe and consistent application deployments.
- Maintain or improve continuous integration and deployment workflows.
- Coordinate release schedules, deployment validation, rollback procedures, and post-release monitoring.
- Help engineering teams identify configuration or infrastructure problems before releases reach production.
- Maintain deployment documentation, technical checklists, and change logs.
- Reduce manual deployment work through automation.
API and Integration Support
- Monitor and troubleshoot third-party APIs, webhooks, postbacks, dialer connections, CRM integrations, payment systems, tracking platforms, and compliance services.
- Investigate failed lead deliveries, missing postbacks, duplicate records, delayed reporting, and authentication problems.
- Support ping-post, real-time bidding, call-routing, SIP, and data-transfer workflows.
- Work with advertisers, publishers, and vendors to diagnose technical integration problems.
- Create clear documentation for common integration methods and troubleshooting procedures.
- Develop alerts that identify integration failures before clients report them.
Call and Lead Operations
- Monitor call-routing, tracking, recording, attribution, and disposition systems.
- Investigate calls that fail to route, connect, record, track, or report correctly.
- Support number provisioning, routing rules, caps, schedules, geographic restrictions, buyer availability, and failover logic.
- Validate that leads, calls, and transactions are properly attributed to the correct advertiser, publisher, campaign, and payout.
- Assist with discrepancies involving call duration, billable events, conversions, payouts, and reporting.
- Help protect revenue by identifying technical leakage and delivery failures.
Data and Reporting Support
- Monitor data pipelines, scheduled jobs, reporting processes, and database performance.
- Investigate discrepancies between platform reporting, billing records, payment records, and third-party systems.
- Write and maintain database queries for troubleshooting, validation, and operational reporting.
- Assist with data corrections using controlled and documented procedures.
- Support dashboards and operational alerts for revenue, margin, consumption, conversion, and platform activity.
- Maintain appropriate controls around production data access and modification.
Security and Access Management
- Support role-based access controls, multifactor authentication, audit logging, encryption, and secure system configuration.
- Provision and remove employee, contractor, client, and vendor access.
- Monitor suspicious activity and report potential security incidents.
- Assist with vulnerability remediation, security reviews, access audits, and incident-response procedures.
- Protect consumer, advertiser, publisher, employee, and company information.
- Follow company policies for credentials, production access, sensitive data, and change management.
Automation and Process Improvement
- Automate repetitive operational tasks using scripts, workflows, APIs, and infrastructure tools.
- Reduce manual work associated with monitoring, deployments, reporting, reconciliation, onboarding, and support.
- Build internal tools that improve visibility and response times.
- Identify operational bottlenecks that affect revenue, margin, client satisfaction, or employee productivity.
- Maintain clear runbooks and standard operating procedures for recurring technical tasks.
Technical Support and Documentation
- Serve as an escalation point for complex platform and integration issues.
- Translate technical problems into clear explanations for nontechnical teams.
- Create and maintain architecture diagrams, system inventories, runbooks, troubleshooting guides, and incident reports.
- Track incidents and technical requests through completion.
- Document known issues, temporary workarounds, permanent resolutions, and system dependencies.
- Participate in an on-call rotation for urgent production incidents.
First 90-Day Priorities
The successful candidate will be expected to:
- Learn EDM’s platforms, infrastructure, integrations, call-routing systems, reporting processes, and revenue workflows.
- Document critical systems, dependencies, credentials ownership, vendor contacts, and escalation procedures.
- Review existing monitoring, alerting, backups, access controls, and deployment procedures.
- Establish baseline metrics for uptime, incident volume, response time, recovery time, deployment success, and integration failures.
- Resolve high-priority recurring production and integration issues.
- Improve alerting for call-routing failures, API errors, delayed data, failed jobs, and reporting discrepancies.
- Create runbooks for the company’s most common and highest-risk technical incidents.
- Identify at least three meaningful automation or cost-saving opportunities.
- Participate in production support and demonstrate ownership of incidents through resolution.
Performance Expectations
Success will be measured by:
- Platform uptime and reliability
- Mean time to acknowledge and resolve incidents
- Reduction in recurring production problems
- Deployment success and rollback rates
- API, postback, webhook, and call-routing reliability
- Reporting and data accuracy
- Backup and recovery readiness
- Quality of technical documentation
- Security and access-control compliance
- Reduction in manual operational work
- Infrastructure costs relative to platform volume
- Responsiveness to internal teams, clients, and technical partners
Required Qualifications
- Three or more years of experience in operations engineering, DevOps, site reliability engineering, cloud infrastructure, systems administration, or production support.
- Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Experience supporting Linux-based production environments.
- Working knowledge of networking, DNS, SSL certificates, firewalls, load balancing, and application security.
- Experience with relational databases and SQL.
- Experience troubleshooting REST APIs, webhooks, authentication, and third-party integrations.
- Familiarity with monitoring, logging, alerting, and incident-management tools.
- Experience with scripting languages such as Python, Bash, JavaScript, or PowerShell.
- Understanding of source control, deployment pipelines, and release management.
- Strong troubleshooting, documentation, and communication skills.
- Ability to prioritize incidents based on business and revenue impact.
- Availability to participate in an on-call rotation.
Preferred Qualifications
- Experience in ad-tech, mar-tech, affiliate marketing, lead generation, pay-per-call, telecommunications, or SaaS.
- Familiarity with SIP, VoIP, dialers, call-tracking platforms, routing systems, and phone-number provisioning.
- Experience with containers, infrastructure as code, and automated deployment tools.
- Experience with Docker, Kubernetes, Terraform, GitHub Actions, or similar technologies.
- Familiarity with payment processing, usage-based billing, reconciliation, and commission systems.
- Experience working with real-time bidding, ping-post, lead distribution, or high-volume transactional systems.
- Understanding of TCPA-related controls, consent records, DNC suppression, data privacy, or regulated marketing environments.
- Experience with security audits, disaster-recovery testing, and compliance documentation.
Ideal Candidate
The ideal candidate:
- Takes ownership instead of waiting for someone else to fix the problem.
- Remains calm and methodical during high-impact incidents.
- Understands the difference between applying a temporary fix and eliminating a root cause.
- Communicates technical problems clearly and directly.
- Recognizes that production reliability is a business and revenue responsibility.
- Automates repetitive work whenever practical.
- Documents systems so the company is not dependent on one person’s memory.
- Protects security and stability without creating unnecessary bureaucracy.
- Is comfortable working in a fast-moving entrepreneurial environment.
- Can manage competing priorities while maintaining attention to detail.
Job Title : DevOps / Infrastructure Engineer
Experience : 3+ Years
Location : Gurugram Sector 48
Work Mode : 6 Days WFO (Monday to Saturday) / 01st & 03rd Saturdays are off
Employment Type : Full-time
Role Overview :
We are looking for a DevOps / Infrastructure Engineer with strong hands-on experience in Linux administration, Jenkins, Docker, networking, bare-metal servers, AWS, Redis, and MongoDB. The candidate should be capable of independently troubleshooting infrastructure, deployment, networking, and application-related issues in production environments.
Mandatory / Non-Negotiable Skills :
- Strong hands-on experience with Linux Administration & Troubleshooting.Strong experience with Jenkins and CI/CD pipelines.
- Hands-on experience with Docker and containerized environments.
- Strong understanding of Networking concepts – TCP/IP, DNS, HTTP/HTTPS, ports, routing, firewalls, load balancing, etc.
- Hands-on experience with Bare Metal Servers / Server Administration.
- Strong hands-on experience with AWS (EC2, VPC, IAM, Security Groups, Load Balancers, S3 & CloudWatch)
- Working knowledge of Redis
- Working knowledge of MongoDB
- Strong production troubleshooting and incident-resolution skills
Key Responsibilities :
- Manage, configure, monitor, and troubleshoot Linux and bare-metal servers
- Build, maintain, and troubleshoot Jenkins CI/CD pipelines
- Deploy, manage, and troubleshoot applications using Docker
- Manage and troubleshoot AWS infrastructure and services
- Configure and maintain networking, security groups, firewalls, ports, and connectivity
- Support and maintain Redis and MongoDB environments
- Perform server health checks, log analysis, performance troubleshooting, and incident resolution
- Work closely with development teams to support application deployments
- Identify root causes of infrastructure and production issues and implement preventive solutions
- Maintain infrastructure security, availability, and reliability
- Automate repetitive operational tasks wherever possible
Required Candidate Profile :
- 3+ years of relevant experience in DevOps, Infrastructure, System Administration, or related roles.
- Strong hands-on / production experience with all mandatory technologies.
- Good understanding of Linux systems and infrastructure.
- Strong troubleshooting and problem-solving abilities.
- Ability to take ownership of production infrastructure and deployment issues.
- Good communication and collaboration skills.
Cloud Infrastructure Engineer – BANG | 10+ Years
Location: Bangalore
Experience: 10+ Years
Job Description:
- Design, implement, and manage cloud infrastructure across AWS/Azure/GCP environments.
- Strong experience in cloud architecture, compute, storage, networking, and security.
- Manage VMs, VPC/VNet, load balancers, DNS, DHCP, firewalls, and IAM.
- Hands-on experience with Windows/Linux servers, VMware, virtualization, and infrastructure operations.
- Automate infrastructure provisioning and configuration using Terraform, Ansible, or similar tools.
- Monitor infrastructure performance, availability, and capacity using tools such as Grafana, Prometheus, or CloudWatch/Azure Monitor.
- Handle incident management, troubleshooting, disaster recovery, backup, and high-availability requirements.
- Work with cross-functional teams to support cloud migration, infrastructure upgrades, and production environments.
- Ensure infrastructure follows security, compliance, and operational best practices.
Must-Have Skills:
Cloud Infrastructure | AWS/Azure/GCP | Networking | Linux/Windows | VMware | Terraform | Ansible | DNS/DHCP | IAM | Monitoring | Backup & DR

Jr Platform Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 0.6-2 years
Shift Timing: 2 PM to 11 PM IST
About Company
It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.
Operating over 70,000 charge points globally,It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.
At this company, we value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.
Role Overview
It is seeking a TechOps Engineer! We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. It engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.
What you’ll do:
- Ensure system reliability, uptime, and performance of global platform.
- Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule)
- Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
- Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
- End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
- Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
- Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
- Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
- Collaborate with development, operations and support teams to build scalable and resilient systems.
- Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
- Participate in capacity planning, performance tuning, and resource optimization.
- Integrate security and compliance best practices into all infrastructure operations.
- Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
- Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
- Flexible to resolve blocking issues during off hours or weekends if required.
What We’re Looking For:
Basic Qualifications and skills
- Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
- Overall 1+ years of experience as a Site Reliability Engineer, DevOps/ Technical project coordinator role.
- Proven experience of DevOps, SRE, or Technical Project Coordination with IoT or connected devices based platforms.
- Hands-on experience with cloud platforms such as AWS and Infrastructure as Code (IaC) tools such as Terraform.
- Experience with incident management and on-call best practices. Provide support to on call engineers.
- Excellent analytical and problem-solving skills with a proactive mindset.
- Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
- Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
- Fluency in English (spoken and written).
- Successfully recommission or decommission chargers following changes in our network.
- Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.
Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
🚀 Hiring – AWS / Kubernetes / OpenShift Engineer
📍 Location: Bangalore
💼 Experience: 7–10 Years
⚡ Joining: Immediate Joiners Only
🔑 Required Skills
- Strong hands-on experience in AWS
- Expertise in Kubernetes & OpenShift
- Strong Linux Administration & Troubleshooting
- Experience in containerized environments and platform operations
- Production support, monitoring and incident troubleshooting
- Good understanding of cloud and infrastructure technologies
📌 Interview Process
2nd Round – Face-to-Face Interview in Bangalore
👉 Please share profiles of candidates who are available for a F2F interview in Bangalore.
#Hiring #AWS #Kubernetes #OpenShift #Linux #CloudEngineer #PlatformEngineer #DevOps #BangaloreJobs #ImmediateJoiners #WFO #ITJobs

Job Title: TechOps Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 6 Month-2 years (Excluding Internship)
Shift Timing: 2 PM to 11 PM IST
Role Overview
We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. Engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.
What you’ll do:
- Ensure system reliability, uptime, and performance of global platform.
- Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule)
- Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
- Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
- End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
- Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
- Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
- Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
- Collaborate with development, operations and support teams to build scalable and resilient systems.
- Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
- Participate in capacity planning, performance tuning, and resource optimization.
- Integrate security and compliance best practices into all infrastructure operations.
- Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
- Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
- Flexible to resolve blocking issues during off hours or weekends if required.
What We’re Looking For:
Basic Qualifications and skills
- Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
- Overall 1 years of experience as a Site Reliability Engineer, Technical project coordinator role.
- Proven experience of SRE or Technical Project Coordination with IoT or connected devices based platforms.
- Experience with incident management and on-call best practices. Provide support to on call engineers.
- Excellent analytical and problem-solving skills with a proactive mindset.
- Hands-on experience with AWS Cloud and IaC tools such as Terraform or Ansible.
- Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
- Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
- Fluency in English (spoken and written).
- Successfully recommission or decommission chargers following changes in our network.
- Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.
Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
Job Description
The engineer will provide direct support to service development teams using the platform and maintain/develop platform components across CI pipelines, tool integrations, deployment architecture, monitoring, documentation, and self-service initiatives.
Responsibilities
- Support onboarding and technical discussions with development teams.
- Provide technical guidance and documentation.
- Support development engineers on CI/CD pipelines, deployments, logging and monitoring.
- Improve platform code, processes and documentation.
- Provide production infrastructure/operations support, including on-call support.
- Research product requirements and new technology rollouts.
Core Mandatory Skills
Kubernetes / Amazon EKS, Terraform, Terragrunt, AWS Cloud, Docker, CI/CD Pipelines, GitOps, Python / Go / Java, Cloud Operations / DevOps, Logging / Monitoring / Tracing, Incident Management
Site Reliability Engineer (SRE) / Production Support Engineer
Experience: 5–10 Years
Location: Hyderabad
Work Mode: Face-to-Face Drive
Shift: Rotational Shifts
Job Description
Looking for an experienced SRE / Production Support Engineer with strong experience in application and production support, incident management, monitoring, troubleshooting, and cloud operations.
Key Skills
Production Support, Incident Management, Splunk, APM, SLI/SLO, Cloud, Kubernetes, Docker, Terraform, Linux/Windows Administration, Shell Scripting and Python.
Good understanding of production deployments, batch monitoring, network/load balancing, and troubleshooting is required.
Candidates from SRE, Production Support, Application Support, Cloud Operations, or DevOps backgrounds are preferred.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Lead Cloud Reliability Engineer
Job Responsibilities
● Lead and manage the Cloud Reliability teams to provide strong Managed Services support to end-customers.
● Isolate, troubleshoot and resolve issues reported by CMS clients in their cloud environment
● Drive the communication with the customer providing details about the issue, current steps, next plan of action, ETA
● Gather client's requirements related to use of specic cloud services and provide assistance in seing them up and resolving issues
● Create SOPs and knowledge articles for use by the L1 teams to resolve common issues
● Identify recurring issues, perform root cause analysis and propose/implement preventive actions
● Follow change management procedure to identify, record and implement changes
● Plan and deploy OS, security patches in Windows/Linux environment and upgrade k8s clusters
● Identify the recurring manual activities and contribute to automation
● Provide technical guidance and educate team members on development and operations. Monitor metrics and develop ways to improve.
● System troubleshooting and problem-solving across plaorm and application domains. Ability to use a wide variety of open-source technologies and cloud services.
● Build, maintain, and monitor conguration standards.
● Ensuring critical system security through using best-in-class cloud security solutions.
Qualifications
● 4-7 years experience in Cloud Infrastructure and Operations domains and IT operational experience preferably in a global enterprise environment.
● Specialize in one or two cloud deployment platforms: AWS, GCP
● Hands on experience with AWS/GCP services (EKS, ECS, EC2, VPC, RDS, Lambda, GKE, Compute Engine)
● Understanding of one or more programming languages (Python, JavaScript, Ruby, Java, .Net)
● Logging and Monitoring tools (ELK, Stackdriver, CloudWatch)
● Knowledge on Conguration Management tools such as Ansible, Terraform, Puppet, Chef
● Experience working with deployment and orchestration technologies (such as Docker, Kubernetes, Mesos)
● Good analytical, communication, problem solving, and learning skills.
● Knowledge on programming against cloud plaorms such as Google Cloud Platform and lean development methodologies.
● Strong service aitude and a commitment to quality.
● Willingness to work in shifts.








