Cutshort logo
For Employers
Our client is a top-tier global executive search and leaders logo
Lead Monitoring
Our client is a top-tier global executive search and leaders

Lead Monitoring at Our client is a top-tier global executive search and leaders · Gurugram · 8 - 15 years · ₹6L - ₹20L / yr · Posted 3 Jul 2025

NeoGenCode Technologies Pvt Ltd's logo

Lead Monitoring

at Our client is a top-tier global executive search and leaders

8 - 15 yrs
₹6L - ₹20L / yr
Gurugram
Skills
Nagios
Site survey
Microsoft SCOM
Site 24x7
Cloud Computing
KPI
Monitoring

Job Title: Lead - Monitoring

Location: Gurgaon

Experience Required: 8+ Years

Job Type: Full Time

Position Summary

We are seeking a highly skilled and motivated individual to join our team as Lead - Monitoring.

As the Lead - Monitoring, you will play a crucial role in overseeing and optimizing our systems and networks. Your responsibilities will include monitoring the performance metrics of our IT infrastructure. Additionally, you will lead troubleshooting efforts, identify and resolve system issues, and implement proactive measures to minimize downtime and disruptions.

This role requires a keen eye for detail, strong analytical skills, and the ability to collaborate effectively with technical teams to implement solutions and improve overall system performance.

Roles and Responsibilities

  • Monitor System Performance:
  • Oversee the monitoring of system performance metrics, including uptime, response times, and resource utilization, using monitoring tools such as Nagios, Microsoft SCOM, Site24X7, and other third-party tools, to ensure optimal performance and availability.
  • Troubleshooting and Issue Resolution:
  • Lead the identification, troubleshooting, and resolution of system issues, working closely with technical teams to implement solutions and minimize downtime.
  • Capacity Planning:
  • Develop and implement capacity planning strategies to forecast future resource needs and optimize system scalability and performance.
  • Incident Response:
  • Develop and maintain incident response protocols and procedures, including escalation paths and response timelines, to address system outages and critical incidents promptly.
  • Monitoring Tools Management:
  • Evaluate, select, and manage monitoring tools and technologies to support efficient and effective monitoring of systems, networks, and applications.
  • Performance Analysis:
  • Conduct performance analysis and trend analysis to identify potential bottlenecks, areas for improvement, and optimization opportunities.
  • Cloud Monitoring:
  • Implement and manage cloud monitoring solutions for platforms such as Azure and AWS, ensuring visibility into cloud-based resources, performance metrics, and cost optimization strategies. Monitor cloud infrastructure, services, and applications to identify and resolve issues proactively.
  • Synthetic Monitoring:
  • Design and implement synthetic monitoring solutions to simulate user interactions and transactions across applications, websites, and services. Analyze synthetic monitoring data to identify performance bottlenecks and optimize user experience.
  • Documentation and Reporting:
  • Maintain accurate documentation of monitoring processes, configurations, and incident reports. Generate regular reports on system performance, uptime, and incident resolution metrics.
  • KPIs and Dashboards:
  • Develop and publish key performance indicators (KPIs), dashboards, and other reporting mechanisms to provide insights into system performance, trends, and areas for improvement. Present findings and recommendations to stakeholders and management.
  • Team Leadership:
  • Provide leadership and guidance to monitoring team members, fostering a culture of collaboration, continuous improvement, and excellence in monitoring practices.

Skills and Qualifications

  • Education:
  • Bachelor’s degree in Computer Science, Information Technology, or a related field. (Master’s degree preferred)
  • Experience:
  • Proven experience (5+ years) in system monitoring, performance analysis, and incident response, preferably in a lead or supervisory role.
  • Strong technical expertise in monitoring tools such as Nagios, Microsoft SCOM, Site24X7, and other third-party tools.
  • Solid understanding of network protocols, server infrastructure, and cloud environments (e.g., AWS, Azure), with experience in cloud monitoring, synthetic monitoring, and optimization.
  • Technical Skills:
  • Experience with scripting languages (e.g., Python, PowerShell) for automation and monitoring tasks.
  • Strong analytical, problem-solving, and decision-making skills.
  • Soft Skills:
  • Strong leadership, communication, and team collaboration abilities.
  • Experience in publishing KPIs, dashboards, and other reporting mechanisms.

Good to Have Certifications (Preferred):

  • ITIL Foundation
  • Certified Monitoring Professional (CMP)
  • Microsoft Certified: Azure Administrator Associate
  • Microsoft Certified: SCOM (if available)


Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

MNC
MNC
Agency job
via by Sandhiya b
Bengaluru (Bangalore), Hyderabad
7 - 12 yrs
₹2L - ₹15L / yr
Windows Azure
Splunk
skill icongrafana
AppDynamics

Job Description

We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.

Responsibilities

  • Design, deploy, and manage Azure cloud infrastructure and services.
  • Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
  • Configure dashboards, alerts, health rules, and monitoring metrics.
  • Perform log analysis, troubleshooting, and root-cause analysis for production issues.
  • Implement observability solutions for applications, cloud infrastructure, and services.
  • Automate monitoring and operational activities using scripting.
  • Support incident, problem, and change management processes.
  • Collaborate with development, DevOps, and SRE teams to improve system reliability.
  • Maintain monitoring standards, documentation, and operational procedures.

Primary Skills

  • Microsoft Azure
  • Splunk
  • Grafana
  • AppDynamics
  • Cloud Monitoring & Observability
  • Application Performance Monitoring (APM)
  • Log Analysis & Troubleshooting

Secondary Skills

  • Azure Monitor / Log Analytics
  • Azure VMs, Storage, Networking
  • Linux
  • Python / PowerShell / Shell Scripting
  • CI/CD
  • Git
  • ITIL / ServiceNow
Read more
MNC
MNC
Agency job
via by aafia parveen
Bengaluru (Bangalore), Hyderabad
7 - 14 yrs
₹2L - ₹15L / yr
Axure
Windows Azure
skill icongrafana
Splunk
AppDynamics

Job Description

We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.

Responsibilities

  • Design, deploy, and manage Azure cloud infrastructure and services.
  • Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
  • Configure dashboards, alerts, health rules, and monitoring metrics.
  • Perform log analysis, troubleshooting, and root-cause analysis for production issues.
  • Implement observability solutions for applications, cloud infrastructure, and services.
  • Automate monitoring and operational activities using scripting.
  • Support incident, problem, and change management processes.
  • Collaborate with development, DevOps, and SRE teams to improve system reliability.
  • Maintain monitoring standards, documentation, and operational procedures.

Primary Skills

  • Microsoft Azure
  • Splunk
  • Grafana
  • AppDynamics
  • Cloud Monitoring & Observability
  • Application Performance Monitoring (APM)
  • Log Analysis & Troubleshooting

Secondary Skills

  • Azure Monitor / Log Analytics
  • Azure VMs, Storage, Networking
  • Linux
  • Python / PowerShell / Shell Scripting
  • CI/CD
  • Git
  • ITIL / ServiceNow


Read more
MNC
MNC
Agency job
via by Sandhiya b
Bengaluru (Bangalore), Hyderabad
8 - 15 yrs
₹2L - ₹24L / yr
Observability
AppDynamics
Splunk
skill iconPython
Scripting

Observability Engineer (AppDynamics)

Hyderabad

Exp: 8+years of exp

Mandate Skills: AppDynamics, Splunk, Python (Scripting knowledge)

Primary Skill Set:

  • AppDynamics administration
  • SPLOC administration
  • Enterprise monitoring and observability
  • Application performance monitoring
  • Alerting and event management
  • Monitoring strategy and design
  • Platform configuration and governance
  • Python scripting
  • Monitoring automation

 

Secondary Skills:

  • Glassbox monitoring
  • Customer journey observability
  • GenAI concepts for operations       
  • Log, metric, and trace telemetry
  • Dashboarding and visualization


Read more
Hiring for DevOps lead
Hiring for DevOps lead
Agency job
via by rincy v
Remote only
7 - 20 yrs
₹12L - ₹18L / yr
DevOps
DataDog
RUM

Hiring: DevOps Lead

📍 Kochi / Trivandrum / Remote

💼 Full-time

🕐 General Shift | Australian Overlap

We are looking for an experienced DevOps Lead to join our team.


🔹 Key Responsibilities

Implement and continually improve the observability platform using Datadog, particularly from a user experience perspective.

Work across engineering squads as a virtual team member, supporting their DevOps, infrastructure, and observability requirements.

Configure and manage Datadog RUM, Session Replay, Distributed Tracing, and APM.

Support campaign readiness activities, including load and performance testing.

Participate in gamedays and incident response activities for production systems.

Liaise closely with the managed infrastructure provider on infrastructure requirements and activities.


🔹 Essential Skills & Requirements

✅ 7+ years of relevant DevOps / Cloud / Observability experience

✅ Strong hands-on experience with Datadog

✅ Strong experience with RUM, Session Replay, Distributed Tracing, and APM

✅ Solid experience with AWS

✅ Strong Infrastructure-as-Code experience using CloudFormation and AWS CDK

✅ Good working knowledge of GitHub Actions and AWS CodePipeline

✅ Real incident response experience on high-traffic systems

✅ Experience with Load & Performance Testing

✅ Strong problem-solving and communication skills

✅ Ability to work closely with infrastructure partners and internal platform teams


🔹 Skills - Good to Have

⭐ Experience with high-traffic platforms

⭐ Experience with campaign readiness and gamedays

⭐ Infrastructure partner management

⭐ Advanced AWS observability

⭐ Performance engineering


📌 Experience: 7+ Years

📌 Work Location: Kochi / Trivandrum / Remote

📌 Shift: General Shift with Australian Overlap

📌 Remote: Mandatory 1 week at office


📩 Interested candidates can share their resume

Read more
Product Based Co
Product Based Co
Agency job
via by Rishika Teja
Hyderabad
18 - 25 yrs
₹70L - ₹80L / yr
SRE
skill iconAmazon Web Services (AWS)

Hiring SRE - Director


Exp : 18 - 25 yrs

Edu : BE/B.Tech

Work Location : Hyd


Must Have Skills :


Must be from product SaaS based companies.


18+ years in Software Engineering, SRE, or reliability roles; 5+ years in leadership(Director). 


Proven ability to leverage software engineering principles and practices to solve reliability and operational challenges.


Expertise in SLI/SLO and monitoring.


Expertise in CI/CD, observability, and incident response. 


Strong AWS knowledge and experience with container orchestration. 


Proven ability to lead reliability programs across multiple SaaS products. 


Experience architecting applications or infrastructure for high-growth cloud platforms. 


Experience in B2B SaaS environments involving large-scale distributed systems. 



Read more
company logo
Harsha Mehrotra
Posted by Harsha Mehrotra
Remote only
3 - 5 yrs
Best in industry
CI/CD
skill iconPostgreSQL
skill iconPython
skill iconAmazon Web Services (AWS)
Artificial Intelligence (AI)
+2 more

SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.


Core responsibilities include:

  • Production monitoring and debugging of live systems.
  • Incident investigation, troubleshooting, and problem resolution.
  • AWS cloud infrastructure support and maintenance.
  • Deployment and operational support activities.
  • Supporting a 24x7 production environment.
  • Working with GitHub-based development workflows.
  • Technical debt remediation and platform improvements.
  • Customer issue investigation and support.
  • Security and compliance-related work, including FedRAMP initiatives.


Preferred Skills:

AWS (especially S3 and EC2)

Strong debugging and troubleshooting skills

Site Reliability Engineering (SRE) experience

GitHub experience

Basic software development skills

TypeScript/JavaScript knowledge

C# preferred

AI experience is a plus.


Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Read more
company logo
Priya Rawat
Posted by Priya Rawat
Gurugram
4 - 5 yrs
₹8L - ₹10L / yr
RCA
SLA
skill icongrafana
ELKI
SOP

About the Role


We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.


Key Responsibilities


  • Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
  • Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
  • Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
  • Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
  • Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
  • Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
  • Participate in incident/severity calls, ensuring clear communication and coordination
  • Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting


Required Skills & Experience


  • Strong understanding of Linux/Unix systems for application support
  • Hands-on experience troubleshooting applications in staging and production environments
  • Ability to monitor system performance and identify root causes using logs and metrics
  • Experience working with Kubernetes and microservices-based architectures
  • Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
  • Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
  • Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
  • Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
  • Experience with ticketing and documentation tools such as Jira and Confluence
  • Minimum 4+ years of experience in application support or reliability engineering


Education & Certifications


  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Relevant certifications (Cloud, Kubernetes, Microservices) are a plus


Work Schedule


  • Willingness to work in a 24x7 environment, including weekends and on-call rotations
Read more
MNC Client
MNC Client
Agency job
via by Gauri Naik
Remote only
5 - 9 yrs
Best in industry
Dynatrace
Active Gate
Monitoring
ServiceNow
JIRA
+2 more

Key Responsibilities

  • Design, deploy, configure, upgrade, and administer enterprise-scale Dynatrace environments.
  • Deploy and manage OneAgent, ActiveGate, extensions, and monitoring configurations across application and infrastructure environments.
  • Implement observability for applications, APIs, microservices, databases, containers, Kubernetes, cloud platforms, and traditional infrastructure.
  • Configure service detection, process groups, management zones, tags, naming rules, metrics, logs, traces, and topology.
  • Develop operational and executive dashboards, notebooks, reports, SLOs, and alerting strategies.
  • Configure and optimize Davis AI problem detection, anomaly detection, baselines, and root-cause analysis.
  • Implement Real User Monitoring (RUM), Synthetic Monitoring, Session Replay, distributed tracing, and log monitoring as required.
  • Analyze application performance issues, service dependencies, transaction traces, response times, resource utilization, and infrastructure bottlenecks.
  • Lead troubleshooting of complex production performance and availability incidents using Dynatrace telemetry.
  • Reduce alert noise through effective event correlation, thresholds, anomaly-detection configuration, and monitoring standards.
  • Integrate Dynatrace with enterprise platforms such as ServiceNow, Jira, PagerDuty, Splunk, CI/CD pipelines, and collaboration/notification tools.
  • Automate Dynatrace configuration and deployment using APIs, configuration-as-code, scripting, and DevOps tooling.
  • Work closely with application, infrastructure, cloud, SRE, DevOps, production support, and operations teams to define observability requirements.
  • Establish Dynatrace monitoring standards, reusable configurations, governance, and best practices.
  • Perform platform health checks, capacity assessments, license/usage optimization, and monitoring coverage reviews.
  • Create technical documentation, runbooks, architecture diagrams, troubleshooting guides, and operational procedures.
  • Mentor junior engineers and provide technical leadership for observability initiatives.

Required Technical Skills

  • 5–8+ years of overall IT experience with significant experience in application/infrastructure monitoring or observability.
  • 3–5+ years of hands-on Dynatrace experience in enterprise environments.
  • Strong knowledge of:
  • Dynatrace OneAgent and ActiveGate
  • Application Performance Monitoring (APM)
  • Infrastructure Monitoring
  • Distributed Tracing
  • Real User Monitoring (RUM)
  • Synthetic Monitoring
  • Log Monitoring and Analytics
  • Metrics, traces, logs, and events
  • Davis AI and automated root-cause analysis
  • Dashboards, SLOs, alerting, and anomaly detection
  • Dynatrace APIs and automation
  • Experience monitoring Java/JVM, .NET, web applications, APIs, microservices, and databases.
  • Experience with Kubernetes, Docker, OpenShift, or other container platforms.
  • Working knowledge of at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.
  • Strong understanding of application architecture, HTTP/HTTPS, REST APIs, networking, operating systems, and middleware.
  • Experience with Linux and Windows environments.
  • Scripting/automation experience using Python, PowerShell, Bash, Ansible, Terraform, or similar technologies.
  • Experience integrating monitoring platforms with ITSM, incident management, and DevOps tools.
  • Strong analytical and production troubleshooting skills.

Preferred Skills

  • Experience with Dynatrace Grail, DQL (Dynatrace Query Language), OpenPipeline, and the latest Dynatrace platform capabilities.
  • Knowledge of OpenTelemetry (OTel) and modern telemetry standards.
  • Experience implementing observability for large-scale Kubernetes and cloud-native environments.
  • Familiarity with SRE practices, including SLIs, SLOs, error budgets, and observability-driven incident management.
  • Knowledge of additional monitoring platforms such as Splunk, AppDynamics, Datadog, New Relic, Grafana, Prometheus, or ELK.
  • Experience with infrastructure-as-code and configuration-as-code approaches.
  • Dynatrace certifications are preferred.

Professional Skills

  • Strong problem-solving and root-cause analysis capabilities.
  • Ability to troubleshoot complex application and infrastructure performance issues independently.
  • Strong written and verbal communication skills.
  • Ability to work effectively with application owners, developers, SREs, infrastructure teams, and senior stakeholders.
  • Ability to translate business and operational requirements into observability solutions.
  • Experience working in enterprise production environments with incident, problem, and change-management processes.
  • Ability to mentor engineers and drive technical standards across teams.

Education & Certification

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent professional experience.
  • Dynatrace Associate/Professional-level certification is desirable.
  • Cloud, Kubernetes, ITIL, SRE, or DevOps certifications are advantageous.

Experience Level

Senior Engineer

  • Overall Experience: 5–8+ years
  • Dynatrace/Observability Experience: 3–5+ years
  • Expected proficiency: Advanced hands-on implementation, administration, troubleshooting, automation, and solution design


Read more
MNC
MNC
Agency job
via by ASHMI GR
Bengaluru (Bangalore), Hyderabad
7 - 12 yrs
₹25L - ₹30L / yr
skill iconMongoDB
Observability
Splunk
Oracle
Thousand Eyes
+1 more

Job Summary

We are looking for an experienced Observability Engineer with strong expertise in ThousandEyes, Splunk, database technologies, and Python automation. The candidate will be responsible for monitoring application, network, and infrastructure performance, identifying issues, and developing automation solutions to improve system visibility, reliability, and operational efficiency.

Required Skills

  • Observability
  • ThousandEyes / Cisco ThousandEyes
  • Splunk
  • OracleDB
  • MongoDB
  • SQL
  • Python Automation

Roles & Responsibilities

  • Implement and support observability and monitoring solutions across applications, networks, and infrastructure.
  • Work with ThousandEyes for network and digital experience monitoring.
  • Configure and maintain monitoring dashboards, alerts, and performance metrics.
  • Use Splunk for log analysis, monitoring, troubleshooting, and reporting.
  • Work with OracleDB, MongoDB, and SQL for data analysis and troubleshooting.
  • Develop Python scripts for monitoring and operational automation.
  • Analyze performance issues and identify root causes across applications and infrastructure.
  • Collaborate with application, infrastructure, and network teams to resolve incidents.
  • Improve monitoring, alerting, and automation processes.
  • Prepare performance and monitoring reports and provide actionable insights.

Mandatory Skills

  • 7+ years of relevant experience
  • Strong experience in Observability / Monitoring
  • Hands-on experience with ThousandEyes
  • Splunk
  • Python Automation
  • SQL
  • OracleDB / MongoDB


Read more
MNC
MNC
Agency job
via by Sandhiya b
Mumbai
6 - 13 yrs
₹1.5L - ₹15L / yr
ITRS

Hands-on experience with ITRS Geneos monitoring platform.

Strong knowledge of Geneos Gateway, Netprobe, Active Console and Web Dashboard.

Experience in configuring Geneos monitoring, rules, alerts, dashboards and Toolkit samplers.

Strong Linux/Unix administration and troubleshooting skills.

Good scripting experience with Shell/Bash for monitoring and automation.

Experience monitoring critical applications, infrastructure and production environments.

Knowledge of Git, Jenkins and CI/CD is an advantage.

Exposure to Kafka, Kubernetes, Cassandra, Splunk/Grafana is preferred.

Experience integrating Geneos with ServiceNow/Jira or other ITSM/alerting tools.

Ability to troubleshoot production incidents and perform root-cause analysis.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos