Incident Manager – Fault Management at NLTS INDIA PVT LTD · Gurugram, Mumbai, Pune · 6 - 10 years · ₹5L - ₹14L / yr · Profitable · Posted 25 Feb 2026

Job Title: Incident Manager – Fault Management
Function: Incident, Problem & Change Management
Department: NOC Operations / Command Centre / Service Operations
Experience: 6–10 Years
Employment Type: Full-Time
Shift: 24x7 Rotational (as per business requirement)
Location: Gurgaon / Mumbai / Pune (or as per project needs)
Job Summary
We are looking for an experienced Incident Manager to lead fault management operations, ensuring rapid restoration of services and minimal business impact. The role focuses on Incident, Problem, and Change Management, acting as a central point of coordination during major incidents and ensuring compliance with ITIL processes, SLAs, and governance standards.
The ideal candidate will have strong operational leadership, stakeholder communication, and escalation management skills within complex IT / Telecom environments.
Key Responsibilities
Incident Management
- Own and manage P1/P2/P3 incidents end-to-end in line with ITIL standards.
- Act as Incident Commander during major incidents, leading bridge calls and coordinating technical teams.
- Ensure timely incident detection, logging, categorization, prioritization, and resolution.
- Drive restoration efforts and ensure adherence to SLAs, OLAs, and KPIs.
- Provide regular incident status updates to customers, management, and stakeholders.
- Ensure proper incident documentation, closure notes, and audit readiness.
Fault Management & Monitoring
- Oversee proactive fault detection through NOC monitoring tools.
- Ensure alarms and alerts are correlated, triaged, and assigned appropriately.
- Coordinate with L2/L3 engineering teams for fault isolation and resolution.
- Identify recurring faults and initiate preventive actions.
Problem Management
- Lead Root Cause Analysis (RCA) for recurring and major incidents.
- Facilitate Post-Incident Reviews (PIRs) and track corrective and preventive actions (CAPA).
- Maintain problem records and trend analysis to reduce repeat incidents.
- Work closely with engineering and vendors to drive permanent fixes.
Change Management
- Govern changes to production environments to minimize risk.
- Review and validate Change Requests (CRs), MOPs, rollback plans, and impact assessments.
- Participate in CAB (Change Advisory Board) meetings.
- Ensure changes are executed as per approved windows with pre/post validation.
- Track change-related incidents and drive improvement actions.
Stakeholder & Vendor Coordination
- Act as a single point of contact during service-impacting events.
- Coordinate with internal teams, service providers, OEMs, and vendors.
- Manage customer communication during outages and critical events.
- Escalate issues appropriately to senior management when required.
Governance, Reporting & Continuous Improvement
- Prepare and publish incident, problem, and change management reports (daily/weekly/monthly).
- Monitor and improve operational KPIs and SLA performance.
- Drive process improvements aligned with ITIL best practices.
- Maintain SOPs, runbooks, escalation matrices, and communication templates.
- Support audits, compliance reviews, and regulatory requirements (if applicable).
Required Skills & Competencies
Technical & Process Skills
- Strong expertise in ITIL Incident, Problem, and Change Management.
- Experience working in NOC / Command Centre / Telecom / Enterprise IT Operations.
- Good understanding of infrastructure domains:
- Network (LAN/WAN/SD-WAN)
- Security (Firewalls, SOC coordination)
- Data Center / Cloud (basic understanding)
- Familiarity with monitoring tools (SolarWinds, Netcool, Splunk, PRTG, etc.).
- Hands-on experience with ITSM tools such as ServiceNow, Remedy, Helix, Jira.
Soft Skills
- Strong leadership and decision-making abilities during high-pressure situations.
- Excellent verbal and written communication skills.
- Strong stakeholder and customer management capability.
- Analytical mindset with attention to detail.
- Ability to work independently and in cross-functional teams.
Education & Certifications
- Bachelor’s degree in Engineering, IT, Computer Science, or related field.
- ITIL Foundation (mandatory); ITIL Intermediate/Expert is a plus.
- PMP / PRINCE2 / Agile certifications are advantageous.
Experience
- 6–10 years of experience in Incident / Problem / Change Management roles.
- Prior experience handling Major Incidents in 24x7 operations environments.
- Experience in Telecom, BFSI, Managed Services, or Large Enterprise IT preferred.
Key Performance Indicators (KPIs)
- Incident response and resolution times.
- SLA and availability compliance.
- Reduction in repeat incidents.
- Quality and timeliness of RCA reports.
- Change success rate and reduction in change-related incidents.

Similar jobs (1)
Position Title: Network Operations Center (NOC) Analyst
Experience Required: 3+ Years
Working Mode: Remote
Preferred Location: Hyderabad / Bangalore / Chennai
Shift Timing: Rotational Shifts (US Timings)
About the Role
We are looking for a proactive and technically strong NOC Analyst to support MSP clients' enterprise network infrastructure. The ideal candidate should have hands-on experience in network monitoring, monitoring tool administration, incident management, basic network troubleshooting, documentation, and coordination with ISP, Field Engineers, and customers while ensuring maximum network availability and SLA compliance.
Mandatory Requirements
· Strong communication skills with customers and technical teams.
· Minimum 3+ years of experience in Enterprise NOC / Network Operations.
· Experience working in 24×7 rotational shift environments.
· Good understanding of Enterprise Network Monitoring and ITIL-based Incident Management.
· MSP environment experience.&ITIL Foundation certification.
· Exposure to AWS/Azure networking is an added advantage.
· Valid CCNA Certification.( Preferred )
Technical Skills Required
Networking: TCP/IP, DNS, DHCP, Subnetting, IPv4 Addressing, LAN/WAN, VPN
Routing & Switching: Basic troubleshooting of OSPF, BGP; good understanding of VLAN, STP and Trunking
Monitoring Protocols: SNMP, Syslog, NetFlow, Telemetry
Monitoring Tools: SolarWinds, LogicMonitor, Auvik, PRTG, Zabbix, Nagios, ManageEngine
Ticketing Tools: ServiceNow, ConnectWise, Jira, Zendesk
Basic Knowledge: Windows Server, Linux, Network Devices, Firewall Concepts
Key Responsibilities
Network Monitoring
· Monitor enterprise network infrastructure using monitoring platforms.
· Integrate and onboard new network devices/nodes into monitoring tools.
· Configure and maintain monitoring templates, thresholds, and alert profiles.
· Ensure continuous availability of monitored infrastructure.
Incident Management
· Monitor alerts and incidents generated from monitoring platforms.
· Perform ticket triaging, categorization and prioritization.
· Perform Level-1 and basic Level-2 network troubleshooting.
· Escalate incidents appropriately and ensure SLA compliance.
Network Troubleshooting
· Basic troubleshooting of OSPF, BGP, Routing & Switching.
· Troubleshoot DNS, DHCP, TCP/IP, LAN/WAN, VPN and network reachability.
· Perform subnet calculations and IP addressing validation.
Monitoring Protocols
· Hands-on knowledge of SNMP, Syslog, NetFlow and Telemetry.
· Understand device discovery, performance monitoring, fault monitoring and alert generation.
Coordination & Documentation
· Coordinate with Field Engineers, ISP providers, Client Technical Teams and internal L2/L3 teams.
· Track incidents until complete resolution.
· Prepare RCA, SOP, Incident Reports, Shift Handover Documents, Knowledge Base Articles and Weekly/Monthly Operational Reports.
Educational Qualification
Bachelor's Degree in Computer Science, Information Technology, Electronics, or a related field.
Notice Period
Immediate Joiners / Short Notice Candidates Preferred.






