We are seeking an experienced Senior Systems Operations Engineer with 6+ years of experience in Application Production Support, Systems Operations, and Incident Management. The ideal candidate will be responsible for maintaining the stability, availability, and performance of production applications and infrastructure while ensuring adherence to ITIL processes and operational excellence.
The candidate should possess strong expertise in Linux/Unix administration, SQL/Oracle Database support, Production Issue Analysis, Incident & Change Management, and monitoring tools such as Splunk, Grafana, and AppDynamics. The role requires excellent troubleshooting skills, stakeholder communication, and the ability to lead critical incident resolution activities in a 16X7 production environment.
Key Responsibilities:
• Provide advanced production support for critical business applications, ensuring high availability and performance.
• Lead incident and change management processes, including root cause analysis and resolution of production issues.
• Monitor application health using tools such as SPLUNK, Grafana, and AppDynamics.
• Collaborate with development, infrastructure, and business teams to drive continuous improvement in system operations.
• Maintain and optimize Linux/Unix environments and SQL/Oracle databases.
• Document operational procedures, troubleshooting steps, and best practices.
Required Skills:
• Strong experience with Linux/Unix system administration.
• Advanced proficiency in SQL and Oracle database management.
• Expertise in incident and change management within enterprise environments.
• Proven ability to analyse and resolve production issues efficiently.
• Hands-on experience with monitoring and alerting tools: SPLUNK, Grafana, AppDynamics.
• Excellent communication and collaboration skills.
• Experience with cloud platforms (AWS, Azure, GCP)
Desired Candidate Profile
• 6+ years of experience in Application Production Support and Systems Operations.
• Proven experience managing mission-critical production environments.
• Strong expertise in Linux/Unix, SQL, Oracle Database, and monitoring tools.
• Demonstrated success in incident resolution, RCA preparation, and service improvement initiatives.
• Ability to work effectively in a fast-paced 16x7 production support environment.
• Exposure to AI tools and implementation as well