Production Support at VyTCDC · Mumbai · 1.5 - 1.8 years · ₹1L - ₹6L / yr · Profitable · Posted 15 May 2025

A Production Support Engineer ensures the smooth operation of software applications and IT systems in a production environment. Here’s a breakdown of the role:
Key Responsibilities
- Monitoring System Performance: Continuously track application health and resolve performance issues.
- Incident Management: Diagnose and fix software failures, collaborating with developers and system administrators.
- Troubleshooting & Debugging: Analyze logs, use diagnostic tools, and implement solutions to improve system reliability.
- Documentation & Reporting: Maintain records of system issues, resolutions, and process improvements.
- Collaboration: Work with cross-functional teams to enhance system efficiency and reduce downtime.
- Process Optimization: Suggest improvements to reduce production costs and enhance system stability.
Required Skills
- Strong knowledge of SQL, UNIX/Linux, Java, Oracle, and Splunk.
- Experience in incident management and debugging.
- Ability to analyze system failures and optimize performance.
- Good communication and problem-solving skills.

Similar jobs (10)
Role Summary:
We are looking for an experienced Application Production Support Engineer with strong expertise in application support, incident and change management, Linux/Unix, SQL, Oracle, and monitoring tools. The candidate will be responsible for maintaining application availability, troubleshooting production issues, monitoring system performance, and coordinating with technical and business stakeholders.
Key Responsibilities
- Provide L2/L3 production support for business-critical applications.
- Monitor applications and infrastructure using Splunk, Grafana, and AppDynamics.
- Analyze and resolve production incidents within defined SLAs.
- Perform incident, problem, change, and service request management.
- Troubleshoot application issues across Linux/Unix, SQL, and Oracle environments.
- Perform SQL queries and database-level troubleshooting to identify application issues.
- Analyze application logs, alerts, and performance metrics to identify root causes.
- Coordinate with development, database, infrastructure, and other technical teams for issue resolution.
- Participate in Root Cause Analysis (RCA) and implement corrective/preventive actions.
- Support application deployments, releases, and production changes.
- Ensure effective communication with business users and stakeholders during critical incidents.
- Identify recurring issues and drive problem management and service improvement initiatives.
- Maintain support documentation, knowledge articles, and operational procedures.
- Participate in on-call/shift support as required.
Mandatory Skills
- 6+ years of experience in Application Production Support.
- Strong experience in Incident & Change Management.
- Hands-on experience with Linux/Unix.
- Good knowledge of SQL and Oracle database support.
- Experience with monitoring and observability tools:
- Splunk
- Grafana
- AppDynamics
- Strong troubleshooting and problem-solving skills.
- Good understanding of application monitoring, logs, alerts, and performance analysis.
- Strong stakeholder management and communication skills
- Linux troubleshooting
- Hands-on AWS
- Production/Application Support
- Bash/Shell/Python
- Monitoring/log analysis
- Incident resolution
- Application deployment/support
- Basic networking and database knowledge
- Production/batch support exposure
- Willingness for rotational weekend/critical production support
Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting
WFO-Immediate
8 to 12 Yrs
Bangalore/Hyderabad
We are looking for an experienced Application Support Engineer with strong expertise in Linux/Unix, Networking, Routing, Load Balancing, and Production Support. The ideal candidate should have hands-on experience troubleshooting application and infrastructure issues using enterprise monitoring and observability tools such as Splunk, Grafana, and AppDynamics.
Key Responsibilities
- Provide L2/L3 production support for enterprise applications and infrastructure.
- Troubleshoot application, Linux/Unix, network, routing, and connectivity-related issues.
- Monitor application and infrastructure health using Splunk, Grafana, and AppDynamics.
- Analyze alerts, logs, performance metrics, and system behavior to identify issues.
- Troubleshoot routing, load balancing, connectivity, and network-related problems.
- Perform incident investigation, troubleshooting, and root cause analysis.
- Coordinate with Network, Infrastructure, Application, and other technical teams for issue resolution.
- Handle production incidents and ensure timely resolution within defined SLAs.
- Participate in problem management and identify recurring issues.
- Perform application health checks and proactively identify potential failures.
- Maintain troubleshooting guides, runbooks, and operational documentation.
- Support planned changes, deployments, and maintenance activities.
Must-Have Skills
Application Support
Linux / Unix
Networking
Routing
Load Balancing
Splunk
Grafana
AppDynamics
Production Support
Incident Management
Troubleshooting
Required Skills & Tools
- 4–6 years of experience in Application / Production Support.
- Strong hands-on experience with Linux/Unix administration and troubleshooting.
- Good understanding of TCP/IP, networking, routing, and connectivity concepts.
- Practical experience troubleshooting load balancing issues.
- Hands-on experience with monitoring and observability tools such as Splunk, Grafana, or AppDynamics.
- Strong log analysis and troubleshooting skills.
- Experience handling production incidents and working within SLA-driven environments.
- Strong communication and coordination skills.
Preferred Experience
- Exposure to ITIL processes including Incident, Problem, and Change Management.
- Experience with scripting/automation using Shell or Python.
- Exposure to cloud environments is an added advantage.
- Experience supporting enterprise-scale applications and infrastructure.
Key Competencies
- Strong analytical and troubleshooting skills.
- Ability to work under pressure during critical production incidents.
- Strong ownership and accountability.
- Excellent communication and stakeholder management.
- Ability to work collaboratively with global technical teams.
Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
Support Engineer
The Support Engineer will be responsible for providing technical support for the Risk Management System (RMS) product, ensuring that both internal and external users can resolve issues quickly and efficiently. This role involves troubleshooting technical problems, providing solutions, and collaborating with engineering and QA teams to enhance the product. Additionally, the Support Engineer will assist with user training, monitor system performance, and contribute to continuous product improvement.
Roles and Responsibilities
- Act as the first point of contact for users experiencing issues with the RMS platform, both internally and externally.
- Troubleshoot and resolve product-related issues, ensuring minimal disruption to users.
- Collaborate with product and engineering teams to identify and resolve recurring technical problems and escalate more complex issues.
- Document support tickets, ensuring detailed tracking of issues, resolutions, and feedback.
- Provide training and guidance to users on how to effectively use the RMS system.
- Ensure customer satisfaction by providing timely and effective solutions to any issues faced by users.
- Assist in the preparation of knowledge base articles, user documentation, and FAQs to empower users to resolve minor issues independently.
- Respond to user queries in a friendly and professional manner, ensuring clear communication.
- Proactively follow up with customers to ensure issues are resolved to their satisfaction.
Required Skills
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
- Prior experience in software support preferably in a fintech or financial services environment.
- Experience with troubleshooting software and hardware issues in a production environment.
- Hands-on experience with common support/ticketing tools (e.g., JIRA, Bugzilla, Selenium, Postman, etc.).
- Ability to manage multiple priorities and deliverables in a fast-paced environment.
- Familiarity with databases like SQL.
Role Summary
We are seeking a proactive and technically skilled Python Application Support Engineer to join our Technical Operations team. This role is crucial for ensuring the stability and reliability of our mission-critical, Python-based applications. You will be responsible for timely incident resolution, deep-dive troubleshooting, implementing permanent fixes, and driving operational efficiency through automation.
🔑 Key Responsibilities
Technical Troubleshooting & Incident Management
- Incident Resolution: Serve as the primary point of contact for complex Level 2 and Level 3 production incidents, diagnosing root causes and resolving issues across our Python application stack.
- Deep-Dive Analysis: Utilize log analysis tools (e.g., Splunk, ELK Stack) and monitoring platforms (e.g., Prometheus, Grafana) to quickly identify and address anomalies in application behavior.
- Code Debugging: Analyze, debug, and fix application issues directly within the Python codebase, including Flask/Django services, worker queues, and custom scripts.
- Database Health: Troubleshoot performance issues and conduct basic SQL/NoSQL query tuning and health checks (e.g., for PostgreSQL, MongoDB, or Redis).
Operational Excellence & Automation
- Monitoring & Alerting: Continuously refine and optimize application monitoring, alerting, and logging configurations to improve mean time to detect (MTTD) and mean time to resolve (MTTR).
- Python Automation: Develop, maintain, and enhance automated scripts (primarily in Python) to streamline routine operational tasks, reporting, health checks, and system recovery processes.
- Documentation: Create and maintain comprehensive documentation, runbooks, and knowledge base articles for application support procedures and recurring issues.
Collaboration & Prevention
- Cross-Functional Fixes: Collaborate closely with the Development and DevOps teams to provide clear technical feedback on recurring issues and implement permanent, scalable solutions.
- Proactive Maintenance: Identify potential system bottlenecks, performance degradation points, and areas prone to failure, recommending and implementing preventative measures.
⚙️ Required Qualifications
- Experience: 3 to 5 years of professional experience in Application Support, Production Support, Site Reliability Engineering (SRE), or a similar technical role.
- Python Expertise (Mandatory): Strong hands-on experience with Python scripting and programming, including the ability to read, debug, and modify application code.
- Operating Systems: Proficient working knowledge of Linux/Unix environments and shell scripting.
- Databases: Solid experience with relational (e.g., PostgreSQL, MySQL) and/or NoSQL (e.g., MongoDB, Redis) databases, focusing on query analysis and performance.
About the Role
We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.
Key Responsibilities
- Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
- Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
- Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
- Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
- Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
- Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
- Participate in incident/severity calls, ensuring clear communication and coordination
- Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting
Required Skills & Experience
- Strong understanding of Linux/Unix systems for application support
- Hands-on experience troubleshooting applications in staging and production environments
- Ability to monitor system performance and identify root causes using logs and metrics
- Experience working with Kubernetes and microservices-based architectures
- Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
- Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
- Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
- Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
- Experience with ticketing and documentation tools such as Jira and Confluence
- Minimum 4+ years of experience in application support or reliability engineering
Education & Certifications
- Bachelor's degree in Computer Science, Information Technology, or a related field
- Relevant certifications (Cloud, Kubernetes, Microservices) are a plus
Work Schedule
- Willingness to work in a 24x7 environment, including weekends and on-call rotations
ROLE:
Address technical issues relating to software implementation, function, and upgrades. Resolve customer complaints or problems and create product problem reports and troubleshoot documents for each issue. Work closely with application support and development teams to identify and resolve any technical problems that might arise during the development of software. Work with Implementation teams and Account managers to recommend solutions to new customer implementations and workflows.
ESSENTIAL DUTIES and RESPONSIBILITIES:
- Understanding the PeopleScout architecture and framework
- Provide L2 support for tickets passed on by L1 support team using advanced knowledge of Java, SQL, Stored Procedures, Functions, and database development.
- Analyze and understand Java code to debug, trace, and communicate technical issues effectively with the Development team.
- Review application logs, perform exception analysis, monitor system performance, and identify root causes to ensure timely issue resolution and optimal application performance.
- Provide technical support to application support team
- Develop solutions to complex customer and software problems
- Assist with software design and development
- Work within an agile environment
- Deliver sprint commitments on time
- Be responsible for your own code, and work with others to improve the quality and deliver of theirs
- Document troubleshooting guides and outcomes of problems, analysis and solutions for future re-use
MUST HAVE SKILLS:
- 5 to 8 years of experience in L2 Application Support, preferably supporting SaaS products.
- Strong troubleshooting skills with Java applications, SQL, and REST APIs.
- Strong hands-on experience with Core Java (Java 8 or above).
- Good understanding of Spring Boot and Java-based enterprise applications.
- Experience troubleshooting production issues, performing root cause analysis (RCA), and resolving incidents within SLA
- Strong SQL skills with databases such as Oracle, MySQL, PostgreSQL, or SQL Server.
- Experience working with REST APIs, JSON, and API testing tools like Postman.
- Ability to analyze application logs using tools such as Splunk, Kibana, ELK, or Grafana.
- Excellent customer communication skills with experience managing client interactions and providing timely updates.
- Familiarity with ServiceNow/Jira, Linux basics, and log analysis for application support
SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.
Core responsibilities include:
- Production monitoring and debugging of live systems.
- Incident investigation, troubleshooting, and problem resolution.
- AWS cloud infrastructure support and maintenance.
- Deployment and operational support activities.
- Supporting a 24x7 production environment.
- Working with GitHub-based development workflows.
- Technical debt remediation and platform improvements.
- Customer issue investigation and support.
- Security and compliance-related work, including FedRAMP initiatives.
Preferred Skills:
AWS (especially S3 and EC2)
Strong debugging and troubleshooting skills
Site Reliability Engineering (SRE) experience
GitHub experience
Basic software development skills
TypeScript/JavaScript knowledge
C# preferred
AI experience is a plus.
Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.
Hiring Platform Engineer
Exp: 6 -- 10 yrs
Edu : BE/B.tech/MCA
Work Location : Pune
Skills :
Platform monitoring ,Incident trouble shooting, Incident recovery, openshift ,kubernetes.
2 years of IT operations, infrastructure, cloud or application support experience.
Exp in Linux command-line knowledge.
Exp in networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.






