Production Support Engineer (AWS, Java, Microservices, Splunk, cs) at Infraveo Technologies Private Limited · Remote only · 3 - 5 years · ₹9.6L - ₹12L / yr · Bootstrapped · Remote only · Posted 30 Sep 2024

Production Support Engineer (AWS, Java, Microservices, Splunk, cs)
We are seeking a Production Support Engineer to join our team.
Responsibilites:
- Be the first line of defense for production and test environment issues.
- Work collaboratively with the team to identify, manage, and resolve ongoing incidents.
- Troubleshoot and connect with appropriate teams to effectively triage issues impacting test and production environments.
- Understand system architecture, upstream, and downstream dependencies to enable effective participation in triage and restoration activities.
- Perform systems monitoring of applications within the IRS domain after service restoration and post patching, maintenance, and upgrades.
- Create necessary service tickets and ensure tickets are routed to the appropriate technical teams.
- Provide weekend support for various activities including patching, release deployments, security updates, and 3rd party updates.
- Keep up with info alerts, patching alerts, and delivery partners' activities.
- Update stakeholders to plan for upcoming maintenance as well as alert them about service issues and restoration.
- Manage and communicate about upcoming maintenance in the test environment on a daily basis.
- Liaise with various stakeholders to gain approval for alert communications, including confirmation before an all-clear communication.
- Work closely with testing and development teams to prepare for infrastructure updates and release readiness.
- Submit Application Redirects tickets for planned maintenance after gaining approval from management.
- Participate in analysis and improvement of system performance.
- Host daily operational standup.
- Provide additional support to existing production support procedures and process improvements.
- Provide regular status reports to management on application status and other metrics.
- Collaborate with management to improve and customize reports related to production support.
- Plan and manage support for incident management tools and processes.
Requirements:
- Bachelor's Degree in computer science, engineering, or related field.
- AWS Cloud certification.
- 3+ years of relevant IT work experience with cloud experience.
- Knowledge of Java and microservice development and deployments.
- Understanding of the business processes behind applications.
- Strong analytical, problem-solving, negotiation, task and project management, and organizational skills.
- Strong oral and written communication skills, including process documentation.
- Proficiency in Microsoft Office applications (Word, PowerPoint, Excel, and Project).
- Proficiency in knowledge of computer systems, databases, and SharePoint.
- Knowledge of Splunk and AppDynamics.
Benefits:
- Work Location: Remote
- 5 days working
You can apply directly through the link: https://zrec.in/gQWFK?source=CareerSite
Explore our Career Page for more such jobs : careers.infraveo.com

Similar jobs (10)
Job Summary:
We are looking for an experienced Production Support Engineer with strong Linux Administration skills to manage production environments, troubleshoot application and infrastructure issues, and ensure system availability and stability. The role involves Linux administration, production support, database support, monitoring, observability, cloud exposure, and automation.
Primary Skills:
- Linux Administration
- Production Support
Secondary Skills:
- Oracle SQL / Database Support
- Splunk
- Grafana
- AppDynamics
- Cloud – OCP / OpenShift
- Monitoring & Observability
- Shell Scripting / Automation
Key Responsibilities:
- Provide L2/L3 production support for critical applications and infrastructure.
- Perform Linux server administration, troubleshooting, and issue resolution.
- Monitor production systems and proactively identify performance and availability issues.
- Analyze application, system, and infrastructure logs to troubleshoot incidents.
- Work with Oracle databases and perform basic SQL queries and database-related troubleshooting.
- Use monitoring and observability tools such as Splunk, Grafana, and AppDynamics.
- Troubleshoot system performance, connectivity, process, disk, memory, and CPU-related issues.
- Participate in incident, problem, and change management activities.
- Support cloud/container platforms such as OpenShift/OCP.
- Create and maintain Shell scripts for operational tasks and automation.
- Coordinate with application, database, cloud, and infrastructure teams for issue resolution.
- Participate in production deployments, maintenance activities, and release support.
- Follow ITIL processes and maintain proper incident and operational documentation.
- Participate in on-call and production support activities as required.
Mandatory Skills:
- 7+ years of relevant IT experience
- Strong Linux Administration
- Strong Production Support experience
- Good troubleshooting and incident management skills
- Experience working in a 24x7 production environment
- Hands-on experience with monitoring/logging tools
Good to Have:
- Oracle SQL
- Splunk / Grafana / AppDynamics
- OpenShift / OCP
- Shell scripting
- Cloud exposure
- Monitoring & Observability
- Linux troubleshooting
- Hands-on AWS
- Production/Application Support
- Bash/Shell/Python
- Monitoring/log analysis
- Incident resolution
- Application deployment/support
- Basic networking and database knowledge
- Production/batch support exposure
- Willingness for rotational weekend/critical production support
Role Summary:
We are looking for an experienced Application Production Support Engineer with strong expertise in application support, incident and change management, Linux/Unix, SQL, Oracle, and monitoring tools. The candidate will be responsible for maintaining application availability, troubleshooting production issues, monitoring system performance, and coordinating with technical and business stakeholders.
Key Responsibilities
- Provide L2/L3 production support for business-critical applications.
- Monitor applications and infrastructure using Splunk, Grafana, and AppDynamics.
- Analyze and resolve production incidents within defined SLAs.
- Perform incident, problem, change, and service request management.
- Troubleshoot application issues across Linux/Unix, SQL, and Oracle environments.
- Perform SQL queries and database-level troubleshooting to identify application issues.
- Analyze application logs, alerts, and performance metrics to identify root causes.
- Coordinate with development, database, infrastructure, and other technical teams for issue resolution.
- Participate in Root Cause Analysis (RCA) and implement corrective/preventive actions.
- Support application deployments, releases, and production changes.
- Ensure effective communication with business users and stakeholders during critical incidents.
- Identify recurring issues and drive problem management and service improvement initiatives.
- Maintain support documentation, knowledge articles, and operational procedures.
- Participate in on-call/shift support as required.
Mandatory Skills
- 6+ years of experience in Application Production Support.
- Strong experience in Incident & Change Management.
- Hands-on experience with Linux/Unix.
- Good knowledge of SQL and Oracle database support.
- Experience with monitoring and observability tools:
- Splunk
- Grafana
- AppDynamics
- Strong troubleshooting and problem-solving skills.
- Good understanding of application monitoring, logs, alerts, and performance analysis.
- Strong stakeholder management and communication skills
Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting
WFO-Immediate
8 to 12 Yrs
Bangalore/Hyderabad
We are looking for an experienced Application Support Engineer with strong expertise in Linux/Unix, Networking, Routing, Load Balancing, and Production Support. The ideal candidate should have hands-on experience troubleshooting application and infrastructure issues using enterprise monitoring and observability tools such as Splunk, Grafana, and AppDynamics.
Key Responsibilities
- Provide L2/L3 production support for enterprise applications and infrastructure.
- Troubleshoot application, Linux/Unix, network, routing, and connectivity-related issues.
- Monitor application and infrastructure health using Splunk, Grafana, and AppDynamics.
- Analyze alerts, logs, performance metrics, and system behavior to identify issues.
- Troubleshoot routing, load balancing, connectivity, and network-related problems.
- Perform incident investigation, troubleshooting, and root cause analysis.
- Coordinate with Network, Infrastructure, Application, and other technical teams for issue resolution.
- Handle production incidents and ensure timely resolution within defined SLAs.
- Participate in problem management and identify recurring issues.
- Perform application health checks and proactively identify potential failures.
- Maintain troubleshooting guides, runbooks, and operational documentation.
- Support planned changes, deployments, and maintenance activities.
Must-Have Skills
Application Support
Linux / Unix
Networking
Routing
Load Balancing
Splunk
Grafana
AppDynamics
Production Support
Incident Management
Troubleshooting
Required Skills & Tools
- 4–6 years of experience in Application / Production Support.
- Strong hands-on experience with Linux/Unix administration and troubleshooting.
- Good understanding of TCP/IP, networking, routing, and connectivity concepts.
- Practical experience troubleshooting load balancing issues.
- Hands-on experience with monitoring and observability tools such as Splunk, Grafana, or AppDynamics.
- Strong log analysis and troubleshooting skills.
- Experience handling production incidents and working within SLA-driven environments.
- Strong communication and coordination skills.
Preferred Experience
- Exposure to ITIL processes including Incident, Problem, and Change Management.
- Experience with scripting/automation using Shell or Python.
- Exposure to cloud environments is an added advantage.
- Experience supporting enterprise-scale applications and infrastructure.
Key Competencies
- Strong analytical and troubleshooting skills.
- Ability to work under pressure during critical production incidents.
- Strong ownership and accountability.
- Excellent communication and stakeholder management.
- Ability to work collaboratively with global technical teams.
Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
Job Description
We are looking for an experienced Application Support Engineer with strong expertise in production support, networking, DNS, Linux/Unix, and application monitoring.
Key Responsibilities
- Provide application and production support for business-critical applications.
- Monitor application performance, availability, and system health.
- Troubleshoot issues related to Network, DNS, Linux, and Unix.
- Analyze application and system logs to identify and resolve production issues.
- Monitor applications using Splunk, Grafana, and AppDynamics.
- Identify and resolve incidents within defined timelines.
- Perform root cause analysis and support issue resolution.
- Coordinate with technical teams for incident investigation and escalation.
- Conduct application health checks and proactively identify potential issues.
- Maintain proper documentation of incidents, troubleshooting steps, and resolutions.
- Follow standard production support and incident management processes.
Required Skills
- 7–11 years of experience in Application/Production Support.
- Strong knowledge of Networking and DNS.
- Hands-on experience with Linux/Unix.
- Experience with monitoring tools such as Splunk, Grafana, and AppDynamics.
- Strong troubleshooting and analytical skills.
- Good understanding of incident management and production support.
- Excellent communication and coordination skills.
- Ability to work in a fast-paced production support environment.
We are looking for an experienced Application Support Engineer with strong knowledge of Network, DNS, Linux/Unix, and application monitoring tools. The candidate will be responsible for production support, incident troubleshooting, monitoring, and ensuring application availability and stability.
Responsibilities:
- Provide L2/L3 application and production support for critical applications.
- Troubleshoot issues related to Network, DNS, Linux/Unix, and application connectivity.
- Monitor application health, performance, and availability using Splunk, Grafana, AppDynamics, or similar tools.
- Analyze logs, alerts, and performance metrics to identify and resolve incidents.
- Handle incident, problem, and change management activities as per ITIL processes.
- Perform root cause analysis (RCA) and implement preventive measures.
- Coordinate with Network, Infrastructure, Development, and other support teams for issue resolution.
- Participate in production deployments, maintenance activities, and on-call support.
Required Skills:
- Strong experience in Application/Production Support.
- Good knowledge of Linux/Unix administration and troubleshooting.
- Strong understanding of Networking concepts and DNS.
- Hands-on experience with Splunk, Grafana, AppDynamics, or similar monitoring tools.
- Good understanding of TCP/IP, HTTP/HTTPS, load balancing, and connectivity troubleshooting.
- Experience with incident, problem, and change management.
- Strong troubleshooting, communication, and stakeholder management skills.
- Good knowledge of SQL and application log analysis is preferred.
Support Engineer
The Support Engineer will be responsible for providing technical support for the Risk Management System (RMS) product, ensuring that both internal and external users can resolve issues quickly and efficiently. This role involves troubleshooting technical problems, providing solutions, and collaborating with engineering and QA teams to enhance the product. Additionally, the Support Engineer will assist with user training, monitor system performance, and contribute to continuous product improvement.
Roles and Responsibilities
- Act as the first point of contact for users experiencing issues with the RMS platform, both internally and externally.
- Troubleshoot and resolve product-related issues, ensuring minimal disruption to users.
- Collaborate with product and engineering teams to identify and resolve recurring technical problems and escalate more complex issues.
- Document support tickets, ensuring detailed tracking of issues, resolutions, and feedback.
- Provide training and guidance to users on how to effectively use the RMS system.
- Ensure customer satisfaction by providing timely and effective solutions to any issues faced by users.
- Assist in the preparation of knowledge base articles, user documentation, and FAQs to empower users to resolve minor issues independently.
- Respond to user queries in a friendly and professional manner, ensuring clear communication.
- Proactively follow up with customers to ensure issues are resolved to their satisfaction.
Required Skills
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
- Prior experience in software support preferably in a fintech or financial services environment.
- Experience with troubleshooting software and hardware issues in a production environment.
- Hands-on experience with common support/ticketing tools (e.g., JIRA, Bugzilla, Selenium, Postman, etc.).
- Ability to manage multiple priorities and deliverables in a fast-paced environment.
- Familiarity with databases like SQL.
SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.
Core responsibilities include:
- Production monitoring and debugging of live systems.
- Incident investigation, troubleshooting, and problem resolution.
- AWS cloud infrastructure support and maintenance.
- Deployment and operational support activities.
- Supporting a 24x7 production environment.
- Working with GitHub-based development workflows.
- Technical debt remediation and platform improvements.
- Customer issue investigation and support.
- Security and compliance-related work, including FedRAMP initiatives.
Preferred Skills:
AWS (especially S3 and EC2)
Strong debugging and troubleshooting skills
Site Reliability Engineering (SRE) experience
GitHub experience
Basic software development skills
TypeScript/JavaScript knowledge
C# preferred
AI experience is a plus.
Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Jr Platform Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 0.6-2 years
Shift Timing: 2 PM to 11 PM IST
About Company
It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.
Operating over 70,000 charge points globally,It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.
At this company, we value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.
Role Overview
It is seeking a TechOps Engineer! We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. It engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.
What you’ll do:
- Ensure system reliability, uptime, and performance of global platform.
- Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule)
- Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
- Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
- End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
- Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
- Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
- Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
- Collaborate with development, operations and support teams to build scalable and resilient systems.
- Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
- Participate in capacity planning, performance tuning, and resource optimization.
- Integrate security and compliance best practices into all infrastructure operations.
- Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
- Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
- Flexible to resolve blocking issues during off hours or weekends if required.
What We’re Looking For:
Basic Qualifications and skills
- Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
- Overall 1+ years of experience as a Site Reliability Engineer, DevOps/ Technical project coordinator role.
- Proven experience of DevOps, SRE, or Technical Project Coordination with IoT or connected devices based platforms.
- Hands-on experience with cloud platforms such as AWS and Infrastructure as Code (IaC) tools such as Terraform.
- Experience with incident management and on-call best practices. Provide support to on call engineers.
- Excellent analytical and problem-solving skills with a proactive mindset.
- Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
- Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
- Fluency in English (spoken and written).
- Successfully recommission or decommission chargers following changes in our network.
- Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.
Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.





