SRE at VY SYSTEMS PRIVATE LIMITED · Bengaluru (Bangalore) · 5 - 9 years · ₹4L - ₹16L / yr · Profitable · Posted 22 Jul 2026

We are seeking an experienced Senior Systems Operations Engineer with 6+ years of experience in Application Production Support, Systems Operations, and Incident Management. The ideal candidate will be responsible for maintaining the stability, availability, and performance of production applications and infrastructure while ensuring adherence to ITIL processes and operational excellence.
The candidate should possess strong expertise in Linux/Unix administration, SQL/Oracle Database support, Production Issue Analysis, Incident & Change Management, and monitoring tools such as Splunk, Grafana, and AppDynamics. The role requires excellent troubleshooting skills, stakeholder communication, and the ability to lead critical incident resolution activities in a 16X7 production environment.
Key Responsibilities:
• Provide advanced production support for critical business applications, ensuring high availability and performance.
• Lead incident and change management processes, including root cause analysis and resolution of production issues.
• Monitor application health using tools such as SPLUNK, Grafana, and AppDynamics.
• Collaborate with development, infrastructure, and business teams to drive continuous improvement in system operations.
• Maintain and optimize Linux/Unix environments and SQL/Oracle databases.
• Document operational procedures, troubleshooting steps, and best practices.
Required Skills:
• Strong experience with Linux/Unix system administration.
• Advanced proficiency in SQL and Oracle database management.
• Expertise in incident and change management within enterprise environments.
• Proven ability to analyse and resolve production issues efficiently.
• Hands-on experience with monitoring and alerting tools: SPLUNK, Grafana, AppDynamics.
• Excellent communication and collaboration skills.
• Experience with cloud platforms (AWS, Azure, GCP)
Desired Candidate Profile
• 6+ years of experience in Application Production Support and Systems Operations.
• Proven experience managing mission-critical production environments.
• Strong expertise in Linux/Unix, SQL, Oracle Database, and monitoring tools.
• Demonstrated success in incident resolution, RCA preparation, and service improvement initiatives.
• Ability to work effectively in a fast-paced 16x7 production support environment.
• Exposure to AI tools and implementation as well

About VY SYSTEMS PRIVATE LIMITED
About
Vy Systems is a Global Technology consulting, Solutions, and Managed Technology Services company. We service our customers with ‘RESPONSIVENESS’ as a key factor and we believe that timely response to any transaction increases the operational efficiency and accelerates the revenue and profitability to our customers.
The Company is founded and managed by a team of professionals having more than two+ decades of global experience in the business of Technology Consulting and Services.
Tech stack
Similar jobs (10)
Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting
WFO-Immediate
8 to 12 Yrs
Bangalore/Hyderabad
Role Summary:
We are looking for an experienced Application Production Support Engineer with strong expertise in application support, incident and change management, Linux/Unix, SQL, Oracle, and monitoring tools. The candidate will be responsible for maintaining application availability, troubleshooting production issues, monitoring system performance, and coordinating with technical and business stakeholders.
Key Responsibilities
- Provide L2/L3 production support for business-critical applications.
- Monitor applications and infrastructure using Splunk, Grafana, and AppDynamics.
- Analyze and resolve production incidents within defined SLAs.
- Perform incident, problem, change, and service request management.
- Troubleshoot application issues across Linux/Unix, SQL, and Oracle environments.
- Perform SQL queries and database-level troubleshooting to identify application issues.
- Analyze application logs, alerts, and performance metrics to identify root causes.
- Coordinate with development, database, infrastructure, and other technical teams for issue resolution.
- Participate in Root Cause Analysis (RCA) and implement corrective/preventive actions.
- Support application deployments, releases, and production changes.
- Ensure effective communication with business users and stakeholders during critical incidents.
- Identify recurring issues and drive problem management and service improvement initiatives.
- Maintain support documentation, knowledge articles, and operational procedures.
- Participate in on-call/shift support as required.
Mandatory Skills
- 6+ years of experience in Application Production Support.
- Strong experience in Incident & Change Management.
- Hands-on experience with Linux/Unix.
- Good knowledge of SQL and Oracle database support.
- Experience with monitoring and observability tools:
- Splunk
- Grafana
- AppDynamics
- Strong troubleshooting and problem-solving skills.
- Good understanding of application monitoring, logs, alerts, and performance analysis.
- Strong stakeholder management and communication skills
We are looking for an experienced Application Support Engineer with strong expertise in Linux/Unix, Networking, Routing, Load Balancing, and Production Support. The ideal candidate should have hands-on experience troubleshooting application and infrastructure issues using enterprise monitoring and observability tools such as Splunk, Grafana, and AppDynamics.
Key Responsibilities
- Provide L2/L3 production support for enterprise applications and infrastructure.
- Troubleshoot application, Linux/Unix, network, routing, and connectivity-related issues.
- Monitor application and infrastructure health using Splunk, Grafana, and AppDynamics.
- Analyze alerts, logs, performance metrics, and system behavior to identify issues.
- Troubleshoot routing, load balancing, connectivity, and network-related problems.
- Perform incident investigation, troubleshooting, and root cause analysis.
- Coordinate with Network, Infrastructure, Application, and other technical teams for issue resolution.
- Handle production incidents and ensure timely resolution within defined SLAs.
- Participate in problem management and identify recurring issues.
- Perform application health checks and proactively identify potential failures.
- Maintain troubleshooting guides, runbooks, and operational documentation.
- Support planned changes, deployments, and maintenance activities.
Must-Have Skills
Application Support
Linux / Unix
Networking
Routing
Load Balancing
Splunk
Grafana
AppDynamics
Production Support
Incident Management
Troubleshooting
Required Skills & Tools
- 4–6 years of experience in Application / Production Support.
- Strong hands-on experience with Linux/Unix administration and troubleshooting.
- Good understanding of TCP/IP, networking, routing, and connectivity concepts.
- Practical experience troubleshooting load balancing issues.
- Hands-on experience with monitoring and observability tools such as Splunk, Grafana, or AppDynamics.
- Strong log analysis and troubleshooting skills.
- Experience handling production incidents and working within SLA-driven environments.
- Strong communication and coordination skills.
Preferred Experience
- Exposure to ITIL processes including Incident, Problem, and Change Management.
- Experience with scripting/automation using Shell or Python.
- Exposure to cloud environments is an added advantage.
- Experience supporting enterprise-scale applications and infrastructure.
Key Competencies
- Strong analytical and troubleshooting skills.
- Ability to work under pressure during critical production incidents.
- Strong ownership and accountability.
- Excellent communication and stakeholder management.
- Ability to work collaboratively with global technical teams.
Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
About the Role
We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.
Key Responsibilities
- Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
- Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
- Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
- Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
- Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
- Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
- Participate in incident/severity calls, ensuring clear communication and coordination
- Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting
Required Skills & Experience
- Strong understanding of Linux/Unix systems for application support
- Hands-on experience troubleshooting applications in staging and production environments
- Ability to monitor system performance and identify root causes using logs and metrics
- Experience working with Kubernetes and microservices-based architectures
- Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
- Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
- Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
- Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
- Experience with ticketing and documentation tools such as Jira and Confluence
- Minimum 4+ years of experience in application support or reliability engineering
Education & Certifications
- Bachelor's degree in Computer Science, Information Technology, or a related field
- Relevant certifications (Cloud, Kubernetes, Microservices) are a plus
Work Schedule
- Willingness to work in a 24x7 environment, including weekends and on-call rotations
SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.
Core responsibilities include:
- Production monitoring and debugging of live systems.
- Incident investigation, troubleshooting, and problem resolution.
- AWS cloud infrastructure support and maintenance.
- Deployment and operational support activities.
- Supporting a 24x7 production environment.
- Working with GitHub-based development workflows.
- Technical debt remediation and platform improvements.
- Customer issue investigation and support.
- Security and compliance-related work, including FedRAMP initiatives.
Preferred Skills:
AWS (especially S3 and EC2)
Strong debugging and troubleshooting skills
Site Reliability Engineering (SRE) experience
GitHub experience
Basic software development skills
TypeScript/JavaScript knowledge
C# preferred
AI experience is a plus.
Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Job Title: TechOps Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 6 Month-2 years (Excluding Internship)
Shift Timing: 2 PM to 11 PM IST
Role Overview
We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. Engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.
What you’ll do:
- Ensure system reliability, uptime, and performance of global platform.
- Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule)
- Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
- Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
- End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
- Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
- Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
- Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
- Collaborate with development, operations and support teams to build scalable and resilient systems.
- Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
- Participate in capacity planning, performance tuning, and resource optimization.
- Integrate security and compliance best practices into all infrastructure operations.
- Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
- Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
- Flexible to resolve blocking issues during off hours or weekends if required.
What We’re Looking For:
Basic Qualifications and skills
- Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
- Overall 1 years of experience as a Site Reliability Engineer, Technical project coordinator role.
- Proven experience of SRE or Technical Project Coordination with IoT or connected devices based platforms.
- Experience with incident management and on-call best practices. Provide support to on call engineers.
- Excellent analytical and problem-solving skills with a proactive mindset.
- Hands-on experience with AWS Cloud and IaC tools such as Terraform or Ansible.
- Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
- Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
- Fluency in English (spoken and written).
- Successfully recommission or decommission chargers following changes in our network.
- Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.
Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
- Linux troubleshooting
- Hands-on AWS
- Production/Application Support
- Bash/Shell/Python
- Monitoring/log analysis
- Incident resolution
- Application deployment/support
- Basic networking and database knowledge
- Production/batch support exposure
- Willingness for rotational weekend/critical production support
Role Summary
We are seeking a proactive and technically skilled Python Application Support Engineer to join our Technical Operations team. This role is crucial for ensuring the stability and reliability of our mission-critical, Python-based applications. You will be responsible for timely incident resolution, deep-dive troubleshooting, implementing permanent fixes, and driving operational efficiency through automation.
🔑 Key Responsibilities
Technical Troubleshooting & Incident Management
- Incident Resolution: Serve as the primary point of contact for complex Level 2 and Level 3 production incidents, diagnosing root causes and resolving issues across our Python application stack.
- Deep-Dive Analysis: Utilize log analysis tools (e.g., Splunk, ELK Stack) and monitoring platforms (e.g., Prometheus, Grafana) to quickly identify and address anomalies in application behavior.
- Code Debugging: Analyze, debug, and fix application issues directly within the Python codebase, including Flask/Django services, worker queues, and custom scripts.
- Database Health: Troubleshoot performance issues and conduct basic SQL/NoSQL query tuning and health checks (e.g., for PostgreSQL, MongoDB, or Redis).
Operational Excellence & Automation
- Monitoring & Alerting: Continuously refine and optimize application monitoring, alerting, and logging configurations to improve mean time to detect (MTTD) and mean time to resolve (MTTR).
- Python Automation: Develop, maintain, and enhance automated scripts (primarily in Python) to streamline routine operational tasks, reporting, health checks, and system recovery processes.
- Documentation: Create and maintain comprehensive documentation, runbooks, and knowledge base articles for application support procedures and recurring issues.
Collaboration & Prevention
- Cross-Functional Fixes: Collaborate closely with the Development and DevOps teams to provide clear technical feedback on recurring issues and implement permanent, scalable solutions.
- Proactive Maintenance: Identify potential system bottlenecks, performance degradation points, and areas prone to failure, recommending and implementing preventative measures.
⚙️ Required Qualifications
- Experience: 3 to 5 years of professional experience in Application Support, Production Support, Site Reliability Engineering (SRE), or a similar technical role.
- Python Expertise (Mandatory): Strong hands-on experience with Python scripting and programming, including the ability to read, debug, and modify application code.
- Operating Systems: Proficient working knowledge of Linux/Unix environments and shell scripting.
- Databases: Solid experience with relational (e.g., PostgreSQL, MySQL) and/or NoSQL (e.g., MongoDB, Redis) databases, focusing on query analysis and performance.
We are seeking a skilled System Administrator and DevOps Engineer to join AmpereHour Energy, where you will play a crucial role in making sure our software systems run without downtime and our infrastructure scales well. This position involves working on cloud and physical servers across multiple operating systems and fundamental understanding of computers and networking.
At AmpereHour Energy, we are dedicated to advancing sustainable energy solutions. Our team thrives on collaboration and innovation, driving impactful projects that contribute to a greener future. We value expertise and creativity, making it an exciting place for tech professionals to grow.
If you are interested in this opportunity and meet the qualifications outlined in the job description, we encourage you to apply and explore how you can contribute to our mission.
Hiring Platform Engineer
Exp: 6 -- 10 yrs
Edu : BE/B.tech/MCA
Work Location : Pune
Skills :
Platform monitoring ,Incident trouble shooting, Incident recovery, openshift ,kubernetes.
2 years of IT operations, infrastructure, cloud or application support experience.
Exp in Linux command-line knowledge.
Exp in networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.





