Cloud Production Support Engineer - Level 2 at Porter.in · Bengaluru (Bangalore) · 3 - 5 years · ₹8L - ₹12L / yr · Profitable · Posted 12 Oct 2023
Job Summary
Cloud Production Support Engineer(PSE) is responsible for fulfilling the day-to-day infrastructure and service requests from the application teams across AWS, CI/CD solutions and observability tools. You will be expected to handle production issues in collaboration with the cloud Infrastructure and application teams.
Responsibilities and Duties
- Troubleshoot production Issues: When technical issues with the cloud infrastructure components arise, PSE must act quickly to analyse the available data and find the root cause of the problem. They may then develop a solution or escalate the problem to other engineering team members while providing stakeholders with progress updates.
- Infrastructure provisioning and modification: Application teams may request to create new infrastructure or modify the existing ones in AWS based on their requirements via the ticketing tool. PSE should ensure that the required data/info is available on the ticket and provide a resolution based on the given SLA.
- Alert Management: Alerts from the observability tools will be received on multiple channels according to the notification settings. PSEs are expected to acknowledge the alerts, troubleshoot the issue, close the alert based on the given SLA, or escalate to the cloud infra/DevOps team for further diagnosis.
- Onboarding, Off-boarding and access management: Whenever an employee joins or leaves the organization, you will receive an onboarding or offboarding request.
- Prepare Technical Documentation: PSEs must prepare documentation when logging product issues, as they must note all details, including their observations, diagnoses, and action steps. Other everyday tasks include weekly reports summarising production performance, upgrade release notes, and troubleshooting guides.
- Product Improvements: Since PSEs have good exposure to the product issues, they should work closely with the PMs+EMs, pass the feedback on the product, and get the improvements/fixes included in the product roadmap.
- Adherence to SLA and timelines: PSEs should always adhere to the timelines shared with other teams for closure of fixes and deliver outcomes as per the SLA guidance agreed with business teams
- Reporting: Report & track weekly regarding SLA metrics, tickets being worked and closed by PSEs/transferred tickets. Identify and devise how productivity can be captured at the individual level and report the same monthly.
Qualifications and Skills
- Degree in Computer Science/Information Technology.
- Two years or more experience in Cloud and system administration.
- Experience troubleshooting in complex environments using monitoring tools.
- Demonstrated experience with containerisation technologies (Docker, Kubernetes, etc.)
- Hands-on experience with the most common AWS services.

About Porter.in
About
Company Overview:
At Porter, we are passionate about improving productivity. We want to help businesses, large and small, optimize their last-mile operations and empower them to unleash the growth of their core functions. Last-mile delivery logistics is one of the biggest and fastest-growing sectors of the economy with a market cap upwards of 50 billion USD and a growth rate exceeding 15% CAGR.
Porter is the fastest-growing leader in this sector with operations in 14 major cities, a fleet size exceeding 1L registered and 50k active driver-partners and a customer base with 3.5M being monthly active. Our industry-best technology platform has raised over 50 million USD from investors including Sequoia Capital, Kae Capital, Mahindra Group and LGT Aspada. We are addressing a massive problem and going after a huge market. We’re trying to create a household name in transportation and our ambition is to disrupt all facets of last-mile logistics including warehousing and LTL transportation. At Porter, we’re here to do the best work of our lives. If you want to do the same and love the challenges and opportunities of a fast-paced work environment, then we believe Porter is the right place for you.
Company URL: https://porter.in
Connect with the team
Similar jobs (6)
Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting
WFO-Immediate
8 to 12 Yrs
Bangalore/Hyderabad
Lead Cloud Reliability Engineer
Job Responsibilities
● Lead and manage the Cloud Reliability teams to provide strong Managed Services support to end-customers.
● Isolate, troubleshoot and resolve issues reported by CMS clients in their cloud environment
● Drive the communication with the customer providing details about the issue, current steps, next plan of action, ETA
● Gather client's requirements related to use of specic cloud services and provide assistance in seing them up and resolving issues
● Create SOPs and knowledge articles for use by the L1 teams to resolve common issues
● Identify recurring issues, perform root cause analysis and propose/implement preventive actions
● Follow change management procedure to identify, record and implement changes
● Plan and deploy OS, security patches in Windows/Linux environment and upgrade k8s clusters
● Identify the recurring manual activities and contribute to automation
● Provide technical guidance and educate team members on development and operations. Monitor metrics and develop ways to improve.
● System troubleshooting and problem-solving across plaorm and application domains. Ability to use a wide variety of open-source technologies and cloud services.
● Build, maintain, and monitor conguration standards.
● Ensuring critical system security through using best-in-class cloud security solutions.
Qualifications
● 4-7 years experience in Cloud Infrastructure and Operations domains and IT operational experience preferably in a global enterprise environment.
● Specialize in one or two cloud deployment platforms: AWS, GCP
● Hands on experience with AWS/GCP services (EKS, ECS, EC2, VPC, RDS, Lambda, GKE, Compute Engine)
● Understanding of one or more programming languages (Python, JavaScript, Ruby, Java, .Net)
● Logging and Monitoring tools (ELK, Stackdriver, CloudWatch)
● Knowledge on Conguration Management tools such as Ansible, Terraform, Puppet, Chef
● Experience working with deployment and orchestration technologies (such as Docker, Kubernetes, Mesos)
● Good analytical, communication, problem solving, and learning skills.
● Knowledge on programming against cloud plaorms such as Google Cloud Platform and lean development methodologies.
● Strong service aitude and a commitment to quality.
● Willingness to work in shifts.
- Linux troubleshooting
- Hands-on AWS
- Production/Application Support
- Bash/Shell/Python
- Monitoring/log analysis
- Incident resolution
- Application deployment/support
- Basic networking and database knowledge
- Production/batch support exposure
- Willingness for rotational weekend/critical production support
Lead Cloud Reliability Engineer
Job Responsibilities
● Lead and manage the Cloud Reliability teams to provide strong Managed Services support to end-customers.
● Isolate, troubleshoot and resolve issues reported by CMS clients in their cloud environment
● Drive the communication with the customer providing details about the issue, current steps, next plan of action, ETA
● Gather client's requirements related to use of specic cloud services and provide assistance in seing them up and resolving issues
● Create SOPs and knowledge articles for use by the L1 teams to resolve common issues
● Identify recurring issues, perform root cause analysis and propose/implement preventive actions
● Follow change management procedure to identify, record and implement changes
● Plan and deploy OS, security patches in Windows/Linux environment and upgrade k8s clusters
● Identify the recurring manual activities and contribute to automation
● Provide technical guidance and educate team members on development and operations. Monitor metrics and develop ways to improve.
● System troubleshooting and problem-solving across plaorm and application domains. Ability to use a wide variety of open-source technologies and cloud services.
● Build, maintain, and monitor conguration standards.
● Ensuring critical system security through using best-in-class cloud security solutions.
Qualifications
● 4-7 years experience in Cloud Infrastructure and Operations domains and IT operational experience preferably in a global enterprise environment.
● Specialize in one or two cloud deployment platforms: AWS, GCP
● Hands on experience with AWS/GCP services (EKS, ECS, EC2, VPC, RDS, Lambda, GKE, Compute Engine)
● Understanding of one or more programming languages (Python, JavaScript, Ruby, Java, .Net)
● Logging and Monitoring tools (ELK, Stackdriver, CloudWatch)
● Knowledge on Conguration Management tools such as Ansible, Terraform, Puppet, Chef
● Experience working with deployment and orchestration technologies (such as Docker, Kubernetes, Mesos)
● Good analytical, communication, problem solving, and learning skills.
● Knowledge on programming against cloud plaorms such as Google Cloud Platform and lean development methodologies.
● Strong service aitude and a commitment to quality.
● Willingness to work in shifts.
Read less
Role Summary:
We are looking for an experienced Application Production Support Engineer with strong expertise in application support, incident and change management, Linux/Unix, SQL, Oracle, and monitoring tools. The candidate will be responsible for maintaining application availability, troubleshooting production issues, monitoring system performance, and coordinating with technical and business stakeholders.
Key Responsibilities
- Provide L2/L3 production support for business-critical applications.
- Monitor applications and infrastructure using Splunk, Grafana, and AppDynamics.
- Analyze and resolve production incidents within defined SLAs.
- Perform incident, problem, change, and service request management.
- Troubleshoot application issues across Linux/Unix, SQL, and Oracle environments.
- Perform SQL queries and database-level troubleshooting to identify application issues.
- Analyze application logs, alerts, and performance metrics to identify root causes.
- Coordinate with development, database, infrastructure, and other technical teams for issue resolution.
- Participate in Root Cause Analysis (RCA) and implement corrective/preventive actions.
- Support application deployments, releases, and production changes.
- Ensure effective communication with business users and stakeholders during critical incidents.
- Identify recurring issues and drive problem management and service improvement initiatives.
- Maintain support documentation, knowledge articles, and operational procedures.
- Participate in on-call/shift support as required.
Mandatory Skills
- 6+ years of experience in Application Production Support.
- Strong experience in Incident & Change Management.
- Hands-on experience with Linux/Unix.
- Good knowledge of SQL and Oracle database support.
- Experience with monitoring and observability tools:
- Splunk
- Grafana
- AppDynamics
- Strong troubleshooting and problem-solving skills.
- Good understanding of application monitoring, logs, alerts, and performance analysis.
- Strong stakeholder management and communication skills
Platform Engineer
Location: Hyderabad, Telangana — On-site
Experience: 3–6 Years
Employment Type: Full-time
Hyderabad-based mobility-tech startup building the technology infrastructure behind student transportation.
We operate a real-time transportation platform that brings together student tracking, routing, parent notifications, driver applications, operations dashboards, and cloud infrastructure to make student mobility safer, more reliable, and easier to manage.
As we scale, we’re looking for a Platform Engineer who can take ownership of the infrastructure and platform layer that powers these systems.
The Role
As a Platform Engineer, you will own the systems that enable our engineering teams to build, deploy, scale, monitor, and operate reliable production services.
This is an early-stage startup role with significant ownership. You will work closely with engineering and product teams to build infrastructure from the ground up, improve deployment velocity, strengthen reliability, and ensure our platform can scale with the business.
You should be comfortable moving between cloud infrastructure, Kubernetes, CI/CD, observability, security, and distributed systems.
What You'll Do
- Design, build, and maintain scalable cloud infrastructure on AWS
- Own production infrastructure across EC2, IAM, RDS, networking, monitoring, and deployments
- Build and maintain Docker and Kubernetes environments for production workloads
- Develop and improve CI/CD pipelines for reliable and rapid deployments
- Manage infrastructure as code using Terraform
- Establish infrastructure standards for scalability, security, reliability, and cost efficiency
- Monitor production systems and proactively identify performance and reliability issues
- Build observability around applications and infrastructure, including metrics, logs, alerts, and incident monitoring
- Troubleshoot production issues and conduct root-cause analysis (RCA)
- Work with backend engineers to design infrastructure for Java/Kotlin-based distributed systems
- Support highly available services involving real-time tracking, routing, notifications, and operational workflows
- Improve deployment processes, release reliability, rollback strategies, and disaster recovery
- Identify infrastructure bottlenecks and continuously improve platform performance
- Help establish engineering practices around reliability, security, and operational excellence
- Work across the stack when required and take end-to-end ownership of infrastructure problems
What We're Looking For
Must Have
- 3–6 years of experience in Platform Engineering, DevOps, SRE, Infrastructure Engineering, or a closely related role
- Strong hands-on experience with AWS
- Strong understanding of EC2, IAM, RDS, networking, monitoring, and production deployments
- Experience with Docker and Kubernetes
- Strong experience building and managing CI/CD pipelines
- Hands-on experience with Terraform / Infrastructure as Code
- Strong Linux and networking fundamentals
- Experience troubleshooting production systems and performing RCA
- Understanding of distributed systems, scalability, availability, and system design
- Experience working with backend services built using Java/Kotlin or similar technologies
- Ability to independently own infrastructure problems from design → implementation → deployment → monitoring
Good to Have
- Experience with Redis, PostgreSQL, or MongoDB
- Experience with microservices architecture
- Experience with AWS security and IAM best practices
- Experience building observability and alerting systems
- Experience with multi-cloud environments such as AWS, Azure, or GCP
- Experience working in an early-stage startup
- Experience with real-time systems, IoT, location services, or mobility platforms
- Open-source contributions or meaningful personal engineering projects
What Makes This Role Different
At ZeroMoblt, you won't be working within a large infrastructure team where responsibilities are narrowly defined.
You'll have the opportunity to:
- Own critical infrastructure decisions
- Build platform capabilities from 0 → 1
- Work directly with engineering and product teams
- Solve real-world scalability and reliability problems
- Influence architecture and engineering practices
- See your work directly impact a platform serving 10,000+ students
- Work in a fast-moving environment with minimal bureaucracy
We're looking for someone who enjoys ownership, ambiguity, and solving problems independently.
This role may not be the right fit if you prefer highly structured processes, narrowly defined responsibilities, or large-company environments with multiple layers of ownership.
Why ZeroMoblt?
- High ownership and autonomy
- Direct exposure to product and engineering decisions
- Opportunity to build infrastructure at an early-stage mobility startup
- Work on real-time transportation and location-based systems
- Hyderabad-based, on-site team







