
Site Reliability Engineer (SRE)
Job Title: Site Reliability Engineer (SRE) – AI & Cloud Infrastructure
Location: Pune (Work From Office)
Experience: 5–8 Years
Employment Type: Full-Time
About the Role
We are looking for an experienced Site Reliability Engineer (SRE) to build and scale AI-powered reliability capabilities from the ground up. In this role, you will drive modern observability, automation, and cloud reliability initiatives while leveraging AI/ML for incident management, forecasting, and infrastructure optimization.
You will own the end-to-end reliability strategy across cloud-native AWS environments, enabling high availability, performance, and operational excellence through automation, intelligent monitoring, and proactive engineering.
Key Responsibilities
- Design, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS.
- Build and maintain an end-to-end observability platform using Open Telemetry, Grafana, Datadog, CloudWatch, and related tools.
- Implement AIOps capabilities, including:
- LLM-assisted incident triage
- AI-powered root cause analysis
- ML-driven forecasting and anomaly detection
- Intelligent alert correlation and noise reduction
- Lead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA).
- Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines.
- Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization.
- Build dashboards, SLOs, SLIs, and error budgets to improve service reliability.
- Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks.
- Monitor application and infrastructure health while ensuring high uptime and service performance.
- Collaborate with Development, DevOps, Security, Platform Engineering, and Product teams across multiple time zones.
- Establish operational best practices for monitoring, incident response, disaster recovery, and resilience engineering.
- Maintain Linux-based production systems and troubleshoot OS, networking, storage, and performance issues.
Required Skills & Qualifications
- 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering.
- Strong experience with AWS services including EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3, and Auto Scaling.
- Hands-on experience with Open Telemetry, Grafana, Datadog, Prometheus, or similar monitoring platforms.
- Strong knowledge of Linux administration, networking, system performance tuning, and troubleshooting.
- Experience with Infrastructure as Code using Terraform or CloudFormation.
- Proficiency in scripting using Python, Bash, or Go.
- Experience with Kubernetes and containerized workloads.
- Strong understanding of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.).
- Experience leading incident management, production support, and RCA processes.
- Knowledge of SRE principles including SLIs, SLOs, and Error Budgets.
- Experience implementing monitoring, logging, alerting, and observability frameworks.
- Strong analytical, troubleshooting, and communication skills.
Preferred Qualifications
- Experience building or implementing AIOps solutions.
- Exposure to Large Language Models (LLMs) for operational automation.
- Experience with machine learning-based forecasting or anomaly detection.
- Hands-on experience administering Adobe Experience Manager (AEM).
- Experience managing Cloudflare CDN, WAF, DNS, and caching strategies.
- Knowledge of FinOps, cloud cost optimization, and capacity planning.
- AWS Solutions Architect, DevOps Engineer, or Kubernetes certifications are a plus.

About EmbarkingOnVoyage Digital Solutions Pvt Ltd
About
Welcome to EmbarkingOnVoyage Digital Solutions: Your leading data and product engineering partner. We empower Technology companies/ISVs, large enterprises, and SMBs to revolutionise technology through innovative solutions in travel, healthcare, banking, finance, and beyond. We've helped countless partners optimize processes, maximize profitability, and achieve their goals. Our dedicated team, driven by ultimate customer satisfaction, delivers expertise in competitive product development, app modernization, UI/UX, data analytics, and more. Led by industry leaders, we seamlessly bridge the gap between services and products, leveraging modern technology to optimize results. Our close-knit team thrives on collaboration, creativity, and mutual motivation, propelling each other to excel and delivering exceptional results reflected in every project. Fuelled by passion for technology and a deep understanding of modern challenges, we bring out the best solutions for our partners. We're more than just a service provider; we're your reliable partner in navigating the ever-evolving digital landscape.
Similar jobs
About the Role :
As the Lead AI/ML Engineer, you will be responsible for leading the design, development, deployment, and continuous improvement of AI-powered solutions that create measurable impact in education. This role combines hands-on technical execution with engineering leadership, requiring you to architect scalable AI systems, mentor engineers, establish best practices, and collaborate closely with cross-functional teams to deliver production-grade solutions.
You will take ownership of the complete AI development lifecycle—from problem definition and data pipeline design to model training, deployment, monitoring, and optimization—while ensuring high standards of reliability, scalability, security, and performance. The role also involves evaluating emerging AI technologies and driving their adoption to enhance product capabilities and user experience.
Roles and Responsibilities:
- Lead the design, development, and deployment of AI-powered products for production environments.
- Drive technical architecture and decision-making across machine learning models, inference systems, cloud infrastructure, and distributed computing.
- Develop and maintain end-to-end machine learning pipelines, including data ingestion, model training, evaluation, deployment, monitoring, and continuous optimization.
- Design scalable, secure, and high-performance AI services capable of handling production-scale workloads.
- Provide technical leadership by mentoring engineers, conducting code and design reviews, and fostering engineering excellence.
- Establish and promote best practices for software quality, testing, performance optimization, observability, reliability, and maintainability.
- Collaborate closely with Product, Backend, Mobile, and Data teams to deliver robust, scalable, and business-focused AI solutions.
- Research, evaluate, and implement emerging AI technologies to enhance product capabilities and drive innovation.
Required Skills & Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, or a related technical discipline.
- Minimum 4 years of hands-on experience developing, deploying, and maintaining production-grade AI/ML systems.
- Strong proficiency in Python with excellent software engineering, object-oriented programming, and system design skills.
- Extensive experience with modern machine learning frameworks such as PyTorch, Hugging Face, TensorFlow, or equivalent.
- Proven expertise in deploying, optimizing, and scaling large language models (LLMs) and other AI models for production inference.
- Strong understanding of containerization and cloud-native technologies, including Docker, Kubernetes, AWS, and distributed system architectures.
- Demonstrated ability to lead technical initiatives, mentor engineering teams, and make sound architectural decisions.
- Excellent analytical, debugging, communication, and problem-solving skills.
Preferred Qualifications
- Hands-on experience with Large Language Models (LLMs), Computer Vision, Speech Recognition, Generative AI, or Multimodal AI systems.
- Experience with model optimization and inference technologies such as Quantization, Distributed Training, vLLM, TensorRT, ONNX Runtime, or similar frameworks.
- Prior experience developing AI solutions for the education, public sector, or other large-scale social impact initiatives.
Details:
- Location: Kothrud, Pune.
- Interested candidates should fill in the application on career website https://careers.vopa.in/
- Salary :- 12 LPA to 14 LPA (Depending on last drawn salary, Interview and Skills)
Job Title: Software Engineer Consultant/Expert 34192
Location: Chennai
Work Type: Onsite
Notice Period: Immediate Joiners only or serving candidates upto 30 days.
Position Description:
- Candidate with strong Python experience.
- Full Stack Development in GCP End to End Deployment/ ML Ops Software Engineer with hands-on n both front end, back end and ML Ops
- This is a Tech Anchor role.
Experience Required:
- 7 Plus Years
Should have aldready worked as Lead and mentored and managed Junior Team Members
4Yrs + on MEAN Stack or if total Experience more than 6Yrs; 3years + Could be MEAN
Worked in Product Based COmpanies
Deisgn,Develop, Debug ,Deploy! All Individually
Micro Service Architecture
Experience in FIn-tech/Finance based companies are a plus!
AWS,Google CLoud & SOcket Programming
Worked on Angular 4+ (5,6,7,8)
Our Tech stack:
· Java 11, Spring Boot, React.JS
· MySQL, Redis, Elastic Stack, MongoDB
· Docker, K8s
· GitHub, Jenkins
· Slack, Jira
What we're looking for:
· Minimum of 5 years of experience as a software developer.
· Excellent knowledge of Java.
· Demonstrated skills in web-based development, including REST API s and JS frameworks such as Vue JS/React.
· Experienced with cloud platforms (AWS / GCP).
· Experienced with SQL and NoSQL databases.
· Experienced with CI/CD environments.
· Familiar with container technologies (Docker, Kubernetes).
· Proficient in agile programming practices.
· Strong analytical mind with a good dose of creativity.
· Meticulous person who strives to constantly improve his/her/their competencies.
· Excellent communication skills.
Nice-to-haves:
· Familiarity with Web sockets.
· Has contributed to open source projects and can point us to his/her/their GitHub account.
- Roles and Responsibilities:
Hands-on experience on Angular, CSS, Scripting, NodeJs/express
Experience in Responsive Web, Automation, CI/CD, Github, Microservices, Postgres (or any other RDBMS)
Experience in AWS (SQS, SNS, Cognito)
Implementing Observability/Monitoring.
Strong experience in REST APIs
The ideal candidate will be responsible for developing high-quality applications. They will also be responsible for designing and implementing testable and scalable code.
Responsibilities :
-
Work with development teams and product managers to ideate software solutions
-
Maintenance of Node.js Backend
-
Working with MongoDB to create various features
-
Troubleshoot, debug and upgrade software
-
Create security and data protection settings Requirements
-
Proven experience as a backend developer in Node.js or similar role
-
Experience developing desktop and mobile applications
-
Familiarity with common stacks
Requirements :
-
Hands on experience building end to end systems
-
Minimum 1 yrs of experience with Javascript, Node.js and Mongo.DB
-
Good architectural & design skills
-
Strong coding, data structures and algorithms
-
The ability to own end to end responsibility - right from requirement to release
-
The ability to produce bug-free and production grade code
-
Experience and fine understanding of cross browser front end development issues
-
Exhibit a deep understanding of server virtualization, networking and storage ensuring that the solution scales and performs with high availability and uptime
Job Description
- Planning with the product advancement group to finish item thoughts
- Utilizing different programming to plan the items. Has hands on coding experience across different technology stacks. Willing to experiment & explore across technologies and identify the best suited
- Performing constant market analysis of competing products
- Testing item models to control configuration blemishes
- Talking with delivery team to work out financially savvy producing methodology and solutions.
- Ability to ideate, code, setup a prototype and on approval design the entire application as per business requirements.
- Complete knowledge of Full stack development. .
- As a product engineer, our Super Hero is expected to participate in all parts of development life cycle. Create user friendly, cost effective product design.
- Ability to research w.r.t features, UX, functionalities and finally testing before being rolled out.
Job Requirements
- Familiarity with all stages of product development life cycle.
- Extensive experience in writing codes, deep understanding of platform architecture, ability to design & develop solutions & knack towards problem solving.
- Deep engineering skills in various technology stacks, ability to handle big data concepts, keen to work in IIOT & machine data domain.
- Ability to effectively present ideas & communicate with team members within & outside organization, ability to write routine reports and correspondence and ability to speak confidently & effectively before groups of customers or other Donaldson colleagues.
- We are happy to have a technology geek take this position as long as s/he is able to align the product roadmap to organization’s growth roadmap.
- Good communication and presentation skills
- Effective written and verbal skills in English Language are mandatory.
- This is a global role and the candidate may be required to travel to Smart Ship Hub’s different delivery centres.
Experience Requirements & Educational Qualification
- Atleast 5 years of proven product engineering experience. Must have: managed product life cycle, knowledge of product roadmap, extensive coding experience.
- Should be open to working in Disruptive Application Design domain, IIOT, Smart Sensors, Big data
Fullstack (Java with AngularJS or Node JS )
Primary Skill: JAVA,J2EE,AgularJS,WebServices
Secondary Skill: Post gres, Angual/React, AWS/GCP, Docker, API, Unix, Node JS, AWS, Docker, quartz, Shell script
Experience: 3 – 7 years
Notice: Immediate to 15 Days Max
Job Description:
Strong organizational and coding skills. Proficiency with fundamental front-end languages such as HTML, CSS and JavaScript. Hands-on experience and strong knowledge in
JavaScript frameworks such as Angular JS/React JS/Node JS and good coding skills in server-side development – Spring Framework, Spring Boot (latest 2.3.1), Core Java (Java 8 & above). Preferable skills in Microservices, Web services, Rest API and design patterns. Hands-on experience in Unix shell script. DevOps – Docker, Any Cloud Platform (AWS/GCP). Strong Integration knowledge in Java preferably in SAP interfacing - Good to have (for eg https://github.com/SAP/cloud-rest-api-client). Familiarity with database technology such as Hibernate/JPA (ORM), PostgreSQL. Good knowledge of designing and developing APIs. Good problem-solving skills. Excellent verbal communication skills.
Full Stack Developer
Skills:
-
Proficiency with JavaScript and HTML5
-
Minimum 2+ years of hands-on experience with AngularJS and Angular Frameworks
-
Experience with Java, JSON, Spring Boot and Hibernate
-
Experience with MYSQL databases
-
Familiarity with Linux environments
-
Experience with GIT
-
Hands on experience with AWS S3 is preferred
-
Experience with web servers & application servers such as Apache and Nginx is good to have
Responsibilities:
-
Design and develop client-side and server-side architecture
-
Develop and manage well-functioning database
-
Implementation of approved user interface
-
Design and construction of REST APIs
-
Server management and deployment for the relevant environment
Are you someone who never intended to play it safe and wants to challenge the status quo in every nook and corner of your life? Do you want to make the maximum utilization of your skills and knowledge? Would you want every decision and move that you make should have a big impact on the company for the better (even if it doesn't, we're a bunch of people who will give you the room to fail and learn from it)? The read on...
We are a start-up into online personal styling & fashion.
We are looking for a core team member as a CTO or Head of Technology. It is much more than a job and the main goals in the first few months will be to build the Android App for the Company, & the Technology team & its strategy, manage systems, and its maintenance regularly. Knowledge of AI & ML is a must. As this product needs to be build AI/ML ready
Don't expect a high salary as this is more of a hands-on plus management role with the right equity that can be provided based upon your performance.









