Senior DevOps Engineer at Designing a generic ML platform as a product. · Bengaluru (Bangalore) · 4 - 8 years · ₹25L - ₹50L / yr · Posted 9 Jun 2022

Requirements
- 3+ years work experience writing clean production code
- Well versed with maintaining infrastructure as code (Terraform, Cloudformation etc). High proficiency with Terraform / Terragrunt is absolutely critical
- Experience of setting CI/CD pipelines from scratch
- Experience with AWS(EC2, ECS, RDS, Elastic Cache etc), AWS lambda, Kubernetes, Docker, ServiceMesh
- Experience with ETL pipelines, Bigdata infra
- Understanding of common security issues
Roles / Responsibilities:
- Write terraform modules for deploying different component of infrastructure in AWS like Kubernetes, RDS, Prometheus, Grafana, Static Website
- Configure networking, autoscaling. continuous deployment, security and multiple environments
- Make sure the infrastructure is SOC2, ISO 27001 and HIPAA compliant
- Automate all the steps to provide a seamless experience to developers.

Similar jobs (10)
We are looking for a hands-on Senior AWS Cloud Engineer to lead the infrastructure build, optimization, automation, and production deployment of a Multi-Agent AI Chatbot Platform hosted on AWS. The development environment is already in place, and the successful candidate will drive the solution through testing, integrations, and production go-live.
Key Responsibilities
- Review, validate, and optimize existing Terraform code and AWS infrastructure.
- Establish and manage integrations with enterprise platforms such as ServiceNow, Workday, and other third-party systems.
- Design, build, and support secure, scalable, and highly available AWS environments.
- Implement and automate CI/CD pipelines and Infrastructure-as-Code practices.
- Lead infrastructure testing, performance tuning, and production readiness activities.
- Drive deployment and operationalization of the platform in the Production environment.
- Implement cloud governance, security, monitoring, and FinOps best practices.
- Troubleshoot and resolve complex cloud infrastructure issues.
Required Skills & Experience
- 10+ years of IT experience with strong expertise in AWS Cloud Engineering.
- Proven experience designing, deploying, and managing AWS production environments.
- Strong hands-on experience with Terraform and Infrastructure-as-Code.
- Experience with CI/CD pipeline automation and DevOps practices.
- Expertise in AWS services including VPC, IAM, EC2, S3, Lambda, CloudWatch, and networking.
- Experience in performance optimization, reliability, and cloud cost management (FinOps).
- Strong scripting and automation skills.
- Experience integrating enterprise applications through APIs and secure connectivity patterns.
Job Title : DevOps Engineer / Site Reliability Engineer (SRE)
Experience : 5+ Years
Location : Gurugram, Haryana
Work Mode : On-site (Full-time)
About the Role :
We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.
Mandatory Skills :
AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux
Key Responsibilities :
- Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
- Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
- Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
- Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
- Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
- Automate operational tasks using Python, Bash, and Groovy scripting.
- Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.
Required Skills & Qualifications :
- Bachelor's degree in Computer Science, IT, Electronics, or a related field.
- 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
- Strong expertise in AWS, with exposure to Azure/GCP.
- Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
- Strong scripting skills in Python and Bash.
- Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
- Good understanding of Linux, networking, SQL, and cloud security best practices.
Preferred Skills :
- Experience with multi-cloud environments and DevSecOps practices.
- Knowledge of disaster recovery, automation, and microservices architecture.
- Strong troubleshooting, communication, and problem-solving skills.
This is a senior role in our Application & Database Modernization pillar, on a specific mission: leading a large-scale VMware-to-AWS migration and landing it in production. VMware estates are exactly where modernization programs go to stall — sprawling dependency graphs, undocumented workloads, and a hypervisor bill that grows while the migration deck gathers dust. Our customers need that estate moved, cut over, and retired, on a date in the contract.
You'll lead the design and implementation of the target infrastructure: Landing Zones built to AWS best practices, EKS-based container platforms, Infrastructure as Code across the stack with Terraform and CloudFormation, and HA, DR, security, and backup strategies that hold up in production. Aedeon absorbs the discovery, dependency-mapping, and validation grind that would otherwise consume the project's first two quarters. You own the judgment calls, architecture tradeoffs, cutover sequencing, risk decisions and you carry them through to production. You'll also mentor engineers and shape long-term infrastructure strategy with customer stakeholders. If you want your migration experience to end in retired VMware clusters rather than revised project plans, this is the role.
What you will do?
- Lead the design and implementation of infrastructure for a large-scale VMware-to-AWS migration — from discovery through production cutover.
- Architect and build secure, scalable, and highly available AWS environments, including Landing Zones that follow AWS best practices.
- Design and implement containerized application platforms on Amazon EKS; exposure to ECS is a plus.
- Implement Infrastructure as Code using Terraform and AWS CloudFormation.
- Define and enforce best practices across security, backups, high availability (HA), disaster recovery (DR), monitoring, and operations.
- Enable configuration management using tools such as Ansible, Chef, or similar.
- Build CI/CD pipelines and manage the complete build and release lifecycle for customer applications.
- Drive automation across provisioning, deployment, and operational workflows.
- Retire legacy infrastructure and land cloud-native architectures in its place.
- Improve and maintain DevOps platforms: Jenkins, Git repositories, monitoring, and observability stacks.
- Provide L3-level support for complex infrastructure and platform issues.
- Work with stakeholders on technical strategy and long-term architecture.
- Mentor engineers, contribute to talent evaluation, and support team development.
What are we looking for?
- 6+ years of hands-on DevOps experience, with strong expertise in designing and managing cloud infrastructure.
- Strong experience in VMware-to-AWS migration projects or large-scale infrastructure migrations.
- 4+ years of Terraform and CloudFormation for Infrastructure as Code (IaC).
- 4+ years in configuration management, systems engineering, and managing production-grade infrastructure.
- Solid Linux and/or Windows administration background.
- Deep understanding of AWS services, including VPC, EC2, IAM, EKS, ECS, RDS, S3, Backup, CloudWatch, etc.
- Experience with Kubernetes (EKS), ECS, and Docker in production environments.
- Hands-on experience designing HA, DR, security, and backup strategies.
- Experience with Landing Zone setup, multi-account strategy, and AWS governance frameworks.
- Proficiency in Git with a strong understanding of branching and merging strategies.
- Experience with CI/CD pipelines, automation, and operational tooling.
- A problem-solving mindset that's proactive and oriented toward reliability, performance, and scalability.
You will be preferred if
- Experience across multiple cloud platforms (AWS, Azure, GCP).
- AWS certifications such as Solutions Architect Associate/Professional or DevOps Engineer Professional.
- Exposure to data platforms such as Amazon EMR, Redshift, Lake Formation, and SageMaker.
- Experience designing cloud architectures at an L3/Architect level.
Job Title : SDE 3 – Infrastructure Platform Engineer
Experience : 5.5 to 8.5 Years
Number of Positions : 2
Employment Type : C2H (Contract to Hire)
Work Mode : Remote during contractual period → 5 Days WFO after conversion
Contract Duration : 3 Months
Post-Conversion Location : Pune
Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred
(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)
Role Overview :
We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.
The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.
Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.
Key Responsibilities :
- Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
- Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
- Manage and optimize Docker and Kubernetes workloads.
- Build internal platform tools to improve developer productivity and engineering efficiency.
- Implement monitoring, logging, alerting, and observability solutions.
- Participate in incident response, RCA, postmortems, and reliability improvements.
- Design and improve CI/CD pipelines and deployment automation.
- Contribute to system design, architecture discussions, scalability, security, and cost optimization.
- Collaborate with application, data, and product engineering teams.
Required Skills :
- 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
- Strong hands-on experience with AWS, GCP, or Azure.
- Strong expertise in Terraform / CloudFormation.
- Experience with CI/CD, Docker, and Kubernetes.
- Strong programming skills in at least one of:
- Python, Go, Java, or Ruby.
- Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
- Experience working with production infrastructure and highly available systems.
- Strong troubleshooting and problem-solving skills.
Nice to Have :
- Experience with SRE practices and production on-call ownership.
- Experience in fintech, payments, banking, or transaction-heavy systems.
- Knowledge of cloud security, compliance, or FinOps/cost optimization.
- Experience building internal developer platforms or productivity tools.
- Previous product company experience.
Interview Process :
Round 1 : Take-Home Coding Assignment – Submit within 48 hours
Round 2 : Coding Assignment Discussion – 1 Hour
Round 3 : Technical Managerial Round – 30 Minutes
Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.
Ideal Candidate :
Strong Platform / Infrastructure Engineer with hands-on experience in :
Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations
Pure DevOps profiles without strong coding and platform engineering experience are not preferred.
Job Title: DevOps Manager
Experience: 8+ Years
Location: Pune
Employment Type: Full-time
About the Role
We are looking for an experienced DevOps Manager with strong hands-on technical expertise and leadership capabilities to manage cloud infrastructure, automation, CI/CD, deployments, and platform reliability. The ideal candidate should have strong experience in Shell Scripting, cloud platforms, Kubernetes, Terraform, and production DevOps environments.
Key Responsibilities
• Lead and mentor the DevOps engineering team.
• Design, implement, and optimize CI/CD pipelines and deployment processes.
• Manage scalable and highly available cloud infrastructure across AWS / Azure / GCP.
• Drive Infrastructure as Code (IaC) using Terraform and configuration automation using Ansible.
• Manage Docker and Kubernetes based environments.
• Develop and maintain automation scripts using Shell/Bash Scripting and Python.
• Implement monitoring, logging, alerting, and observability.
• Handle production incidents, troubleshooting, root cause analysis, and preventive actions.
• Drive infrastructure security, performance, scalability, and cloud cost optimization.
• Collaborate with Development, QA, Security, Architecture, and Delivery teams.
Required Skills & Qualifications
• 8+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or related roles.
• Strong hands-on experience in Shell / Bash Scripting.
• Experience with AWS / Azure / GCP cloud platforms.
• Strong knowledge of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
• Hands-on experience with Docker, Kubernetes, and Helm.
• Strong experience with Terraform and Ansible.
• Experience with Linux administration, Git, networking, IAM, and cloud security.
• Experience with monitoring tools such as Prometheus, Grafana, ELK, or equivalent.
• Strong troubleshooting, communication, stakeholder management, and team leadership skills.
Preferred Skills (Nice to Have)
• Python scripting experience.
• Cloud / Kubernetes / Terraform certifications.
• Experience with SRE, observability, high-availability architecture, and cloud cost optimization.
Key Responsibilities
- Automate application deployments from Bitbucket to servers using CI/CD pipelines.
- Design and manage scalable, highly available AWS infrastructure.
- Implement Auto Scaling, ELB, and Route 53 for traffic management and high availability.
- Work with AWS services including IAM, RDS, DynamoDB, EC2, and other cloud services.
- Build and manage Docker containers and server images.
- Deploy and manage applications using Kubernetes.
- Implement Infrastructure as Code using Terraform, CloudFormation, or Ansible.
- Develop automation scripts using Python and Bash.
- Implement monitoring and logging using tools such as Prometheus, Grafana, and ELK.
- Integrate security and compliance practices into CI/CD pipelines.
- Optimize infrastructure for security, scalability, performance, and cost.
Required Skills
- 3+ years of experience in DevOps or a similar role.
- Strong knowledge of AWS beyond EC2.
- Hands-on experience with Jenkins or similar CI/CD tools.
- Experience with Docker and Kubernetes.
- Good understanding of Terraform/IaC and automation.
- Proficiency in Python and/or Bash scripting.
- Knowledge of DevSecOps, security, and compliance best practices.
- Strong troubleshooting and problem-solving skills.
About the Role
The non-negotiable is deep, hands-on infrastructure expertise spanning hybrid cloud and self-hosted systems. You will architect and maintain our unique infrastructure combining AWS CDN, bare metal servers, and Kubernetes clusters. This is not just maintenance work: you will establish company-wide DevOps policy, build security-hardened environments, and create the documentation and processes that scale with us. You will design our CI/CD pipelines, implement zero-trust networking, and guide the technical team on how to keep our production systems robustly online. This role is hands-on, autonomous, and sets the standard for how we approach infrastructure as the company grows.
What You'll Build
- Hybrid Infrastructure Management: Architect and maintain our unique infrastructure spanning AWS CDN, bare metal servers, and Kubernetes clusters for computationally intensive facial analysis workloads.
- CI/CD Pipeline Architecture: Design and implement CircleCI or Jenkins pipelines with comprehensive build testing, versioning, and change logging.
- Zero-Trust Networking: Build and maintain mesh topology networks using Tailscale or Wireguard to securely connect our hybrid infrastructure.
- Security-First Culture: Establish and enforce security policies including key rotation, access controls, compliance frameworks, and employee security management.
- Infrastructure as Code: Document and codify all infrastructure decisions, creating repeatable, auditable deployments.
- Containerization Strategy: Implement and optimize Docker/K8s deployments for our AI/ML workloads.
- Cost Optimization: Continue our approach of strategic compute placement using owned, rented or borrowed infrastructure where it makes financial sense without sacrificing security or reliability.
- Observability & Monitoring: Implement comprehensive logging, monitoring, and alerting across our distributed systems.
What We're Looking For
- 5+ years of DevOps or infrastructure engineering experience, with a track record of building from scratch
- Hybrid infrastructure expertise: experience managing both cloud (AWS) and self-hosted infrastructure, understanding the tradeoffs and security risks of each
- Kubernetes production experience: deep knowledge of cluster design, operations, and scaling
- Networking mastery: strong understanding of VPCs, mesh networks, VPNs, and zero-trust architectures
- Security-first mindset: experience with security compliance, key management, IAM policies, and hardening production systems
- CI/CD expertise: hands-on experience building robust pipelines for build testing before deployment
- Infrastructure as Code: proficiency with Terraform, Ansible, or similar tools
- Scripting and automation: strong Python, Bash, or Go skills for tooling and automation
- Policy and documentation: ability to establish best practices and document them clearly for team adoption
- Leadership mentality: comfortable setting standards and directing technical decisions, not just executing them
Nice to Have
- Experience architecting infrastructure for AI/ML workloads
- Background in a fast-moving startup or scale-up environment
- H ands-on experience with cost optimization across cloud and on-premises infrastructure
Why Join
- Opportunity to solve real healthcare problems with cutting-edge technology
- Well-funded startup with a strong market presence
- Work with advanced AI technology in a healthcare context
- Collaborate with a talented team in a fast-paced environment
- Competitive salary with equity options
- Performance and quarterly bonuses
- Professional development opportunities
Compensation and Logistics
- Remote, full-time
- Reports to: Head of Engineering
- Competitive based on experience
Job Description:
Pre-requisite skills required for a DevOps Engineer role include:
- 6+yrs exp in DevOps
- Experience working on Linux based infrastructure
- knowledge in AWS, docker, CI/CD tools
- Hands on exp in Python/shell scripting language
- hands on exp in AWS and Azure
- Work exp in Docker, Terraform, Ansible, Kubernetes, LINUX
- Excellent understanding of Ruby, Python, Perl, and Java
- Configuration and managing databases such as MySQL, Mongo
- Excellent troubleshooting
- Working knowledge of various tools, open-source technologies, and cloud services
Roles & Responsibilities
- Own end-to-end customer deployments across cloud, BYOC, air-gapped, and data-center environments
- Deploy and operate production environments using Kubernetes, Helm, and Terraform
- Troubleshoot complex Kubernetes, networking, infrastructure, application, and deployment issues
- Build and maintain observability, monitoring, alerting, and reliability
- Handle customer escalations and drive issues to resolution with a "fix first, optimize later" mindset
- Automate deployment and operational workflows
- Work closely with Engineering, Product, and Customer teams to resolve production challenges
- Participate in on-call and shift rotations, including critical customer escalations outside standard working hours
Ideal Candidate
1.Strong Hands-On DevOps Engineer Profile with deep Kubernetes and multi-cloud experience
2.Mandatory (Experience 1): Must have 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.
3.Mandatory (Experience 2): Must be able to own end-to-end customer deployments across cloud and enterprise environments (the role covers cloud, BYOC, air-gapped, and data-center deployments).
4.Mandatory (Tech skill 1): Must have strong hands-on Kubernetes experience — troubleshooting, networking, workloads, storage, RBAC, and security
5.Mandatory (Tech skill 2): Must have hands-on Helm experience — deployment, templating, and troubleshooting.
6.Mandatory (Tech skill 3): Must have hands-on Terraform experience — infrastructure provisioning and automation
7.Mandatory (Tech skill 4): Must have strong Linux and cloud infrastructure troubleshooting skills
8.Mandatory (Tech skill 5): Must have hands-on experience with observability — metrics, logs, traces, and alerting
9.Mandatory (Tech skill 6): Must have hands-on experience with one or more of AWS / GCP / Azure / OCI
10.Mandatory (Tech skill 7): Must have a strong incident-management and production-troubleshooting mindset, with a "fix first, optimize later" approach to customer escalations.
11.Mandatory (Company): B2B SaaS product companies
12.Preferred (Enterprise deployments): Prior experience with BYOC, air-gapped, or restricted-network environments.
13.Preferred (Observability tools): Prometheus / Grafana / OpenTelemetry / ELK / OpenSearch / Datadog
14.Preferred (Other): GitOps / CI-CD; Python or Golang for automation; multi-tenant SaaS infrastructure scaling; exposure to AI/ML pipeline deployments or iPaaS / reverse ETL connectors; SRE concepts (SLIs/SLOs, DR, high availability).
About Us
CLOUDSUFI is a Silicon Valley-based specialist Data Engineering & Cloud Technologies player with top-tier clients, favorable revenue mix, strong financial performance, and robust management. We pride ourselves in helping in the Data Discovery, Insights and Monetization for organizations. We offer quality of work, opportunities to learn new platforms/technologies that will help young engineers put themselves ahead in their careers compared to their peers in the IT Services industry. CLOUDSUFI is a Data Science and Product Engineering company building Products/Solutions for Technology and Enterprise industries leveraging the advent of Cloud Hyper Scalers and AI/ML, NLP technologies.
The organization is built to scale with strong external/ internal tech capabilities and governance standards. Started in 2019, CLOUDUSUFI is a family of 250 members working towards a common goal of making the enterprise data dance.
To know more, please visit https://cloudsufi.com
Our Values
We are a passionate and empathetic team that prioritizes human values. Our purpose is to elevate the quality of lives for our family, customers, partners and the community.
Equal Opportunity Statement
CLOUDSUFI is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified candidates receive consideration for employment without regard to race, colour, religion, gender, gender identity or expression, sexual orientation and national origin status. We provide equal opportunities in employment, advancement, and all other areas of our workplace. Please explore more at https://www.cloudsufi.com/
Role : Full-Time Individual Contributor (IC)
Reporting to : Solution Architect / Program Manager
Location : India Remote (Quarterly visits to Noida office)
Shift : 2PM-11PM IST
12x5 (On call duty)
Experience : 8-14 Years
ABOUT YOU
- 5+ years’ experience with AWS orchestration via Terraform script
- 5+ years’ experience with CloudWatch/CloudTrail/Guard Duty
- 5+ years’ experience with AWS WAF
- 4+ years’ experience with CloudFlare
- 3+ years’ experience with DataDog
- Experience with PagerDuty
- Ability to make nuanced threat assessments
- Experience in SOPHOS.
- Significant experience with PCI, SOC2, SOX, HIPAA, or other compliance regimes
- Experience in Infrastructure As Code – Ansible / Terraform/ CloudFormation
- Hands-on experience implementing various security tools in CI/CD pipeline
- Strong experience with any cloud service provider (AWS Preferred)
- Implement and oversee technological upgrades, improvements and major changes to the cloud security environment.
- Develop solutions, install/configure/integrate IT tools and security processes within an application or organization to help improve the overall IT security posture.
- Set up Static and Dynamic Code Analysis tools, review the results and explain any gaps and potential impact to the teams (development and operations).
- Penetration testing and container security.
- Evaluate and analyze threat, vulnerability, impact and risk to security issues discovered from security assessments.
- Assess current technology architecture for vulnerabilities, weaknesses and for possible upgrades or improvement
- Creating and managing security strategies
- Oversee information security audits, whether performed by organization or third-party personnel
- Develop, maintain and publish up-to-date information security policies, standards and guidelines.
- Preferred Certification - AWS Security/CISSP/CISM (Certified Information Security Manager)
ABOUT THE ROLE
- Work independently with vendors and collaborate with colleagues
- Experience negotiating remediation timelines and/or remediate found issues independently
- Ability to implement vendor platforms within CI/CD pipelines
- Experience managing/responding to incidents, collecting evidence, and making decisions.
- Working with vendors and HM Teams to deploy criteria within WAF and fine tuning it according to applications’ needs
- Multitasking and continuous ability to provide a high level of concentration for assigned projects.
- Good working knowledge of AWS security in general and familiarity of the AWS native security tools
- The candidate should be experienced and articulate, who is not going to get discouraged, despite meeting roadblocks, and will continue promoting security within the company.
- Ability to create DevSecOps security requirements while working on a project
- Ability to articulate security requirements during the Architecture meetings and working hand in hand with HM Applications and DevOps Principal Engineers







