Senior Cloud & ML Infrastructure Engineer at Hunarstreet technologies pvt ltd · Bengaluru (Bangalore), Mumbai, Hyderabad, Pune, Mohali, Panchkula, Delhi · 6 - 10 years · ₹10L - ₹15L / yr · Posted 16 Sep 2025

Senior Cloud & ML Infrastructure Engineer
at Hunarstreet technologies pvt ltd
Senior Cloud & ML Infrastructure Engineer
Location: Bangalore / Bengaluru, Hyderabad, Pune, Mumbai, Mohali, Panchkula, Delhi
Experience: 6–10+ Years
Night Shift - 9 pm to 6 am
About the Role:
We’re looking for a Senior Cloud & ML Infrastructure Engineer to lead the design,scaling, and optimization of cloud-native machine learning infrastructure. This role is ideal forsomeone passionate about solving complex platform engineering challenges across AWS, witha focus on model orchestration, deployment automation, and production-grade reliability. You’llarchitect ML systems at scale, provide guidance on infrastructure best practices, and work cross-functionally to bridge DevOps, ML, and backend teams.
Key Responsibilities:
● Architect and manage end-to-end ML infrastructure using SageMaker, AWS StepFunctions, Lambda, and ECR
● Design and implement multi-region, highly-available AWS solutions for real-timeinference and batch processing
● Create and manage IaC blueprints for reproducible infrastructure using AWS CDK
● Establish CI/CD practices for ML model packaging, validation, and drift monitoring
● Oversee infrastructure security, including IAM policies, encryption at rest/in-transit, andcompliance standards
● Monitor and optimize compute/storage cost, ensuring efficient resource usage at scale
● Collaborate on data lake and analytics integration
● Serve as a technical mentor and guide AWS adoption patterns across engineeringteams
Required Skills:
● 6+ years designing and deploying cloud infrastructure on AWS at scale
● Proven experience building and maintaining ML pipelines with services like SageMaker,ECS/EKS, or custom Docker pipelines
● Strong knowledge of networking, IAM, VPCs, and security best practices in AWS
● Deep experience with automation frameworks, IaC tools, and CI/CD strategies
● Advanced scripting proficiency in Python, Go, or Bash
● Familiarity with observability stacks (CloudWatch, Prometheus, Grafana)
Nice to Have:
● Background in robotics infrastructure, including AWS IoT Core, Greengrass, or OTA deployments
● Experience designing systems for physical robot fleet telemetry, diagnostics, and control
● Familiarity with multi-stage production environments and robotic software rollout processes
● Competence in frontend hosting for dashboard or API visualization
● Involvement with real-time streaming, MQTT, or edge inference workflows
● Hands-on experience with ROS 2 (Robot Operating System) or similar robotics frameworks, including launch file management, sensor data pipelines, and deployment to embedded Linux devices

Similar jobs (10)
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
Job Description: Lead - Cloud Engineering (AWS / Azure)
Role Title: Lead - Cloud Engineering
Experience Level: 10+ Years
Domain Focus: Healthcare AI & Cloud Infrastructure
Location: Remote
Job Overview
We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.
The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.
Key Responsibilities
Cloud Architecture & Migration
- Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
- Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
- Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.
Security & Healthcare Compliance
- Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
- Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.
Leadership & Team Management
- Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
- Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
- Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.
Operations & FinOps
- Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
- Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.
Key Requirements
- Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
- Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
- DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
- Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
- AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
- Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Title: AWS Cloud Engineer (Terraform & Monitoring Tools)
Experience: 8–10 Years
Location: Bangalore / Hyderabad
Joining: Immediate Joiners Preferred - 20 day serving
Job Description:
We are looking for an experienced AWS Cloud Engineer with 8+ years of experience in cloud infrastructure, automation, and monitoring. The ideal candidate should have strong expertise in AWS, Terraform, and at least one monitoring/observability tool such as Splunk, AppDynamics, or Grafana.
Required Skills:
- 8–10 years of IT experience with strong AWS cloud expertise
- Hands-on experience with Terraform for Infrastructure as Code (IaC)
- Experience with Splunk, AppDynamics, or Grafana (any one mandatory)
- Knowledge of cloud monitoring, logging, and performance optimization
- Experience with CI/CD pipelines and cloud infrastructure automation
- Strong troubleshooting, analytical, and communication skills
Example Responsibilities:
- Build and optimize model serving infrastructure with a focus on inference latency and cost optimization
- Architect efficient inference pipelines that balance latency, throughput, and cost across various acceleration options
- Develop monitoring and observability solutions for ML systems
- Collaborate with ML Engineers to establish best practices for optimized model deployment
- Implement cost-efficient, enterprise-scale solutions
- Collaborate in a cross-functional, distributed team for continuous system improvement
- Work with MLEs, QA Engineers, and DevOps Engineers
- Evaluate and implement new technologies and tools
- Contribute to architectural decisions for distributed ML systems
Experience and Qualifications:
- 5+ years of experience in software engineering with Python
- Experience with ML frameworks, particularly PyTorch
- Experience optimizing ML models with hardware acceleration (AWS Neuron , ONNX, TensorRT)
- Experience with AWS ML services and hardware-accelerated instances (Sagemaker, Inferentia,Trainium)
- Proven experience building and operating AWS serverless architectures
- Deep understanding of event-driven processing patterns, SQS/SNS and serverless caching solutions
- Experience with containerization using Docker and orchestration tools
- Strong knowledge of RESTful API design and implementation
- Proficiency in writing good quality & secure code and be familiar with static code analysis tools
- Excellent analytical, conceptual and communication skills in spoken and written English
- Experience applying Computer Science fundamentals in algorithm design, problem solving, and complexity analysis
Great to have Experience and Qualifications:
- Experience with any of the following: model compilation and quantization, performance profiling and benchmarking ML inference systems
- Experience working in regulated industries with strict compliance requirements for cloud-native solutions
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.
Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.
You Will:
- Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
- Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
- CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
- Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
- Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
- Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
- Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
- Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
- Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
- Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
- Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
- Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
- Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
- Perform other duties as assigned
You Have:
- Enterprise SaaS software solutions with high availability and scalability
- Solution handling large scale structured and unstructured data from varied data sources
- Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
- Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
- In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
- AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
- Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
- Programming languages like Python and SQL
- Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
- Solution Cost Optimisations and design to cost
- Legally eligible to work in India on an ongoing basis
Get to Know Us:
At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.
Equal Opportunity Employer:
Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.
If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.
Job application link : https://grnh.se/z7qx2ehx1us
Job Title: Senior AI/ML Engineer
Company: Timble Technologies Pvt. Ltd
Location: Gurugram (Hybrid)
Experience: 2 TO 5 Years
About Us
Timble Glance is a high-growth AI RegTech and B2B SaaS company catering to top-tier BFSI and enterprise clients. We build cutting-edge systems powering 30+ high-scale APIs for digital identity verification, fraud detection, document intelligence, and compliance automation.
Role Overview
We are looking for a hands-on Senior AI/ML Engineer to design, develop, and productionize high-throughput AI/ML and Generative AI systems. You will own the full lifecycle—from problem formulation and data pipelines to deep learning architectures, RAG systems, LLMOps, and model governance—delivering sub-second latency and high reliability across our enterprise products.
Key Responsibilities
· Model Architecture & Deployment: Design, train, and deploy production-scale ML/Deep Learning and GenAI systems (computer vision, document intelligence, OCR, NLP, fraud risk classification, and LLM applications).
· GenAI & LLM Solutions: Develop robust LLM workflows including prompt engineering, fine-tuning, RAG pipelines, semantic search, vector indexing (Pinecone/Milvus/Chroma), and safety guardrails.
· Pipelines & Engineering: Build performant feature extraction and data pipelines; write modular, vectorized, production-grade Python (NumPy, Pandas) and advanced SQL.
· MLOps & Monitoring: Establish end-to-end MLOps/LLMOps standards—model registries, CI/CD, experiment tracking, drift detection, A/B testing, latency optimization, and cost governance.
· Responsible AI & Security: Ensure model decisions comply with enterprise data security, privacy standards, and auditability required by the BFSI sector.
· Collaboration & Ownership: Translate complex business requirements into technical roadmaps, conduct rigorous code reviews, and mentor junior engineers.
Required Qualifications & Skills
· Education: B.Tech / M.Tech in Computer Science, AI/ML, Mathematics, or a related field—Tier-1 institutes (IIT, IIIT, NIT) strongly preferred.
· Experience: 2+ years of hands-on experience developing, deploying, and maintaining ML/Deep Learning or GenAI models in production environments.
· GenAI & NLP Stack: Hands-on experience with LLMs, embeddings, RAG architectures, and frameworks such as LangChain, LlamaIndex, or Hugging Face.
· Deep Learning Frameworks: Strong proficiency in PyTorch or TensorFlow, with deep knowledge of transformer architectures and modern NLP/CV models.
· Software & Data Engineering: Expert-level Python skills (pytest, Git, OOP, asynchronous programming), solid SQL proficiency, and familiarity with data workflows.
· Deployment & Cloud: Practical exposure to cloud platforms (AWS/GCP), containerization (Docker), API frameworks (FastAPI/Flask), and basic orchestration (Kubernetes).
Preferred Qualifications
· Prior domain experience in Fintech, RegTech, Identity Verification (KYC/AML), Fraud Intelligence, or B2B SaaS.
· Experience optimizing models for low latency and inference cost (e.g., ONNX, TensorRT, model quantization).
· Familiarity with workflow orchestrators such as Airflow, Prefect, or Kubeflow.
We're hiring a Cloud Architect (Contract) to work with our Equity Partners who builds profitable growth by acquiring and operating enterprise software companies. Refining a proprietary operating model across 40+ acquisitions and two decades of hands-on experience, now supercharged by our patented agentic AI platform . In this role, you'll take full architectural control of our CI/CD, observability, and event streaming infrastructure, build the standards every new acquisition plugs into, and use AI-assisted automation to keep 20+ products reliable without proportionally scaling headcount.
Job title: Cloud/Platform Architect (SRE)
Type: Global Remote | Contract
What You Bring
- 8–12 years in platform engineering, DevOps, or SRE, with growing ownership over time
- Deep Terraform experience across multi-account, multi-env setups
- Real production experience with event streaming at scale
- Hands-on Grafana, Prometheus, Loki, and strong AWS depth (ECS, EKS, IAM, VPC, RDS)
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortems
- Bonus: acquisition or greenfield platform-building experience
Roles and Responsibilities
- Own everything outside core AWS infra: CI/CD, observability, event streaming, deployment, incidents
- Define the standards every future acquisition will plug into
- Keep 20+ enterprise products running at serious scale (millions–billions of requests)
- Build self-service tooling so product teams never wait on you
- Use AI/automation to kill toil — not to replace engineering judgement
Ready to build the platform that scales an entire portfolio? — let's connect.
Cloud Expertise(Azure):
• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.
• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.
• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).
• Understanding on policies and security aspects of cloud.
Kubernetes & Helm:
• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.
• Understand of Kubernetes templates and its deployment.
• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.
• Implement Kubernetes best practices, including security, networking, and scaling.
• Concepts of Docker and Containers
CI/CD & Programming:
• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).
• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.
• Proven skill in python programming and concepts.
Monitoring and Observability:
Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.
• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.
Job Summary
We are looking for an experienced AWS Cloud Engineer with strong expertise in AWS infrastructure, deployment, migration, and cloud operations. The candidate should have hands-on experience managing AWS services such as EC2, EBS, S3, EFS, and FSx, along with infrastructure provisioning, migration, troubleshooting, and optimization.
Key Responsibilities
- Design, deploy, configure, and manage AWS infrastructure environments.
- Perform application and infrastructure migration to AWS.
- Provision and manage EC2 instances, including configuration, scaling, patching, and troubleshooting.
- Manage EBS volumes, snapshots, backups, and storage performance.
- Configure and administer S3 buckets, storage policies, lifecycle management, and access controls.
- Manage EFS for scalable shared file storage.
- Implement and manage Amazon FSx file systems based on application requirements.
- Monitor AWS infrastructure performance, availability, and capacity.
- Troubleshoot infrastructure, networking, storage, and deployment-related issues.
- Implement security best practices including IAM, security groups, encryption, and access controls.
- Support backup, disaster recovery, high availability, and business continuity requirements.
- Optimize AWS resources for performance, scalability, reliability, and cost.
Mandatory Skills
- Strong hands-on experience in AWS Cloud Infrastructure.
- Expertise in EC2, EBS, S3, EFS, and FSx.
- Experience in AWS deployment and migration projects.
- Strong knowledge of AWS networking concepts such as VPC, Subnets, Route Tables, Security Groups, and Load Balancers.
- Experience with IAM and AWS security best practices.
- Good knowledge of AWS monitoring and troubleshooting.
- Experience with cloud infrastructure automation using Terraform or CloudFormation is preferred.
- Strong Linux administration and troubleshooting skills.
- Good understanding of backup, disaster recovery, and high-availability concepts
WowPe is a leading fintech company revolutionizing the way businesses handle financial transactions. Our suite of innovative products includes a secure Payment Gateway for seamless online transactions, robust Payouts solutions to streamline bulk payments, and a versatile Point of Sale (POS) system for efficient in-store transactions. At WowPe, we’re dedicated to providing user-friendly, scalable, and reliable solutions that empower businesses to grow and succeed in today’s fast-paced digital economy.
We are looking for a highly skilled Cloud Infrastructure & Cloud Network Engineer to design, build, and manage secure, scalable hybrid cloud environments at WowPe. This role will focus on cloud networking, hybrid connectivity, infrastructure automation, security, reliability, and performance across on-prem and cloud platforms (AWS/Azure). You will play a critical role in ensuring high availability, security, and performance of our fintech platforms.
A Day in the Life
- Design and review cloud network architectures (VPC/VNet, routing, segmentation)
- Troubleshoot latency, connectivity, VPN, and performance issues
- Work with DevOps and application teams to support deployments and scalability
- Automate infrastructure provisioning and security guardrails
- Monitor network health, traffic flow, and system performance
- Ensure security, compliance, and disaster recovery readiness
- Support hybrid connectivity between on-prem data centers and cloud environments.
Key Responsibilities
Cloud Networking & Hybrid Architecture
- Design and implement VPC/VNet architectures, subnetting, routing tables, NAT, gateways, and secure segmentation
- Build and manage hybrid connectivity between on-prem data centers and cloud using Site-to-Site VPN, ExpressRoute, Direct Connect
- Configure and manage Layer 4 & Layer 7 load balancers for high availability and traffic distribution
- Architect DNS, CDN, and edge networking strategies for low-latency global access
Security & Zero-Trust
- Implement Zero-Trust networking using security groups, NACLs, identity-aware proxies, mTLS
- Design secure access controls and network isolation
- Work closely with security teams to ensure PCI, ISO, SOC compliance readiness
Automation, Platform & Reliability
- Automate infrastructure provisioning using Infrastructure as Code (Terraform/ARM/CloudFormation)
- Build self-healing, auto-scaling architectures with health checks and fault tolerance
- Support and manage Kubernetes clusters (EKS/AKS) including networking and service communication
- Integrate infra with CI/CD pipelines for safe and frequent deployments
Operations & Observability
- Perform advanced troubleshooting for latency, packet loss, MTU, routing loops
- Implement and manage monitoring, logging, and observability (Prometheus, Grafana, ELK, Azure Monitor, CloudWatch)
- Design and maintain disaster recovery architectures, multi-region networking, and replication strategies
- Ensure high availability, performance optimization, and cost efficiency
Basic Qualifications & Skills
- 4+ years of experience in Cloud Infrastructure & Cloud Networking
- Strong hands-on experience with AWS and/or Azure
- Deep understanding of VPC/VNet, routing, NAT, gateways, load balancers, DNS
- Experience with Hybrid connectivity (VPN, ExpressRoute, Direct Connect)
- Solid knowledge of network security, firewalls, access control, segmentation
- Hands-on experience with Infrastructure as Code (Terraform preferred)
- Experience in monitoring, troubleshooting, and incident handling
- Strong understanding of high availability, DR, and performance optimization
Preferred Qualifications
- Experience in fintech, BFSI, or high-compliance environments
- Exposure to Zero-Trust architecture and security best practices
- Experience with Kubernetes (EKS/AKS), service mesh, microservices networking
- Familiarity with CI/CD pipelines and DevOps practices
- Knowledge of CDN, edge networking, and global traffic management
- Experience with compliance frameworks (PCI-DSS, ISO 27001, SOC2)
- Ability to design large-scale, resilient, production-grade architectures







