DevOps Engineer at Bell Techlogix · Hyderabad · 5 - 10 years · ₹15L - ₹20L / yr · Profitable · Posted 22 Apr 2026

The DevOps Engineer will play a critical role in operationalizing artificial intelligence across Bell Techlogix client environments. This role focuses on building and supporting cloud infrastructure, CI/CD pipelines, and automation frameworks that power AI and machine learning workloads. The ideal candidate has experience supporting AI platforms such as Azure AI, Azure Machine Learning, Azure OpenAI, and ServiceNow or conversational AI platforms, and understands the operational requirements of production AI systems, including reliability, scalability, and security.
Key Responsibilities
•Design, build, and operate cloud infrastructure and platform services that support AI and machine learning workloads in production, SLA-driven managed services environments
•Implement CI/CD and MLOps pipelines to enable automated training, testing, deployment, and rollback of AI and ML models
•Develop and maintain Infrastructure as Code to provision AI-ready environments consistently across dev/test/prod
•Support AI platform operations including monitoring model health, pipeline execution, compute utilization, and data dependencies
•Partner with Machine Learning Engineers and Data Engineers to standardize deployment patterns for AI services and LLM-based solutions
•Enable secure and scalable AI integrations using APIs, messaging, and event-driven architectures
•Implement observability solutions for AI platforms, including logging, metrics, alerting, and drift detection integrations
•Troubleshoot AI platform incidents, perform root cause analysis, and implement remediation to improve reliability and automation coverage
•Apply security best practices for AI environments including secrets management, identity and access controls, network isolation, and policy enforcement
•Support AI-driven automation use cases across platforms such as Microsoft Copilot, ServiceNow, and conversational AI tools
•Collaborate with service desk, security, and architecture teams to continuously improve AI service delivery and operational maturity
Required Qualifications
•Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience
•5+ years of experience in DevOps, cloud engineering, or platform operations, with exposure to AI or data workloads
•Hands-on experience with Microsoft Azure, including compute, networking, storage, and monitoring services
•Experience building CI/CD pipelines using Azure DevOps, GitHub Actions, or similar tools
•Working knowledge of Infrastructure as Code (Terraform and/or Bicep/ARM)
•Scripting experience using PowerShell and/or Python
•Experience supporting production platforms with incident management, change control, and root cause analysis
•Understanding of cloud security fundamentals and enterprise governance requirements
Preferred Qualifications
•Experience with Azure Machine Learning, Azure AI Services, Azure OpenAI, or MLOps frameworks
•Exposure to containerization and orchestration technologies (Docker, Kubernetes, AKS)
•Experience supporting data pipelines or feature stores used by machine learning systems
•Familiarity with ServiceNow, AI-driven ITSM workflows, or automation platforms
•Experience with observability tools
•Knowledge of Responsible AI, data governance, and compliance considerations for AI systems
•Relevant certifications (Microsoft Azure Administrator, Azure DevOps Engineer, Azure AI Engineer)

About Bell Techlogix
About
Similar jobs (10)

Position Overview
The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems.
Responsibilities
- Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms
- Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning
- Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection
- Deploy, maintain and monitor containerized ML workloads
- Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows
- Implement AI-driven data validation, schema and concept drift detection and metadata management.
- Establish governance frameworks for AI systems, including bias detection, explainability, and auditability
- Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention
- Develop automated test suites for model performance, regression, edge cases and bias validation
- Monitor model KPIs (accuracy, precision, recall, latency, calibration)
- Ensure reproducability of experiments and production models
- Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption
Qualifications
- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
- Strong hands-on experience in Microsoft Azure Cloud.
- Good understanding of Azure services such as Compute, Storage, Event Hub, Event Subscription, Storage Queue, and PaaS services.
- Basic understanding of Azure AI Foundry and AI-related Azure service setup.
- Good Azure networking basics: VNet, subnet, routing, and basic troubleshooting.
- Strong knowledge of Terraform, especially:
- Terraform state
- plan / apply
- troubleshooting failures
- migration risks
- Terraform Enterprise concepts
- Strong Python coding capability, not just basic scripting.
- Experience using Python for API integration, automation, JSON/YAML handling, and internal tooling.
- Good understanding of CI/CD pipelines.
- Ability to troubleshoot pipeline failures.
- Comfortable with YAML and JSON.
- Ability to troubleshoot Azure infrastructure/platform issues.
- Ability to collect logs/evidence and coordinate with network/app/Microsoft support teams.
- Basic awareness of agentic AI / LLM concepts.
- Awareness of security and cost best practices.
Good to Have Skills
- Hands-on experience with Harness.
- Hands-on experience with Terraform Enterprise.
- Exposure to LangGraph / LangChain.
- Exposure to agentic AI workflows or skill creation.
- Exposure to Claude or enterprise LLM integrations.
- Knowledge of Azure ML Workspace, model registry, and managed endpoints.
- MLOps / LLMOps knowledge.
- FinOps / Azure cost optimization experience.
- Azure certifications: AZ-104, AZ-305, AZ-400, AZ-500.
Screening Priority:
Azure Cloud + Terraform + Python Coding + CI/CD Troubleshooting + YAML/JSON + Basic Agentic AI Awareness
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Role: DevOps Infrastructure
Location: Pune
Experience: 8–12 years
Job Summary:
We are looking for an experienced DevOps Infrastructure professional with strong expertise in Microsoft Azure, cloud architecture, automation, and infrastructure engineering.
Key Responsibilities:
- Design and manage scalable, secure Azure infrastructure and platform solutions.
- Implement CI/CD pipelines using Azure DevOps or GitHub Actions.
- Automate infrastructure using Terraform, Bicep, or ARM and scripting with PowerShell/Azure CLI/Python/Bash.
- Work with Azure networking, Entra ID, RBAC/PIM, storage, backup, and disaster recovery.
- Support containerized workloads using Docker and AKS.
- Define cloud standards, governance, reference architectures, and security best practices.
- Implement monitoring, observability, DevSecOps, and policy-as-code practices.
- Collaborate with architecture, security, and application teams on hybrid/multi-cloud initiatives.
Must-Have Skills:
- 3+ years hands-on experience with Microsoft Azure.
- Strong DevOps experience with CI/CD, Git, and Infrastructure as Code.
- 2+ years in infrastructure design, platform engineering, or architecture.
- Strong understanding of Azure Well-Architected Framework, networking, IAM, security, and governance.
- Experience with Azure DevOps, Terraform/Bicep, Docker, and preferably AKS.
Preferred: Azure/Azure DevOps certifications, Terraform, ITIL, CISSP, multi-cloud/hybrid cloud, GitOps, and DevSecOps experience.
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
We are looking for a hands-on Azure DevOps & Infrastructure Engineer to manage enterprise Azure environments, automate infrastructure, build CI/CD pipelines, and support cloud modernization initiatives.
🔑 Key Responsibilities
• Design, implement & support CI/CD pipelines and deployment automation
• Perform TFS / Azure DevOps Server → Azure DevOps Services migration
• Build & manage Azure infrastructure using Infrastructure as Code (IaC)
• Support Azure cloud infrastructure, networking, security, monitoring & production environments
• Drive cloud migration, application modernization & platform transformation initiatives
• Work with Terraform, containerization, DevSecOps & cloud automation
• Troubleshoot infrastructure/deployment issues and support incident & on-call operations
• Create technical documentation, runbooks and deployment standards
🎯 Mandatory Requirements
✅ AZ-204 – Azure Developer Associate – Mandatory
✅ 5–8 years total IT experience
✅ 3+ years hands-on Azure DevOps / Azure Infrastructure experience
✅ Strong experience in CI/CD & Azure DevOps
✅ Hands-on experience with TFS / Azure DevOps Server to Azure DevOps Services migration
✅ Strong understanding of Azure cloud infrastructure & cloud operations
✅ Experience with IaC, provisioning & infrastructure automation
✅ Exposure to cloud/data migration and application modernization
✅ Experience supporting Azure networking, security & monitoring
✅ Experience in production support / managed services / cloud operations
✅ Willingness to work UK/US shifts and participate in on-call support
⭐ Preferred / Optional
• AZ-400 – Microsoft Certified DevOps Engineer Expert
• AZ-104 – Azure Administrator Associate
• HashiCorp Terraform Associate
• Container platforms, orchestration & DevSecOps experience
🎓 Education: Bachelor's degree in Computer Science / IT / Engineering or related field.
Azure DevOps Engineer
Experience
5-10 years of hands-on experience in Platform Engineering, Cloud Engineering, SRE, DevOps or Infrastructure Engineering roles.
Priority 1 – Must Have (Hands-On)
Azure Cloud Platform
· Strong hands-on experience supporting Azure workloads in production environments.
· Experience designing, building and supporting Azure infrastructure using Terraform.
· Good understanding of Azure networking and connectivity patterns.
· Experience supporting:
Ø AKS
Ø Application Gateway
Ø Azure Traffic Manager
Ø Key Vault
Ø Azure Monitor / Log Analytics
Ø Managed Identities
Ø Service Principals
Ø Private Endpoints
Ø VNets, NSGs and Route Tables
Kubernetes / AKS
- Strong practical experience operating and supporting AKS.
- Ability to troubleshoot:
Ø Pod failures
Ø Ingress issues
Ø DNS issues
Ø SSL/TLS certificate issues
Ø Network routing issues
Ø Performance and availability incidents
- Experience with:
Ø Helm
Ø Ingress Controllers
Ø Cluster upgrades
Ø Scaling
Ø Monitoring
Terraform
· Strong hands-on experience writing and maintaining Terraform.
· Experience creating reusable modules.
· Experience managing:
· State files
· Remote backends
· Environment promotion
· Infrastructure lifecycle
Linux & Scripting
· Strong Linux administration fundamentals.
· Practical experience troubleshooting production issues.
· Bash scripting mandatory.
· Python desirable.
Application Support / Troubleshooting
Must be comfortable supporting business applications end-to-end.
Priority 2 – Highly Desirable
GitHub & DevOps Platform
Hands-on experience with:
· GitHub Enterprise
· GitHub Actions
· Shared workflows
· Reusable pipelines
· Repository onboarding
· Branch protections
· GitHub security features
Experience supporting:
· Runner issues
· Disk space issues
· Network connectivity issues
· Dependency failures
· Self-hosted runners lifecycle management
API Management
Pipeline failures
GitOps
Experience with:
· ArgoCD
· GitOps deployment models
· Kubernetes deployment automation
Monitoring & Observability
Experience working with:
· Prometheus
· Grafana
· Azure Monitor
· Log Analytics
· Application Insights
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.
Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.
You Will:
- Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
- Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
- CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
- Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
- Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
- Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
- Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
- Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
- Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
- Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
- Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
- Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
- Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
- Perform other duties as assigned
You Have:
- Enterprise SaaS software solutions with high availability and scalability
- Solution handling large scale structured and unstructured data from varied data sources
- Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
- Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
- In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
- AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
- Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
- Programming languages like Python and SQL
- Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
- Solution Cost Optimisations and design to cost
- Legally eligible to work in India on an ongoing basis
Get to Know Us:
At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.
Equal Opportunity Employer:
Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.
If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.
Job application link : https://grnh.se/z7qx2ehx1us
The Role
As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that
keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and
observability that let a small, fast-moving team ship confidently — and you'll keep our AI and
data workloads reliable and affordable at scale.
This is a hands-on role with real ownership: you won't be maintaining someone else's setup,
you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make
deployment boring, incidents rare, and scaling a non-event. ---
What You'll Own
**CI/CD & developer experience**
- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with
confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible
automated testing, rollbacks, and release controls.
**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or
similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and
resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch
processing for the speech pipeline.
**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting,
on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and
processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in
transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security
requirements) for a product that handles sensitive customer conversations.
**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability
and spend. ---
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production
systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and
Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI,
Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong
automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog,
OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team.
Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch
pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data
pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud
ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. ---
Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real
scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data,
and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an
Hiring for AI Engineer
Exp: 6 - 8 yrs
Edu : BE/B.Tech/MCA
Work Location : Pune
Skill Set:
- Total experience ranging from 6–8 years in software engineering/AI roles
- Min 5 years strong programming experience in Python is a MUST
- Min 3.5 years hands-on experience in AI with LLMs, RAG pipelines, and AI frameworks
- Experience with cloud platforms (AWS/Azure/GCP)
We are seeking a highly skilled Senior DevOps Engineer with 8+ years of professional experience to join our team. In this role, you will design, implement, and optimize cloud infrastructure and CI/CD processes.
You will collaborate closely with development, QA, and operations teams to deliver scalable, secure, automated, and reliable solutions on AWS.
The ideal candidate will have strong hands-on experience with AWS, Terraform, Git/GitHub, Jenkins, PowerShell, Python, AWS Systems Manager (SSM) Documents, and Active Directory (AD) administration, along with a passion for automation, efficiency, and operational excellence.
Key Responsibilities
- Design, build, and maintain scalable cloud infrastructure on AWS.
- Develop and manage Infrastructure as Code (IaC) using Terraform.
- Build, maintain, and optimize CI/CD pipelines using Jenkins and GitHub.
- Use AWS Systems Manager (SSM) to support operational automation and system administration.
- Automate system tasks and administrative workflows using PowerShell and other scripting languages.
- Create, maintain, and execute custom SSM Documents for configuration management, patching, automation, and troubleshooting.
- Manage and administer Active Directory (AD), including users, groups, permissions, policies, authentication, and integration with AWS services.
- Implement and manage version-control workflows in Git and GitHub.
- Ensure infrastructure and deployment processes follow best practices for security, reliability, scalability, and cost optimization.
- Monitor and troubleshoot production systems to ensure high availability, performance, and reliability.
- Collaborate with development teams to improve software delivery processes and release management.
- Mentor junior engineers and contribute to the development of DevOps standards and best practices.
Required Skills and Experience
- 8+ years of professional experience, including at least 4 years in DevOps or Site Reliability Engineering (SRE) roles.
- Strong hands-on experience with AWS services, including EC2, VPC, IAM, S3, EKS, Lambda, and related services.
- Proven expertise in Terraform for Infrastructure as Code.
- Experience administering both Linux and Windows operating systems.
- Experience designing and managing CI/CD pipelines using Jenkins and GitHub.
- Proficiency with Git workflows and source-code management best practices.
- Strong PowerShell scripting skills; familiarity with Python and/or Bash is a plus.
- Experience with AWS Systems Manager (SSM), including creating and managing SSM Documents.
- Hands-on experience managing Active Directory, including users, groups, policies, authentication, permissions, and AWS integration.
- Solid understanding of cloud networking, security, monitoring, and troubleshooting.
- Excellent problem-solving, communication, collaboration, and decision-making skills.
Nice-to-Have Skills
- Experience with containerization and orchestration technologies, such as Docker, Kubernetes, and Amazon EKS.
- Knowledge of monitoring and observability tools, such as Amazon CloudWatch, New Relic, and Sumo Logic.
- An AWS certification, such as AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect.
Who You Are
- You are eager to learn new technologies and continuously improve your skills.
- You make sound decisions and take ownership of your work.
- You are proactive, self-motivated, and comfortable taking initiative.
- You are a strong communicator who enjoys collaborating with cross-functional teams.
- You are committed to improving processes, automation, reliability, and operational efficiency.





