Cutshort logo
For Employers
NeoGenCode Technologies Pvt Ltd logo
AI Infrastructure Architecture
AI Infrastructure Architecture

AI Infrastructure Architecture at NeoGenCode Technologies Pvt Ltd · Remote only · 8 - 14 years · ₹25L - ₹35L / yr · Raised funding · Remote only · Posted 25 Jun 2025

NeoGenCode Technologies Pvt Ltd's logo

AI Infrastructure Architecture

Divya Sharma's profile picture
Posted by Divya Sharma
8 - 14 yrs
₹25L - ₹35L / yr
Remote only
Skills
IT infrastructure
Cloud Computing

Education:

•Bachelor’s degree in Computer Science, IT Infrastructure Engineering, or a related field

•Certifications in cloud computing (AWS, Azure, GCP), MLOps, or DevOps (preferred)

Experience:

•8+ years of experience in IT Infrastructure, Cloud Computing, or AI Systems Architecture

•Hands on experience in AI infrastructure design, cloud-based AI solutions, or MLOps

•Expertise in managing AI workloads across cloud and hybrid environments.

•Proven track record in scaling AI infrastructure for large enterprises

•Strong experience in Kubernetes, containerization, and orchestration tools

•Experience in optimizing AI workloads for performance and cost efficiency

Skills :

•Expertise in AI Infrastructure & Cloud Architecture

•Strong Understanding of AI Model Deployment & MLOps

•Advanced Proficiency in Kubernetes & AI Workload Orchestration

•Hands-on Experience with Cloud Platforms (AWS, Azure, GCP)

•Proficiency in Infrastructure as Code (Terraform, Ansible)

•AI Security & Compliance Knowledge

•AI Infrastructure Cost Optimization Strategies

•Performance Tuning for AI Systems & Workloads

 

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About NeoGenCode Technologies Pvt Ltd

Founded :
2023
Type :
Services
Size :
20-100
Stage :
Raised funding

About

Welcome to Neogencode Technologies, an IT services and consulting firm that provides innovative solutions to help businesses achieve their goals. Our team of experienced professionals is committed to providing tailored services to meet the specific needs of each client. Our comprehensive range of services includes software development, web design and development, mobile app development, cloud computing, cybersecurity, digital marketing, and skilled resource acquisition. We specialize in helping our clients find the right skilled resources to meet their unique business needs. At Neogencode Technologies, we prioritize communication and collaboration with our clients, striving to understand their unique challenges and provide customized solutions that exceed their expectations. We value long-term partnerships with our clients and are committed to delivering exceptional service at every stage of the engagement. Whether you are a small business looking to improve your processes or a large enterprise seeking to stay ahead of the competition, Neogencode Technologies has the expertise and experience to help you succeed. Contact us today to learn more about how we can support your business growth and provide skilled resources to meet your business needs.

Read more

Candid answers by the company

What does the company do?
What is the location preference of jobs?

IT & Engineering Talent Staffing

  • Provides full-time and contract-based hiring, delivering handpicked, pre‑screened developers across tech stacks—ranging from web, mobile, AI/ML, Web3/blockchain.
  • Maintains a bench o vetted candidates, offering fast delivery of interview-ready profiles—often within 24 hours.
  • Offers payroll management, handling compliance, tax, attendance, and documentation for both contractors and full-time employees.

2. End-to-End Project Delivery

  • Delivers full-stack development solutions: web, mobile, cloud, AI/ML, Blockchain/Web3.
  • Manages entire project lifecycle—requirements gathering, design (UI/UX), development, deployment, and ongoing support .

3. Additional Offerings

  • Expands into cybersecurity consulting, digital marketing, and cloud platform services (like AWS, GCP, Azure) .
  • Provides strategic IT consulting to align technology solutions with business objectives

Company social profiles

bloginstagramlinkedinfacebook

Similar jobs (10)

Amura Health
at Amura Health
3 candid answers
1 video
Swathi S
Posted by Swathi S
Chennai
7 - 12 yrs
₹30L - ₹55L / yr
skill iconAmazon Web Services (AWS)
skill iconPython
CI/CD
DevOps
Platform as a Service (PaaS)
+7 more

Amura’s Vision 


We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.


Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.


Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.


These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.

We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence. 


Role Overview 


We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.


This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.


You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability. 


Key Responsibilities 


Cloud Infrastructure & Platform Engineering (AWS) 

  • Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
  • Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
  • Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
  • Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
  • Build reusable platform templates and shared infrastructure modules. 


AI/ML Infrastructure & MLOps 

  • Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
  • Support GPU-based workloads and optimize compute/storage usage.
  • Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
  • Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines. 


CI/CD, Automation & Developer Productivity 

  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
  • Automate deployments, environment provisioning, and release workflows.
  • Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
  • Implement automated patching, scaling, backups, cleanup workflows, and drift detection. 


Containers, Kubernetes & Platform Reliability

  • Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
  • Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
  • Optimize infrastructure for performance, resilience, and cost-efficiency.
  • Implement progressive deployment strategies including blue/green, canary, and rolling deployments. 


Observability, Incident Response & SRE Practices

  • Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
  • Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
  • Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.

FinOps, Cost Governance & Security

  • Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
  • Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
  • Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.

Collaboration, Leadership & Platform Culture

  • Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
  • Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.

Skills & Qualifications


Must-Have:

  • 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
  • Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Here are answers to some questions you may have

Where is your office?

Chennai (Velachery)

Work Model

Work from Office – because great stories are built in person!

Do you have an online presence?

https://amura.ai (we are @AmuraHealth on all social media)


Read more
Mango Sciences
Remote only
10 - 15 yrs
₹30L - ₹35L / yr
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Generative AI
Implementation
System deployment
+9 more

Job Description: Lead - Cloud Engineering (AWS / Azure)

Role Title: Lead - Cloud Engineering

Experience Level: 10+ Years

Domain Focus: Healthcare AI & Cloud Infrastructure

Location: Remote

Job Overview

We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.

The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.

Key Responsibilities

Cloud Architecture & Migration

  • Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
  • Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
  • Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.

Security & Healthcare Compliance

  • Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
  • Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.

Leadership & Team Management

  • Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
  • Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
  • Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.

Operations & FinOps

  • Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
  • Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.

Key Requirements

  • Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
  • Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
  • DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
  • Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
  • AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
  • Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).


Read more
Remote only
10 - 15 yrs
₹28L - ₹34L / yr
GenAI,
DevOps
skill iconPython
Infrastructure architecture
Infrastructure-as-Code
+4 more


Job Description: AI Engineer – GenAI Platform Automation

Experience: 10+ Years

Location: Remote – Pan India

Employment Type: Haparz Payroll

Work Mode: Remote

Notice Period: Immediate / Short Notice Preferred


About the Role


We are looking for a senior AI Engineer – GenAI Platform Automation to lead automation initiatives across enterprise Generative AI, Data Science, Data Engineering, and Analytics platforms.

The role focuses on building scalable, secure, and self-service automation capabilities across infrastructure provisioning, CI/CD, cloud environments, AI workload deployment, governance, observability, and operational excellence. The ideal candidate will have strong hands-on experience in platform engineering, cloud automation, DevOps, Infrastructure-as-Code, Python, and enterprise GenAI ecosystems.


Key Responsibilities

  • Lead end-to-end automation initiatives for enterprise GenAI, Data Science, Data Engineering, Metadata, Data Quality, Event Streaming, and Analytics platforms.
  • Design self-service automation for infrastructure provisioning, environment onboarding, deployment, governance, monitoring, and operational workflows.
  • Build automation capabilities supporting the AI lifecycle, including experimentation, model training, deployment, inference, observability, and lifecycle management.
  • Develop scalable Infrastructure-as-Code solutions using Terraform and cloud-native automation frameworks.
  • Design and maintain enterprise CI/CD pipelines, automated testing, deployment, and release processes using modern DevOps toolchains.
  • Automate Kubernetes, containers, serverless, and distributed computing environments in collaboration with cloud and platform engineering teams.
  • Develop automation solutions for GenAI and Agentic AI applications, including MCP-enabled services, API integrations, workflow automation, and event-driven architectures.
  • Implement observability, monitoring, logging, tracing, alerting, automated remediation, and reliability engineering practices.
  • Work with architecture, security, governance, engineering, and business teams to ensure enterprise standards and compliance requirements are met.
  • Conduct technical design reviews, automation assessments, code reviews, and establish engineering best practices.
  • Provide technical leadership and mentorship to engineering teams adopting automation-first and platform engineering practices.

What We’re Looking For

  • 10+ years of hands-on experience in platform engineering, automation engineering, cloud engineering, DevOps, or distributed systems.
  • Strong experience building enterprise self-service platforms supporting AI/ML, Data Science, Data Engineering, or Advanced Analytics workloads.
  • Strong expertise in automation frameworks, CI/CD, DevOps, Infrastructure-as-Code, and software delivery lifecycle automation.
  • Hands-on experience with Terraform and cloud-native infrastructure automation.
  • Strong experience with Python for automation, orchestration, scripting, tooling, and platform engineering.
  • Experience with Bitbucket, Bamboo, Jira, Confluence, or similar enterprise DevOps toolchains.
  • Experience working with Kubernetes, containers, serverless platforms, YARN, and distributed processing environments.
  • Knowledge of Generative AI and Agentic AI architectures, MCP frameworks, APIs, workflow automation, and enterprise AI platforms.
  • Experience with event-driven architectures and technologies such as Kafka and streaming platforms.
  • Strong understanding of cloud engineering, networking, security, scalability, resilience, and cost optimization.
  • Experience implementing observability solutions covering monitoring, logging, tracing, alerting, and operational dashboards.
  • Understanding of metadata management, data lineage, data governance, and semantic-layer concepts is highly valuable.

Good to Have

  • Experience supporting enterprise GenAI platforms, AI governance, model management, and AI operationalization.
  • Experience with GitOps, DevSecOps, Platform Engineering, and Reliability Engineering practices.
  • Exposure to data governance, data quality, metadata management, and model lifecycle automation.
  • Experience creating reusable internal developer platforms and self-service engineering tools at enterprise scale.
  • Banking, AML, fraud detection, financial crime, or risk analytics domain experience is an advantage.


Read more
o9 Solutions, Inc.
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Bengaluru (Bangalore)
8 - 16 yrs
₹1L - ₹2L / yr (ESOP available)
Large Language Models (LLM)
Agentic AI
Applied mathematics

Key Responsibilities:

·      Architectural Leadership: Design and lead the development of robust, scalable AI architectures, ensuring high performance, reliability, and security.

·      Applied Mathematics & Statistics: Apply statistical analysis, numerical computation, and mathematical modeling to derive insights from large-scale data and optimize model performance.

·      Deep Learning Development: Design, train, and deploy advanced Deep Learning (DL) models.

·      Technical Mentorship: Mentor engineering teams on best practices for AI/ML, coding standards, and architectural design.

·      Model Optimization: Optimize models for speed, efficiency, and accuracy using techniques like pruning, quantization, or GPU acceleration.

·      Strategy & Innovation: Evaluate and select appropriate AI frameworks, tools, and platforms, staying abreast of cutting-edge research and industry trends.

Qualifications:

Required:

·      Education: Master's or PhD in Computer Science, Applied Mathematics, Statistics, Physics, or a related quantitative field.

·      Experience: 10+ years of experience in software development, with at least 3-5 years in a Applied Mathematics and Deep learning.

·      AI/ML Expertise: Proven experience designing and deploying deep learning models in production using frameworks.

·      Mathematics/Statistics: Strong proficiency in linear algebra, calculus, probability, and statistical methods.

·      Programming Skills: Expert-level coding skills in Python (NumPy, Pandas, Scikit-learn) and experience with languages like Java or C++.

Key Competencies:

  • Strategic mindset with deep operational awareness.
  • Excellent communication and stakeholder management skills.
  • Ability to simplify complex technical concepts for executive reporting.
  • Strong leadership, people development, and cross-functional influencing skills.

Bias for action and a relentless focus on continuous improvement.

Read more
The industry’s only Manufacturing Operating System
The industry’s only Manufacturing Operating System
Agency job
via Cutshort Lightning by Ariba Khan
Hyderabad
10 - 15 yrs
Best in industry
Artificial Intelligence (AI)
Generative AI (GenAI)
Retrieval Augmented Generation (RAG)
Large Language Models (LLM)

We’re on hunt for AI Architect


Responsibilities:

  • 10–15+ years overall experience, with recent hands-on AI/GenAI architecture ownership.
  • Must have architected enterprise AI platforms/solutions end-to-end, not just individual ML models or PoCs.
  • Strong GenAI/LLM production experience: RAG, embeddings, vector DBs, hybrid search, reranking, evaluation, guardrails.
  • Strong Agentic AI understanding: agents, tool calling, workflows, orchestration, human-in-the-loop.
  • Experience taking AI solutions from architecture → production → scale, ideally across multiple business teams/use cases.
  • Strong cloud architecture — Azure/AWS preferred; hybrid/on-prem experience is a plus.
  • Must understand enterprise security, governance, Responsible AI, observability and LLMOps/MLOps.
  • Should be able to articulate build-vs-buy, MVP-vs-target architecture, cost/performance/security tradeoffs.
  • Strong stakeholder-facing / consulting ability — can work with business leaders, engineering, security and data teams and influence without authority.


There is scope to move to the US for this role if you are aligned for the same, else this will be a WFO role from Hyderabad location

Read more
Ampera Technologies
Chennai, Bengaluru (Bangalore), Mumbai, Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Pune, Hyderabad, Kolkata
6 - 15 yrs
Best in industry
skill iconMachine Learning (ML)
MLOps
skill iconKubernetes
Large Language Models (LLM)
openshift
+1 more

Title                                 : Senior AI Platform / MLOps Engineer

Experience                    : 6+ years

Work type                      : Chennai - Work from Office/other locations - Remote

Employment Type      : Full Time

Notice Period              : Immediate

Work Day                     :Mon to Fri

 

Key Responsibilities:

  • Install, configure and operate OpenShift, NVIDIA GPU operator, OpenShift AI, and NIM microservices on 12× RTX PRO 6000 across two servers; single-node and HA control-plane topologies
  • Serving configuration and tuning: quantized model deployment (FP8/FP4), replica balancing, batching, KV-cache and context management
  • Azure GPU build environments: provisioning, cost control, parity with the on-prem stack via pinned container/model versions; cloud-to-factory migration with parity regression
  • GitOps CI/CD, container registry, artifact/model versioning, environment promotion; observability and audit wiring (Splunk, Prometheus/Grafana)
  • Benchmark automation: load harness, p50/p95/p99 latency, tokens/sec, GPU utilization; the capacity report data pipeline
  • Platform upgrade procedure with evaluation-regression gates; deployment runbook as a first-class deliverable

Technical Skills:

  • 6+ years infrastructure/platform engineering with 3+ years production Kubernetes; OpenShift experience strongly preferred
  • Hands-on GPU inference serving in production: NIM, Triton, vLLM, or TensorRT-LLM — you have sized, deployed, and tuned LLM serving on real GPUs and can talk memory-bandwidth trade-offs
  • GitOps fluency (ArgoCD/Flux), infrastructure-as-code, container internals; comfortable in air-gapped/proxy-restricted enterprise networks
  • Observability depth: metrics, traces, log pipelines; has built performance test harnesses, not just run them


  • Azure or AWS GPU compute operations experience

Strongly preferred

  • NVIDIA GPU operator and AI Enterprise stack specifics; KServe; Milvus or pgvector operations; VAST/NFS/S3 storage integration; banking or other regulated-environment delivery

 



About Ampera: 

Ampera Technologies, a purpose driven Digital IT Services with primary focus on supporting our client with their Data, AI / ML, Accessibility and other Digital IT needs. We also ensure that equal opportunities are provided to Persons with Disabilities Talent. Ampera Technologies has its Global Headquarters in Chicago, USA and its Global Delivery Center is based out of Chennai, India. We are actively expanding our Tech Delivery team in Chennai and across India. We offer exciting benefits for our teams, such as 1) Hybrid and Remote work options available, 2) Opportunity to work directly with our Global Enterprise Clients, 3) Opportunity to learn and implement evolving Technologies, 4) Comprehensive healthcare, and 5) Conducive environment for Persons with Disability Talent meeting Physical and Digital Accessibility standards 

Read more
KnackLabs
at KnackLabs
5 recruiters
Stuti Jain
Posted by Stuti Jain
Hyderabad
7 - 10 yrs
₹25L - ₹35L / yr
Retrieval Augmented Generation (RAG)
skill iconAmazon Web Services (AWS)

Location: Hyderabad, India. Based at the KnackLabs headquarters, with occasional travel to client locations for workshops and reviews. This role does not involve extended onsite deployments.

About the Role

You will work as an AI Architect who designs the systems behind our client engagements: AI agents, RAG systems, automation platforms, and the conventional backend systems around them.

This is a hands-on design role, not a slideware role. You will scope architectures with clients, make the hard technical decisions, defend them in review, and stay accountable for how the systems perform in production.


You will work directly with clients. Everyone at KnackLabs does. You will sit in design discussions with client engineering teams, present architecture decisions to technical and business stakeholders, and answer for the choices you make.


A full KnackLabs engineering team in Hyderabad builds with you. You own the technical design and the quality of what ships.

What you'll own

  1. Architecture - Design AI agents, RAG systems, integrations, and the scalable backend systems around them, for multiple client engagements.
  2. Technical scoping - Work directly with clients to turn a business problem into a system design, with clear trade-offs and clear reasons.
  3. Scale and reliability - Make sure what we build handles real load: data stores, queues, caching, horizontal scaling, and fault tolerance.
  4. Design reviews - Review designs and builds across engagements. Set the technical bar and hold it.
  5. Evaluation strategy - Define how we measure accuracy, safety, latency, and cost for the AI systems we ship.
  6. Guiding engineers - Raise the level of the engineers building with you, through reviews and direct pairing.
  7. Feedback to the platform - Feed what you learn across engagements back into our platform and internal tools.

What we are looking for

  1. Around 7 or more years of software engineering experience, including direct work with customers on design or delivery.
  2. Full-stack development experience with strength in backend technologies.
  3. Experience designing and building scalable applications. You understand how large-scale distributed systems work: data partitioning, queues, caching, horizontal scaling, and fault tolerance.
  4. At least 2 years of strong, hands-on AI experience with large language models in production.
  5. You build with AI coding tools like Claude Code or Codex as your default way of working. You understand Claude Skills, have written skills yourself, use them actively, and have contributed to them.
  6. Hands-on experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
  7. Hands-on experience building AI agents.
  8. Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
  9. Experience with at least one cloud platform (AWS, Azure, or GCP).
  10. Clear communication. You can explain an architecture decision to an engineer and to a business leader, and defend it under questioning.
  11. High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a design.

Nice to have

  1. Experience building evaluations to measure accuracy, safety, latency, and cost.
  2. Experience with observability and tracing tools such as LangSmith or Braintrust.
  3. Experience with on-premises or private cloud (VPC) deployments.
  4. Experience deploying AI systems in regulated industries such as insurance, banking, or the public sector.
  5. Experience with data engineering and pipelines.
  6. A history of side projects, open source contributions, or products you shipped end-to-end.
  7. Experience working at a consulting or professional services firm in a client-facing delivery role.

Stack and tools

  1. Languages: Python and TypeScript.
  2. Models: Claude and other frontier or open-source models, chosen to fit the customer.
  3. AI patterns: RAG, agents, prompt engineering, skills, and evaluations.
  4. Vector and retrieval: vector databases and retrieval pipelines.
  5. Cloud: AWS, Azure, or GCP, on public or private cloud.
  6. Integration: REST APIs and enterprise system connectors.


Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Akilandeswari Panneerselvam
Mumbai
8 - 10 yrs
₹7L - ₹15L / yr
skill iconJava
skill iconAmazon Web Services (AWS)
skill iconKubernetes
DevOps
Artificial Intelligence (AI)
+3 more

Job Title: Java AWS Kubernetes DevOps AI/ML

Experience: 8–10 Years

The candidate should have at least 1 year of experience in AI/ML and hands-on experience with the below technologies:

Java

AWS

Kubernetes

DevOps

AI/ML

MongoDB / PostgreSQL

Key-Value Caching

Vector Databases – ChromaDB / Pgvector

Read more
Wissen Technology
at Wissen Technology
4 recruiters
Shakthi M
Posted by Shakthi M
Bengaluru (Bangalore)
5 - 14 yrs
Best in industry
skill iconPython
Azure
Terraform
DevOps
  • Strong hands-on experience in Microsoft Azure Cloud.
  • Good understanding of Azure services such as Compute, Storage, Event Hub, Event Subscription, Storage Queue, and PaaS services.
  • Basic understanding of Azure AI Foundry and AI-related Azure service setup.
  • Good Azure networking basics: VNet, subnet, routing, and basic troubleshooting.
  • Strong knowledge of Terraform, especially:
  • Terraform state
  • plan / apply
  • troubleshooting failures
  • migration risks
  • Terraform Enterprise concepts
  • Strong Python coding capability, not just basic scripting.
  • Experience using Python for API integration, automation, JSON/YAML handling, and internal tooling.
  • Good understanding of CI/CD pipelines.
  • Ability to troubleshoot pipeline failures.
  • Comfortable with YAML and JSON.
  • Ability to troubleshoot Azure infrastructure/platform issues.
  • Ability to collect logs/evidence and coordinate with network/app/Microsoft support teams.
  • Basic awareness of agentic AI / LLM concepts.
  • Awareness of security and cost best practices.

Good to Have Skills

  • Hands-on experience with Harness.
  • Hands-on experience with Terraform Enterprise.
  • Exposure to LangGraph / LangChain.
  • Exposure to agentic AI workflows or skill creation.
  • Exposure to Claude or enterprise LLM integrations.
  • Knowledge of Azure ML Workspace, model registry, and managed endpoints.
  • MLOps / LLMOps knowledge.
  • FinOps / Azure cost optimization experience.
  • Azure certifications: AZ-104, AZ-305, AZ-400, AZ-500.

 

Screening Priority:

Azure Cloud + Terraform + Python Coding + CI/CD Troubleshooting + YAML/JSON + Basic Agentic AI Awareness

 

Read more
Smartsheet
Sandeep Selvan
Posted by Sandeep Selvan
Bengaluru (Bangalore)
4 - 12 yrs
Best in industry
MLOps
databricks
skill iconMachine Learning (ML)
MLFlow
LangGraph
+4 more

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.


Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.


You Will:

  • Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
  • Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
  • CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
  • Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
  • Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
  • Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
  • Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
  • Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
  • Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
  • Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
  • Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
  • Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
  • Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
  • Perform other duties as assigned


You Have:

  • Enterprise SaaS software solutions with high availability and scalability
  • Solution handling large scale structured and unstructured data from varied data sources
  • Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
  • Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
  • In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
  • AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
  • Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
  • Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
  • Programming languages like Python and SQL
  • Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
  • Solution Cost Optimisations and design to cost
  • Legally eligible to work in India on an ongoing basis

 

Get to Know Us:

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.


Equal Opportunity Employer:

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information. 

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.



Job application link : https://grnh.se/z7qx2ehx1us

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos