AI/ML Engineer - AI Observability(ML Ops Engineer) at Hiring for IT Consulting Firm (MNC) · Pune, Nagpur · 5 - 10 years · ₹20L - ₹30L / yr · Posted 30 Sep 2026

AI/ML Engineer - AI Observability(ML Ops Engineer)
at Hiring for IT Consulting Firm (MNC)
Position Overview
The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems.
Responsibilities
- Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms
- Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning
- Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection
- Deploy, maintain and monitor containerized ML workloads
- Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows
- Implement AI-driven data validation, schema and concept drift detection and metadata management.
- Establish governance frameworks for AI systems, including bias detection, explainability, and auditability
- Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention
- Develop automated test suites for model performance, regression, edge cases and bias validation
- Monitor model KPIs (accuracy, precision, recall, latency, calibration)
- Ensure reproducability of experiments and production models
- Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption
Qualifications
- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn

Similar jobs (10)
- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 8-10 years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.
Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.
You Will:
- Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
- Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
- CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
- Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
- Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
- Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
- Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
- Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
- Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
- Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
- Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
- Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
- Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
- Perform other duties as assigned
You Have:
- Enterprise SaaS software solutions with high availability and scalability
- Solution handling large scale structured and unstructured data from varied data sources
- Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
- Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
- In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
- AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
- Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
- Programming languages like Python and SQL
- Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
- Solution Cost Optimisations and design to cost
- Legally eligible to work in India on an ongoing basis
Get to Know Us:
At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.
Equal Opportunity Employer:
Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.
If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.
Job application link : https://grnh.se/z7qx2ehx1us
Greetings!
Hiring For Large Product Based Company!
Role- Mlops Engineer
Experience- 8-12 years
Location- Pune, Nagpur
JD-
- 8-10 years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
Title : Senior AI Platform / MLOps Engineer
Experience : 6+ years
Work type : Chennai - Work from Office/other locations - Remote
Employment Type : Full Time
Notice Period : Immediate
Work Day :Mon to Fri
Key Responsibilities:
- Install, configure and operate OpenShift, NVIDIA GPU operator, OpenShift AI, and NIM microservices on 12× RTX PRO 6000 across two servers; single-node and HA control-plane topologies
- Serving configuration and tuning: quantized model deployment (FP8/FP4), replica balancing, batching, KV-cache and context management
- Azure GPU build environments: provisioning, cost control, parity with the on-prem stack via pinned container/model versions; cloud-to-factory migration with parity regression
- GitOps CI/CD, container registry, artifact/model versioning, environment promotion; observability and audit wiring (Splunk, Prometheus/Grafana)
- Benchmark automation: load harness, p50/p95/p99 latency, tokens/sec, GPU utilization; the capacity report data pipeline
- Platform upgrade procedure with evaluation-regression gates; deployment runbook as a first-class deliverable
Technical Skills:
- 6+ years infrastructure/platform engineering with 3+ years production Kubernetes; OpenShift experience strongly preferred
- Hands-on GPU inference serving in production: NIM, Triton, vLLM, or TensorRT-LLM — you have sized, deployed, and tuned LLM serving on real GPUs and can talk memory-bandwidth trade-offs
- GitOps fluency (ArgoCD/Flux), infrastructure-as-code, container internals; comfortable in air-gapped/proxy-restricted enterprise networks
- Observability depth: metrics, traces, log pipelines; has built performance test harnesses, not just run them
- Azure or AWS GPU compute operations experience
Strongly preferred
- NVIDIA GPU operator and AI Enterprise stack specifics; KServe; Milvus or pgvector operations; VAST/NFS/S3 storage integration; banking or other regulated-environment delivery
About Ampera:
Ampera Technologies, a purpose driven Digital IT Services with primary focus on supporting our client with their Data, AI / ML, Accessibility and other Digital IT needs. We also ensure that equal opportunities are provided to Persons with Disabilities Talent. Ampera Technologies has its Global Headquarters in Chicago, USA and its Global Delivery Center is based out of Chennai, India. We are actively expanding our Tech Delivery team in Chennai and across India. We offer exciting benefits for our teams, such as 1) Hybrid and Remote work options available, 2) Opportunity to work directly with our Global Enterprise Clients, 3) Opportunity to learn and implement evolving Technologies, 4) Comprehensive healthcare, and 5) Conducive environment for Persons with Disability Talent meeting Physical and Digital Accessibility standards
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
About the role
We are building AI systems that read, understand and act on real business documents, bank statements, financial reports, policy documents and forms and putting them into production where accuracy and cost both matters.
This is not a research role and it is not a prompt-writing role. You will own features end to end: pick and deploy open-source models, build the pipelines around them, measure whether they actually work on our documents, drive the cost per document down, and keep the whole thing running in production.
You will work closely with the engineering and product teams, and your work will be directly used by business users from day one.
What you will do
Deploy and evaluate open-source models
- Select, deploy and benchmark open-source LLMs and vision-language models for specific, narrow use cases not general chat.
- Build evaluation sets from real documents and define what "good" means numerically (field-level accuracy, extraction recall, hallucination rate) before shipping.
- Run structured comparisons between models and approaches, and write up the trade-offs so the team can make a decision.
- Apply quantization, batching and other optimizations to fit models into a sensible GPU budget.
Build and optimize AI orchestration
- Design multi-step pipelines that combine deterministic code, ML models and LLM calls and know when not to use an LLM.
- Optimize for latency, cost and reliability: caching, batching, request routing, fallback tiers, retries and graceful degradation.
- Instrument pipelines so failures are visible and traceable rather than silent.
Ship to production
- Package models and services with Docker, expose them behind clean APIs, and deploy them to our GPU and CPU infrastructure.
- Handle the unglamorous production concerns: cold starts, timeouts, concurrency limits, versioning, rollback and monitoring.
- Own on-call-style responsibility for the AI features you build, including cost tracking.
Must-have skills
Programming & engineering
- Strong Python: type hints, async/await, dataclasses/Pydantic, clean module design, testing.
- REST API development with FastAPI (or Flask/Django with a willingness to move to FastAPI).
- Git, code review discipline, and the ability to write code someone else can maintain.
- Comfortable in Linux and on the command line.
Machine learning fundamentals
- Working knowledge of PyTorch and the Hugging Face ecosystem (transformers, tokenizers, accelerate).
- Understanding of inference-time concepts: tokenization, context windows, batching, precision (FP16/BF16/INT8), memory footprint.
- Ability to read a model card and a paper well enough to judge whether a model fits a use case.
Document processing
- Hands-on experience with at least two of: pypdfium2, PyMuPDF, pdfplumber, pdfminer.six, Docling, Unstructured, Surya, DocTR, LayoutLM family.
- Practical OCR experience (Tesseract, PaddleOCR, or a cloud OCR) and an understanding of when OCR is the wrong tool.
- Experience extracting tables from PDFs and dealing with merged cells, multi-line rows, and inconsistent column layouts.
Strongly preferred
You will be a much stronger candidate with any of these. We do not expect all of them.
Model serving & optimization
- vLLM, TGI, Ollama, llama.cpp, or Triton Inference Server.
- Quantization formats and tooling: GGUF, AWQ, GPTQ, bitsandbytes, ONNX Runtime, INT8 export.
- Serverless GPU platforms: Modal, RunPod, Replicate, Baseten including cold-start and container-lifecycle management.
- LoRA / QLoRA fine-tuning with PEFT for narrow, task-specific improvements.
Vision-language models
- Practical use of open VLMs: Qwen2.5-VL, InternVL, Granite Vision, Molmo, Phi-Vision, or similar.
- Awareness of where VLMs hallucinate especially on numeric and financial content and patterns for constraining them (using the model for layout only, sourcing values from the text layer, constrained decoding).
Orchestration & pipelines
- Workflow orchestration: Dagster, Airflow, Prefect, or Temporal.
- Async job patterns: Celery, RQ, or platform-native spawn/poll patterns.
- LLM orchestration frameworks (LangGraph, LlamaIndex, Haystack) with the judgement to know when plain Python is a better answer.
- Structured output enforcement: Instructor, Outlines, XGrammar, JSON schema / tool-use modes.
Evaluation & observability
- Building golden datasets and regression suites for extraction tasks.
- Eval tooling: promptfoo, DeepEval, Ragas, or in-house harnesses.
- LLM tracing and monitoring: Langfuse, Arize Phoenix, LangSmith, OpenTelemetry.
Nice extras
- Rule engines and policy evaluation (Open Policy Agent / Rego, Drools, rule-engine).
- Experience in fintech, lending, insurance or accounting documents.
- Handling of PII and data-security practices in document pipelines.
- Contributions to open-source ML or document-processing projects.
Why join us
- Real production ownership from month one your work goes to actual users, not a demo.
- Genuinely hard technical problems in document AI, not wrappers over an API.
- Small team, short decision cycles, direct access to leadership.
- Budget and freedom to evaluate and adopt new open-source models as they land.
To apply: send your CV along with a short note on one AI system you have taken to production what it did, what the accuracy was, and what broke.
Job Title: Senior AI/ML Engineer
Company: Timble Technologies Pvt. Ltd
Location: Gurugram (Hybrid)
Experience: 2 TO 5 Years
About Us
Timble Glance is a high-growth AI RegTech and B2B SaaS company catering to top-tier BFSI and enterprise clients. We build cutting-edge systems powering 30+ high-scale APIs for digital identity verification, fraud detection, document intelligence, and compliance automation.
Role Overview
We are looking for a hands-on Senior AI/ML Engineer to design, develop, and productionize high-throughput AI/ML and Generative AI systems. You will own the full lifecycle—from problem formulation and data pipelines to deep learning architectures, RAG systems, LLMOps, and model governance—delivering sub-second latency and high reliability across our enterprise products.
Key Responsibilities
· Model Architecture & Deployment: Design, train, and deploy production-scale ML/Deep Learning and GenAI systems (computer vision, document intelligence, OCR, NLP, fraud risk classification, and LLM applications).
· GenAI & LLM Solutions: Develop robust LLM workflows including prompt engineering, fine-tuning, RAG pipelines, semantic search, vector indexing (Pinecone/Milvus/Chroma), and safety guardrails.
· Pipelines & Engineering: Build performant feature extraction and data pipelines; write modular, vectorized, production-grade Python (NumPy, Pandas) and advanced SQL.
· MLOps & Monitoring: Establish end-to-end MLOps/LLMOps standards—model registries, CI/CD, experiment tracking, drift detection, A/B testing, latency optimization, and cost governance.
· Responsible AI & Security: Ensure model decisions comply with enterprise data security, privacy standards, and auditability required by the BFSI sector.
· Collaboration & Ownership: Translate complex business requirements into technical roadmaps, conduct rigorous code reviews, and mentor junior engineers.
Required Qualifications & Skills
· Education: B.Tech / M.Tech in Computer Science, AI/ML, Mathematics, or a related field—Tier-1 institutes (IIT, IIIT, NIT) strongly preferred.
· Experience: 2+ years of hands-on experience developing, deploying, and maintaining ML/Deep Learning or GenAI models in production environments.
· GenAI & NLP Stack: Hands-on experience with LLMs, embeddings, RAG architectures, and frameworks such as LangChain, LlamaIndex, or Hugging Face.
· Deep Learning Frameworks: Strong proficiency in PyTorch or TensorFlow, with deep knowledge of transformer architectures and modern NLP/CV models.
· Software & Data Engineering: Expert-level Python skills (pytest, Git, OOP, asynchronous programming), solid SQL proficiency, and familiarity with data workflows.
· Deployment & Cloud: Practical exposure to cloud platforms (AWS/GCP), containerization (Docker), API frameworks (FastAPI/Flask), and basic orchestration (Kubernetes).
Preferred Qualifications
· Prior domain experience in Fintech, RegTech, Identity Verification (KYC/AML), Fraud Intelligence, or B2B SaaS.
· Experience optimizing models for low latency and inference cost (e.g., ONNX, TensorRT, model quantization).
· Familiarity with workflow orchestrators such as Airflow, Prefect, or Kubeflow.
- Strong hands-on experience in Microsoft Azure Cloud.
- Good understanding of Azure services such as Compute, Storage, Event Hub, Event Subscription, Storage Queue, and PaaS services.
- Basic understanding of Azure AI Foundry and AI-related Azure service setup.
- Good Azure networking basics: VNet, subnet, routing, and basic troubleshooting.
- Strong knowledge of Terraform, especially:
- Terraform state
- plan / apply
- troubleshooting failures
- migration risks
- Terraform Enterprise concepts
- Strong Python coding capability, not just basic scripting.
- Experience using Python for API integration, automation, JSON/YAML handling, and internal tooling.
- Good understanding of CI/CD pipelines.
- Ability to troubleshoot pipeline failures.
- Comfortable with YAML and JSON.
- Ability to troubleshoot Azure infrastructure/platform issues.
- Ability to collect logs/evidence and coordinate with network/app/Microsoft support teams.
- Basic awareness of agentic AI / LLM concepts.
- Awareness of security and cost best practices.
Good to Have Skills
- Hands-on experience with Harness.
- Hands-on experience with Terraform Enterprise.
- Exposure to LangGraph / LangChain.
- Exposure to agentic AI workflows or skill creation.
- Exposure to Claude or enterprise LLM integrations.
- Knowledge of Azure ML Workspace, model registry, and managed endpoints.
- MLOps / LLMOps knowledge.
- FinOps / Azure cost optimization experience.
- Azure certifications: AZ-104, AZ-305, AZ-400, AZ-500.
Screening Priority:
Azure Cloud + Terraform + Python Coding + CI/CD Troubleshooting + YAML/JSON + Basic Agentic AI Awareness
EMBEDDED AI ENGINEERING POD
AI Implementation Engineer Role
Level: AI Implementation Engineer Senior / Advanced - 6+ years
Practice: Wissen GenAI
Locations: Mumbai / Bengaluru / New York - hybrid, embedded with delivery teams
Reports to: EMBEDDED AI PRACTICE Senior AI Engineering Specialist (Architect); Wissen GenAI Program Lead
Embedded inside enterprise delivery teams, you work closely with global, cross-regional teams to turn prioritized GenAI use cases into production software - building, integrating, and hardening Azure-based AI solutions and accelerating adoption within the teams you join.
You deliver production software and help the teams you join work faster.
As an embedded AI Implementation Engineer, you help convert prioritized use cases into shipped, governed, measurable software.
Key responsibilities
1. Build and ship.
Implement GenAI features end to end on Azure - RAG pipelines, agents, APIs, and UI integrations - against enterprise systems and data.
2. Embed and enable.
Work inside the delivery pods: pair with their engineers, remove blockers, and transfer GenAI skills so adoption sticks after you move on.
3. Productionize.
Add evaluation, observability, guardrails, caching, and CI/CD so prototypes become reliable, cost-efficient services.
4. Integrate securely.
Connect to enterprise data with correct access control, secrets management, and compliance with enterprise security standards and handling of sensitive data.
5. Iterate on quality.
Use evaluation results and user feedback to improve grounding, accuracy, latency, and cost.
6. Measure.
Track delivery and quality metrics that roll up to the program's targets.
Must-have qualifications
- 6+ years in software engineering, with 2+ years building GenAI/LLM applications in production.
- Strong Python (incl. async) and Java (the primary enterprise application stack; Spring a plus); solid API and systems design.
- Azure GenAI hands-on: Azure OpenAI, Azure AI Foundry, Azure AI Search for RAG, Azure AI Document Intelligence (IDP), and Prompt Flow.
- Agent frameworks: Microsoft Agent Framework / Semantic Kernel / AutoGen (or LangChain / LangGraph) and tool / function calling.
Preferred
- RAG fundamentals: embeddings, chunking, vector search, reranking, and grounding.
- Data platforms: Snowflake including Cortex AI (Cortex Search, LLM functions) and SQL, for accessing and grounding on enterprise data.
- Prompt engineering as versioned code; building and running evaluations.
- DevOps: Azure DevOps / GitHub Actions, Docker, AKS / Azure Functions, and observability.
- Financial services or other regulated environments.
- Front-end (React) for AI-assisted UX; streaming and token level operations.
- Azure AI Content Safety and responsible-AI practices.
- Certification: Azure AI Engineer Associate.
What success looks like - first 6 to 12 months
- Multiple GenAI features shipped to production within the embedded delivery pods.
- Measurable adoption and productivity uplift in the teams you support.
- Reusable components adopted from the architects' reference framework.
- Clear contribution to faster time-to-market and lower defect rates.
Design and develop Agentic AI systems using LLMs, tools, memory,
workflows, and MCP.
Build production-grade RAG pipelines, including ingestion, chunking,
embeddings, retrieval, reranking, and evaluation.
Implement context engineering strategies for improving LLM accuracy,
relevance, and reliability.
Develop and integrate MCP-based tools and services for AI agents.
Work with LLMs, SLMs, quantized models, and model optimization
techniques for efficient inference.
Develop scalable backend services and APIs for AI applications.
Design databases and data models supporting AI/agentic applications.
Implement AI observability covering latency, token usage, cost, failures,
quality, and agent/tool execution.
Apply AI governance and responsible AI practices, including security,
access control, data privacy, and auditability.
Optimize AI systems for latency, scalability, cost, and reliability.
Collaborate with engineering and product teams to take AI solutions from
POC to production.
Strong hands-on experience with GenAI, LLMs, and Agentic AI.
Experience building RAG applications.
Strong understanding of Context Engineering and prompt/context
optimization.
Role Overview
We are looking for a hands-on AI/ML Engineer to design, develop, and deploy
production-ready GenAI and Agentic AI applications. The role involves building
intelligent agents, RAG pipelines, AI APIs, backend services, and scalable AI
infrastructure with a strong focus on context engineering, observability,
governance, and model optimisation.
Key Responsibilities
Required Skills
Practical experience with MCP (Model Context Protocol).
Experience with frameworks such as LangChain, LangGraph,
LlamaIndex, or equivalent.
Knowledge of LLM/SLM deployment and quantization techniques.
Strong Python backend development experience.
Experience developing REST APIs using FastAPI/Flask or equivalent.
Strong understanding of SQL/NoSQL databases and database design.
Experience with vector databases such as Qdrant, Pinecone, Weaviate,
ChromaDB, or FAISS.
Understanding of AI observability, evaluation, monitoring, and
governance.
Experience with cloud platforms and production deployment is preferred.
Strong understanding of software engineering principles, Git, testing, and
CI/CD.















