Cutshort logo
For Employers
Its for a IT Service MNC logo
AI Platform Engineer
Its for a IT Service MNC
AI Platform Engineer

AI Platform Engineer at Its for a IT Service MNC · Remote, Bengaluru (Bangalore), Chennai, Mumbai · 6 - 10 years · Remote friendly · Posted 20 Jul 2026

Freelancer's logo

AI Platform Engineer

at Its for a IT Service MNC

Agency job
6 - 10 yrs
Best in industry
Remote, Bengaluru (Bangalore), Chennai, Mumbai
Skills
ELK
skill icongrafana
Observability
Agentic AI
Retrieval Augmented Generation (RAG)
LangChain
skill iconJava
skill iconPython

Primary Skills

  • Observability: ELK (Elasticsearch/Kibana), Prometheus, Grafana, PromQL
  • Automation: Java/Vert.x or Python (FastAPI), Shell/Bash, REST/SOAP APIs
  • Cloud & Platform: Docker, Kubernetes, Kafka, Redis
  • Reliability Engineering: Distributed Systems, Microservices, Event-Driven Architecture, DR & Incident Management
  • Stakeholder Management & Cross-functional Collaboration

Secondary Skills

  • Agentic AI: LangChain, LangGraph, RAG, MCP
  • LLM Integration & AI Frameworks
  • Python (FastAPI)

Key Responsibilities

  • Build and deploy LLM-powered Agentic AI solutions with tool calling and autonomous workflows.
  • Integrate AI capabilities into existing applications using modern AI frameworks.
  • Own platform reliability through SLAs, SLOs, error budgets, MTTD/MTTR, and operational governance.
  • Enhance observability using ELK, Prometheus, Grafana, and advanced alerting.
  • Lead incident response, RCA, disaster recovery, and resiliency initiatives.
  • Drive production readiness, automation, platform stability, and infrastructure optimization.
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

GYTWorkz Technologies Pvt Ltd
Hyderabad
5 - 10 yrs
₹10L - ₹40L / yr
Retrieval Augmented Generation (RAG)
LLM Evaluation Frameworks
Model Context Protocol (MCP)
Large Language Models (LLM) tuning
Fine-tuning LLMs
+6 more

Design and develop Agentic AI systems using LLMs, tools, memory,

workflows, and MCP.

Build production-grade RAG pipelines, including ingestion, chunking,

embeddings, retrieval, reranking, and evaluation.

Implement context engineering strategies for improving LLM accuracy,

relevance, and reliability.

Develop and integrate MCP-based tools and services for AI agents.

Work with LLMs, SLMs, quantized models, and model optimization

techniques for efficient inference.

Develop scalable backend services and APIs for AI applications.

Design databases and data models supporting AI/agentic applications.

Implement AI observability covering latency, token usage, cost, failures,

quality, and agent/tool execution.

Apply AI governance and responsible AI practices, including security,

access control, data privacy, and auditability.

Optimize AI systems for latency, scalability, cost, and reliability.

Collaborate with engineering and product teams to take AI solutions from

POC to production.

Strong hands-on experience with GenAI, LLMs, and Agentic AI.

Experience building RAG applications.

Strong understanding of Context Engineering and prompt/context

optimization.

Role Overview

We are looking for a hands-on AI/ML Engineer to design, develop, and deploy

production-ready GenAI and Agentic AI applications. The role involves building

intelligent agents, RAG pipelines, AI APIs, backend services, and scalable AI

infrastructure with a strong focus on context engineering, observability,

governance, and model optimisation.

Key Responsibilities

Required Skills

Practical experience with MCP (Model Context Protocol).

Experience with frameworks such as LangChain, LangGraph,

LlamaIndex, or equivalent.

Knowledge of LLM/SLM deployment and quantization techniques.

Strong Python backend development experience.

Experience developing REST APIs using FastAPI/Flask or equivalent.

Strong understanding of SQL/NoSQL databases and database design.

Experience with vector databases such as Qdrant, Pinecone, Weaviate,

ChromaDB, or FAISS.

Understanding of AI observability, evaluation, monitoring, and

governance.

Experience with cloud platforms and production deployment is preferred.

Strong understanding of software engineering principles, Git, testing, and

CI/CD.

Read more
Hyderabad, Bengaluru (Bangalore)
5 - 8 yrs
₹5L - ₹17L / yr
skill iconPython
skill iconJava
skill iconReact.js
Generative AI
skill iconSpring Boot
+1 more

Position: Senior/Lead Full Stack Engineer – Gen AI / Agentic AI

Experience: 7+ Years

Employment: Permanent Position

Location: Banglore / Hyderabad

Job Summary

We are looking for a Senior/Lead Full Stack Engineer – Gen AI / Agentic AI with strong hands-on experience in Python, React.js, MongoDB, Java/Spring Boot and Generative AI/Agentic AI.

The candidate should have experience designing and developing scalable enterprise applications and implementing production-grade LLM, RAG, AI Agent and multi-agent solutions.

Key Responsibilities

  • Design, develop and maintain scalable full-stack applications using Python, React.js, MongoDB and Java/Spring Boot.
  • Build production-grade Generative AI and Agentic AI applications using LLMs and modern AI frameworks.
  • Develop RAG pipelines, AI agents, tool calling, memory management, planning and agent orchestration.
  • Work with LangChain, LangGraph, MCP, vector databases and semantic search.
  • Develop Python-based APIs, microservices and asynchronous applications using FastAPI/Flask/Django.
  • Build REST APIs and event-driven microservices with focus on scalability, performance and resilience.
  • Integrate LLMs, embeddings, vector stores and external enterprise tools/services.
  • Implement prompt engineering, LLM evaluation, guardrails and AI observability.
  • Develop responsive front-end applications using React.js.
  • Work with MongoDB, SQL and hybrid data models.
  • Implement CI/CD pipelines and support cloud/OCP deployments.
  • Follow secure coding, testing, code quality and performance best practices.
  • Participate in architecture, technical design, code reviews and mentoring of team members.
  • Collaborate with business and technical stakeholders to translate requirements into scalable solutions.

Mandatory Skills

  1. Python
  2. React.js
  3. Gen AI / Agentic AI
  4. RAG + LLM
  5. LangChain / LangGraph
  6. MongoDB
  7. Java + Spring Boot
  8. REST APIs / Microservices
  9. Vector Databases / Embeddings
  10. MCP / AI Agent orchestration

Good to Have

  • FastAPI / Flask / Django
  • Kafka / Solace
  • Docker / Kubernetes
  • AWS / Azure / GCP / OCP
  • CI/CD – Jenkins / GitHub Actions
  • LLMOps / AI evaluation / observability
  • ELK / Grafana / Splunk / AppDynamics
  • SQL / NoSQL
  • Agile/Scrum
Read more
EMB Global
Rishu Dutta
Posted by Rishu Dutta
Gurugram
7 - 12 yrs
₹20L - ₹50L / yr
Retrieval Augmented Generation (RAG)
Agentic AI
Multi-agent Systems

Role Overview 

We are looking for an AI Engineer to design, build, and ship production AI systems, including agentic AI applications, for enterprise clients. This is a hands-on engineering role: you will write production code, build and evaluate models and agents, and work closely with architects and product teams to take solutions from prototype to scale. 


Key Responsibilities 

Design and build agentic AI systems: agent workflows, tool/function-calling, memory, and human-in-the-loop patterns. Build and productionise RAG pipelines, prompt-based applications, and LLM integrations across providers. Develop and maintain data and ML pipelines: feature engineering, model training, evaluation, and monitoring. Integrate AI systems with enterprise applications (CRMs, ERPs, ITSM tools) via APIs, events, and MCP-based tool servers. Implement guardrails, prompt-injection defences, and evaluation frameworks to keep AI systems safe and reliable in production. 

Write clean, tested, production-grade code and participate actively in code and design reviews. 

Collaborate with architects, product managers, and delivery teams to translate requirements into working AI solutions. Troubleshoot and optimise AI systems for accuracy, latency, and cost in production. 


Required Qualifications 

8–12 years of hands-on software engineering experience, with a strong, unbroken technical track record. Hands-on experience building and shipping AI/ML systems in production, not just POCs. 

Practical experience with agentic AI systems and at least one major agent framework (LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, Bedrock Agents/Strands, or Semantic Kernel). 

Experience with LLM/GenAI systems: RAG pipelines, prompt engineering, structured outputs, and tool calling across providers. 

Strong Python skills (TypeScript/Node.js a plus), with production-grade testing, CI/CD, and API design practices. Working knowledge of ML fundamentals: model evaluation, feature engineering, and experimentation. Cloud-native experience on AWS and/or Azure: containers, serverless, event backbones, and vector databases. Understanding of LLM safety and reliability practices: guardrails, prompt-injection defences, and observability. 



Read more
Leadsquared
Leadsquared
Agency job
via Right Hire by Vrishali Mishra
Bengaluru (Bangalore)
2 - 4 yrs
₹25L - ₹45L / yr
Large Language Models (LLM) tuning

About LeadSquared

LeadSquared is a leading sales execution and marketing automation platform trusted by 2,000+ businesses globally, including healthcare, education, financial services, and real estate. Headquartered in Bengaluru with offices across the US, UK, UAE, and Southeast Asia, we empower sales teams to close faster, smarter, and at scale.

Our AI team is at the forefront of integrating cutting-edge large language model capabilities into enterprise workflows — building intelligent agents, copilots, and automation systems that redefine how businesses operate.

Role Overview

We are looking for a Senior AI Engineer with hands-on experience building LLM-powered agents and agentic AI systems. You will design, develop, and deploy autonomous AI pipelines that solve complex, multi-step business problems — from lead qualification and follow-up automation to intelligent CRM workflows and beyond.

This role is ideal for someone who is deeply excited about the frontier of AI, can move fast, and wants their work to directly impact millions of sales professionals worldwide.

Key Responsibilities

•

Design and build LLM-powered agentic systems using frameworks such as LangChain, LlamaIndex, AutoGen, or CrewAI to automate complex, multi-step workflows.

•

Develop and maintain Retrieval-Augmented Generation (RAG) pipelines with vector databases (Pinecone, Weaviate, Chroma, pgvector) for domain-specific knowledge grounding.

•

Build and integrate tool-use and function-calling capabilities into AI agents, enabling dynamic interaction with internal APIs, databases, and third-party services.

•

Implement prompt engineering strategies including chain-of-thought, few-shot prompting, and structured output parsing to ensure reliable agent behavior.

•

Design evaluation frameworks and observability pipelines (LangSmith, Helicone, custom metrics) to monitor agent performance, accuracy, and cost.

•

Collaborate with product, sales, and domain teams to translate business requirements into AI-driven solutions and features.

•

Optimize LLM inference for latency and cost using techniques like caching, model distillation, quantization, and batching.

•

Stay current with the rapidly evolving LLM ecosystem and proactively propose improvements and new approaches.

•

Contribute to internal best practices, documentation, and knowledge-sharing across the engineering org.

Required Qualifications

Experience

•

2–4 years of professional software engineering experience, with at least 1–2 years focused on LLM/AI systems.

•

Proven experience shipping LLM-based products or agentic AI systems into production environments.

Technical Skills

•

Strong proficiency in Python and familiarity with async programming patterns for AI pipelines.

•

Hands-on experience with LLM APIs: OpenAI (GPT-4o), Anthropic (Claude), Google (Gemini), or open-source models (Llama, Mistral).

•

Experience with agentic frameworks: LangChain, LangGraph, LlamaIndex, AutoGen, CrewAI, or similar.

•

Solid understanding of RAG architectures, embedding models, and semantic search.

•

Experience with vector databases and similarity search infrastructure.

•

Knowledge of REST APIs, microservices architecture, and containerization (Docker/Kubernetes).

Problem-Solving & Mindset

•

Strong ability to decompose ambiguous, open-ended problems into structured AI system designs.

•

Experience with prompt debugging, LLM evaluation, and iterative refinement workflows.

•

Ability to balance research exploration with engineering pragmatism to ship reliable systems.

Preferred Qualifications

•

Experience with multi-agent orchestration and agent memory systems (short-term and long-term).

•

Familiarity with fine-tuning or RLHF workflows for domain adaptation.

•

Background in NLP, information retrieval, or conversational AI.

•

Prior experience in B2B SaaS or CRM domain is a plus.

•

Contributions to open-source AI/ML projects or published research/blogs.

•

Experience with cloud platforms: AWS, GCP, or Azure — particularly AI/ML services

Read more
Hiring for IT Consulting Firm (MNC)
Hiring for IT Consulting Firm (MNC)
Agency job
Pune, Nagpur
5 - 10 yrs
₹20L - ₹30L / yr
MLOps
DevOps
Artificial Intelligence (AI)
skill iconData Science
Data engineering
+5 more

Position Overview 

The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems. 

 

Responsibilities 

  • Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms 
  • Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning 
  • Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection 
  • Deploy, maintain and monitor containerized ML workloads 
  • Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows 
  • Implement AI-driven data validation, schema and concept drift detection and metadata management. 
  • Establish governance frameworks for AI systems, including bias detection, explainability, and auditability 
  • Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention 
  • Develop automated test suites for model performance, regression, edge cases and bias validation 
  • Monitor model KPIs (accuracy, precision, recall, latency, calibration) 
  • Ensure reproducability of experiments and production models 
  • Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption 

Qualifications 

  • Bachelor’s degree in computer science, Data Science, Information Systems, or a related field 
  • 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering 
  • Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure 
  • Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn 
Read more
Treosoft IT
Anish N
Posted by Anish N
Bengaluru (Bangalore)
3 - 5 yrs
₹10L - ₹20L / yr
skill iconPython
Generative AI
Agentic AI
LangChain
LlamaIndex
+3 more

Job Description:

We are looking for a hands-on AI Engineer with experience in Generative AI and Agentic AI to build and deploy production-ready AI solutions.

Key Responsibilities:

  • Develop and deploy GenAI and Agentic AI applications.
  • Build RAG pipelines, LLM workflows, and AI agents.
  • Develop solutions using Python, LangChain, LangGraph, LlamaIndex, or similar frameworks.
  • Implement tool calling, context retrieval, and LLM orchestration.
  • Integrate AI solutions with APIs and cloud platforms.
  • Work with AWS/Azure/GCP, Docker, and CI/CD.

Required Skills:

  • Strong Python programming skills.
  • 3+ years of GenAI/Agentic AI experience.
  • RAG and LLM orchestration.
  • LangChain / LangGraph / LlamaIndex / AutoGen / CrewAI / Semantic Kernel.
  • MCP and A2A knowledge.
  • Cloud, APIs, Docker, and CI/CD experience.

Preferred Experience:

Hands-on experience building and deploying production-ready AI solutions.

Read more
Remote only
4 - 8 yrs
₹10L - ₹30L / yr
Generative AI (GenAI)
Large Language Models (LLM) tuning
Fine-tuning LLMs
Retrieval Augmented Generation (RAG)
skill iconPython
+2 more

Forward-deployed engineers (FDEs) are Mactores' services layer. You embed with the customer's team, own outcomes from discovery through the production cutover, and personally carry the delivery commitment.

The agent platform we deploy absorbs 60–70% of engagement work, discovery, assessment, design, and testing. You absorb the judgment: target architecture, refactoring trade-offs, model selection, cutover strategy, and the decisions an agent platform cannot make. The agent absorbs scale. You absorb judgment. 

This is not a staff-augmentation seat and not an advisory role. You ship.

 

What you will do?

  • Deliver production agentic AI systems and AWS modernization engagements on committed dates across three pillars: Data Platform Modernization, Application & Database Modernization, and AI Agents for Apps.
  • Build and productionize AI agents, orchestration, retrieval pipelines, evaluation harnesses, observability running against real customer data, not demo data.
  • Convert existing products into agents: expose product functionality as callable tools for agent-to-agent composition, or replace form-and-click UX with agent-native, intent-driven interfaces.
  • Convert existing Business processes into agents: expose process functionality as callable tools for agent-to-agent composition, or replace form-and-click UX with agent-native, intent-driven interfaces.
  • Embed directly with customer engineering teams. Run architecture sessions, defend design decisions, and align stakeholders from VP Engineering to CTO.
  • Make agent decisions traceable and defensible, validation runs in parallel with live workloads, and outputs hold up to internal audit and regulators (HIPAA, PCI-DSS, FSI-grade governance where the vertical demands it).
  • Feed field experience back into the platform and practice: your deployment patterns, integration playbooks, and edge cases shape how we deliver.


What are we looking for?

  • Excellent communication skills (English) — verbal and written. Non-negotiable. You will present architecture to customer CTOs, write documents that hold up in audit, and defend judgment calls in the room. If you can build but not explain, this role is not a fit.
  • You have shipped production agentic AI systems on AWS. Not POCs, not notebooks — systems running in production for real users. This is the primary qualification. Be prepared to walk through what you shipped, the decisions you made, and what broke.
  • Deep understanding of agentic architecture — you can design an agent system from first principles and explain why each component exists:
  • Agent design patterns: single-agent vs. multi-agent systems, supervisor/orchestrator patterns, hierarchical agent topologies, planner–executor separation, and when each applies.
  • Orchestration: building and operating orchestrator agents that decompose tasks, route work to specialist agents or tools, and manage state across multi-step workflows (LangGraph, Strands Agents, CrewAI, or equivalent).
  • Memory: short-term/working memory (context management, conversation state) and long-term memory (episodic and semantic stores, vector- and graph-backed retrieval), and the production trade-offs of each.
  • Reflection and self-correction: critique loops, self-evaluation, retry-with-feedback patterns, and evaluation harnesses that catch agent failures before customers do.
  • Tool use and function calling: schema design, tool-selection reliability, error handling, and agent-to-agent composition.
  • RAG and retrieval pipelines: chunking, embedding, hybrid retrieval, reranking, and grounding agent decisions in customer data.
  • Strong AWS production experience: Amazon Bedrock and AWS AI services, plus core platform services (Lambda, API Gateway, DynamoDB, RDS/Aurora, Glue, EMR, Redshift, Kinesis, or similar depending on specialization).
  • Solid software engineering fundamentals Python, TypeScript, CI/CD, infrastructure-as-code, testing-driven development discipline.
  • Experience with data or application modernization (database migration, legacy refactoring, data platform builds) is a strong plus, since agents run against these workloads.
  • Indicative experience: roughly 3–10 years in engineering roles, with agentic AI / GenAI as your current day job. We have demonstrated agent-native expertise over tenure — an engineer with 3–4 years of hands-on agentic AI work typically outperforms a 12-year generalist on this work.


You'll be preferred if you've:

  • US English verbal and written fluency 
  • Delivery experience in one or more of our verticals: Financial Services, Healthcare & Life Sciences, Internet & Software, Manufacturing, or Telco/Media/Entertainment/Gaming/Sports.
  • Model tuning and fine-tuning: systematic prompt engineering and optimization; parameter-efficient fine-tuning (LoRA/QLoRA or similar); instruction tuning; working knowledge of RLHF/DPO; sound judgment on when to fine-tune vs. prompt vs. RAG; and evaluation of tuned models against baselines. Fine-tuning experience on Amazon Bedrock or SageMaker is a plus.
  • Experience with compliance-sensitive AI systems (HIPAA, PCI-DSS, SOC 2, data residency).
  • Knowledge graph, code-analysis (AST), or CDC/streaming experience (Debezium, Kafka/MSK).
  • Solid software engineering fundamentals — Java, C++, Go Lang, .Net, Rust
  • Prior customer-facing consulting or forward-deployed experience.
  • AWS certifications (Solutions Architect Professional, Machine Learning Specialty, or Data Analytics).


Why This Role?

  • You own outcomes, not tickets. FDEs carry the delivery commitment personally — architecture, judgment, and cutover are yours.
  • You work agent-native from day one. Our delivery model would not function without agents. You build with the platform, not around it.
  • You ship. Engagements measured in weeks to production, legacy retired, outcomes named. No archived pilots.
  • You compound. Field delivery informs the Aedeon platform roadmap; the platform's growth expands what you can deliver. Few engineering roles sit in that loop.


Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Akilandeswari Panneerselvam
Mumbai
8 - 10 yrs
₹7L - ₹15L / yr
skill iconJava
skill iconAmazon Web Services (AWS)
skill iconKubernetes
DevOps
Artificial Intelligence (AI)
+3 more

Job Title: Java AWS Kubernetes DevOps AI/ML

Experience: 8–10 Years

The candidate should have at least 1 year of experience in AI/ML and hands-on experience with the below technologies:

Java

AWS

Kubernetes

DevOps

AI/ML

MongoDB / PostgreSQL

Key-Value Caching

Vector Databases – ChromaDB / Pgvector

Read more
Unico Connect Private Limited
Mumbai
5 - 8 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Artificial Intelligence (AI)
Prompt engineering
LangGraph
+6 more

Senior AI Engineer

Code Generation, Agent Architecture & LLM Systems

📍 Mumbai (On-site) | Full-time | 5+ years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.

The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.


The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.

You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.


Responsibilities:


Agent Architecture

Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.

Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.


Code Generation Pipeline

Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.

Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.

Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.


Self-Repair Loop

Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.

This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.


Context Management

Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.

Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.

Own token budgeting and prompt caching strategy.


Prompt Engineering as a Discipline

Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).

Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.

Version prompts like code with changelogs and rollback capability.


Evaluation and Quality Measurement

Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.

Define regression gates that block quality-degrading changes from shipping.

Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.

This responsibility is non-negotiable at this level.


Model Strategy and Cost

Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.

Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.

Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.


Safety and Reliability of Agent Behaviour

Defend against prompt injection from user content and fetched web content.

Ensure secrets never appear in generated client code.

Define what the agent's tools may and may not do in collaboration with the platform team.

Contribute to output moderation and abuse-pattern awareness.


Mentorship and Engineering Standards

Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.

Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.


Requirements:


Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)

Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.

POCs, internal demos, and tutorial-grade work do not qualify.


5+ Years of Professional Software or AI Engineering Experience

With at least 3 years focused on LLM applications, AI engineering, or production AI systems.

Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.


Strong Python Proficiency and Service Development

Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.

Not notebook-only.


Depth Across LLM APIs and Agent Systems

Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).

Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Hands-on with tool calling, structured outputs, and multi-step reasoning.


Demonstrated, Systematic Evaluation Practice - Non-Negotiable

Must have built evaluation harnesses that gate production releases, not ad-hoc testing.

Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.

Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.


Cost Discipline for Production AI

Track record of measurable cost optimisation on production AI features.

Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.


AWS Working Knowledge

Hands-on with EC2, S3, IAM, and Docker.

Comfort with CI/CD workflows and deploying AI services.


Awareness of LLM Security Failure Modes

Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.


Nice to Have

  • Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
  • MCP server authoring
  • Open-source AI contributions
  • Published technical writing on LLM systems
  • Multi-modal model experience
  • Fine-tuning exposure (LoRA, QLoRA, PEFT)
Read more
Bengaluru (Bangalore), Hyderabad
5 - 8 yrs
₹2L - ₹15L / yr
skill iconData Science
skill iconPython
skill iconMachine Learning (ML)
Statistics/ML Fundamentals
GenAI/LLM
+8 more

Role Overview

We are looking for an experienced Data Scientist – Agentic AI with strong expertise in Python, Machine Learning, Generative AI, Large Language Models (LLMs), RAG and Agentic AI.

The ideal candidate should have hands-on experience in developing, fine-tuning, evaluating and deploying machine learning and GenAI solutions. The candidate should be comfortable working with open-source LLMs, LangChain/LangGraph, PySpark and AI observability/tracing frameworks.

The role involves building intelligent AI systems that can reason, use tools, retrieve information and execute multi-step tasks using Agentic AI architectures.

Mandatory Technical Skills

1. Data Science & Python

  • 5+ years of experience in Data Science / Machine Learning / AI.
  • Strong programming experience in Python.
  • Strong understanding of data analysis, feature engineering and statistical techniques.
  • Experience with Python ML and data science libraries such as:
  • NumPy
  • Pandas
  • Scikit-learn
  • Matplotlib / Seaborn
  • Good understanding of data preprocessing, exploratory data analysis and experimentation.

2. Machine Learning & Statistics

  • Strong understanding of Machine Learning fundamentals.
  • Experience with supervised and unsupervised learning techniques.
  • Knowledge of:
  • Regression
  • Classification
  • Clustering
  • Feature Engineering
  • Model Selection
  • Hyperparameter Tuning
  • Cross-validation
  • Strong understanding of Statistics / ML fundamentals.
  • Ability to interpret model performance and statistical results.

3. Generative AI / LLM

  • Strong hands-on experience with Generative AI and Large Language Models (LLMs).
  • Understanding of Transformer architecture and modern LLM-based applications.
  • Experience working with commercial or open-source LLMs.
  • Strong understanding of:
  • Prompt Engineering
  • Context Management
  • Embeddings
  • Tokenization
  • LLM inference
  • Hallucination mitigation

4. RAG – Retrieval Augmented Generation

  • Strong hands-on experience developing RAG applications.
  • Experience with:
  • Document ingestion
  • Chunking
  • Embeddings
  • Vector search
  • Semantic search
  • Retrieval pipelines
  • Context retrieval
  • Reranking
  • Ability to optimize RAG pipelines for relevance, accuracy and latency.
  • Experience integrating LLMs with enterprise knowledge sources.

5. Agentic AI

  • Hands-on experience building Agentic AI / AI Agent solutions.
  • Understanding of agent architecture and multi-step reasoning workflows.
  • Experience with:
  • AI Agents
  • Multi-Agent systems
  • Tool Calling
  • Function Calling
  • Agent orchestration
  • Planning and reasoning workflows
  • Memory
  • Workflow automation
  • Ability to build agents that can interact with tools, APIs, databases and external systems.

6. LangChain / LangGraph

  • Strong hands-on experience with LangChain and/or LangGraph.
  • Experience building LLM workflows and agent-based applications.
  • Understanding of:
  • Chains
  • Agents
  • Tools
  • State management
  • Graph-based workflows
  • Agent orchestration
  • Retrieval workflows
  • Experience designing scalable Agentic AI workflows.

7. LLM Fine-Tuning

  • Hands-on experience with LLM fine-tuning.
  • Understanding of techniques such as:
  • Supervised Fine-Tuning (SFT)
  • Parameter-Efficient Fine-Tuning
  • LoRA
  • QLoRA
  • Experience preparing datasets for fine-tuning.
  • Ability to evaluate fine-tuned models against baseline models.
  • Understanding of model optimization and inference considerations.

8. BERT / LLaMA / Open-Source LLMs

Experience working with one or more open-source / transformer-based models such as:

  • BERT
  • LLaMA / Llama
  • Mistral
  • Gemma
  • Qwen
  • Other open-source LLMs

Candidate should understand model loading, inference, fine-tuning and evaluation.

9. PySpark

  • Strong experience with PySpark for large-scale data processing.
  • Experience working with large datasets and distributed data processing.
  • Knowledge of:
  • Data transformations
  • Data cleaning
  • Aggregations
  • Joins
  • Spark SQL
  • Performance optimization
  • Ability to build scalable data processing pipelines.

10. Model Validation & Evaluation

  • Experience validating and evaluating ML and GenAI models.
  • Understanding of traditional ML evaluation metrics.
  • Experience evaluating LLM/RAG applications using relevant quality metrics.
  • Ability to compare model performance and identify areas for improvement.
  • Experience with:
  • Accuracy
  • Precision
  • Recall
  • F1 Score
  • ROC-AUC
  • Retrieval metrics
  • LLM response quality
  • Groundedness / relevance
  • Experience designing evaluation datasets and test cases is preferred.

11. AI Tracing / Observability

  • Experience with AI/LLM tracing and observability.
  • Ability to monitor AI applications in production.
  • Experience tracking:
  • LLM requests/responses
  • Latency
  • Token usage
  • Errors
  • Retrieval performance
  • Agent/tool execution
  • Model performance
  • Exposure to tools/frameworks such as LangSmith, OpenTelemetry, Arize Phoenix, MLflow or similar is preferred.

12. Model Deployment

  • Experience deploying ML/LLM/GenAI solutions into production.
  • Exposure to cloud and/or on-premise model deployment.
  • Experience with model serving, APIs and production inference.
  • Knowledge of deployment environments such as:
  • AWS
  • Azure
  • GCP
  • On-premise infrastructure
  • Experience with Docker, APIs and CI/CD is an advantage.

Key Responsibilities

  • Design, develop and deploy Data Science, Machine Learning and GenAI solutions.
  • Build production-ready RAG and Agentic AI applications.
  • Develop intelligent agents capable of tool calling, reasoning and multi-step task execution.
  • Build LLM-powered applications using LangChain/LangGraph.
  • Work with open-source LLMs including BERT, LLaMA and other transformer-based models.
  • Fine-tune LLMs for specific business use cases.
  • Develop scalable data processing pipelines using PySpark.
  • Perform data analysis, feature engineering and statistical modeling.
  • Develop and maintain model validation and evaluation frameworks.
  • Evaluate ML and LLM models using appropriate performance and quality metrics.
  • Implement AI tracing, monitoring and observability for production GenAI systems.
  • Deploy models and AI applications in cloud or on-premise environments.
  • Optimize model performance, response quality, latency and cost.
  • Troubleshoot issues related to model inference, retrieval, agents and LLM workflows.
  • Collaborate with Data Scientists, ML Engineers, Software Engineers and business stakeholders.
  • Convert business requirements into scalable AI/ML solutions.

Good to Have

  • Experience with Vector Databases such as:
  • FAISS
  • Pinecone
  • Weaviate
  • Milvus
  • Chroma
  • Azure AI Search
  • Experience with MLflow or similar ML lifecycle tools.
  • Experience with Docker/Kubernetes.
  • Experience with REST APIs / FastAPI.
  • Knowledge of cloud AI/ML services.
  • Experience with MLOps / LLMOps.
  • Experience with multi-agent frameworks other than LangChain/LangGraph.
  • Experience working with enterprise GenAI applications.

Ideal Candidate Profile

The ideal candidate should be a Data Scientist / ML Engineer with strong GenAI and Agentic AI experience, rather than a pure Python developer.

A strong candidate would typically have:

Data Science + Python + ML + Statistics + GenAI/LLM + RAG + Agentic AI + LangChain/LangGraph + LLM Fine-Tuning + Open-Source LLMs + PySpark + Model Evaluation + AI Observability + Model Deployment.

Core Mandatory Skills

Data Science, Python, Machine Learning, Statistics/ML Fundamentals, GenAI/LLM, RAG, Agentic AI, LangChain/LangGraph, LLM Fine-Tuning, BERT/LLaMA/Open-Source LLMs, PySpark, Model Validation/Evaluation, AI Tracing/Observability, Cloud/On-Prem Model Deployment.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos