Cutshort logo
For Employers
IAI solution  logo
AI/ML Engineer (LLMs, RAG & Agent Systems)
AI/ML Engineer (LLMs, RAG & Agent Systems)

AI/ML Engineer (LLMs, RAG & Agent Systems) at IAI solution · Bengaluru (Bangalore) · 1 - 2 years · ₹5L - ₹6L / yr · Raised funding · Posted 12 Jan 2026

IAI solution 's logo

AI/ML Engineer (LLMs, RAG & Agent Systems)

Anajli Kanojiya's profile picture
Posted by Anajli Kanojiya
1 - 2 yrs
₹5L - ₹6L / yr
Bengaluru (Bangalore)
Skills
Large Language Models (LLM)
LangChain
skill iconMongoDB
skill iconDocker

Job Title: AI/ML Engineer (LLMs, RAG & Agent Systems)

Location: Bangalore, India (On-site)

Type: Full-time


About the Role

As an AI/ML Engineer, you’ll be part of a small, fast-moving team focused on developing LLM-powered agentic systems that drive our next generation of AI products.

You’ll work on designing, implementing, and optimizing pipelines involving retrieval-augmented generation (RAG), multi-agent coordination, and tool-using AI systems.


Responsibilities

  • Design and implement components for LLM-based systems (retrievers, planners, memory, evaluators).
  • Build and maintain RAG pipelines using vector databases and embedding models.
  • Experiment with reasoning frameworks like ReAct, Tree of Thought, and Reflexion.
  • Collaborate with backend and infra teams to deploy and optimize agentic applications.
  • Research and experiment with open-source LLM frameworks to identify best-fit architectures.
  • Contribute to internal tools for evaluation, benchmarking, and scaling AI agents.


Required Skills

  • Strong foundation in ML/DL theory and implementation (PyTorch preferred).
  • Understanding of transformer architectures, embeddings, and LLM mechanics.
  • Practical exposure to prompt engineering, tool calling, and structured output design.
  • Experience in Python, Git/GitHub, and data processing pipelines.
  • Familiarity with RAG systems, vector databases, and API-based model inference.
  • Ability to write clean, modular, and reproducible code.


Preferred Skills

  • Experience with LangChain, LangGraph, Autogen, or CrewAI.
  • Hands-on with Hugging Face ecosystem (transformers, datasets, etc.).
  • Working knowledge of Redis, PostgreSQL, or MongoDB.
  • Experience with Docker and deployment workflows.
  • Familiarity with OpenAI, Anthropic, vLLM, or Ollama inference APIs.
  • Exposure to MLOps concepts like CI/CD, model versioning, or cloud (AWS/GCP/Azure).

What We Value

  • Deep understanding of core principles over surface-level familiarity with tools.
  • Ability to think like a researcher and execute like an engineer.
  • Collaborative mindset, building together, learning together.


Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About IAI solution

Founded :
2025
Type :
Product
Size :
20-100
Stage :
Raised funding

About

N/A

Company social profiles

N/A

Similar jobs (10)

Treosoft IT
Anish N
Posted by Anish N
Bengaluru (Bangalore)
3 - 5 yrs
₹10L - ₹20L / yr
skill iconPython
Generative AI
Agentic AI
LangChain
LlamaIndex
+3 more

Job Description:

We are looking for a hands-on AI Engineer with experience in Generative AI and Agentic AI to build and deploy production-ready AI solutions.

Key Responsibilities:

  • Develop and deploy GenAI and Agentic AI applications.
  • Build RAG pipelines, LLM workflows, and AI agents.
  • Develop solutions using Python, LangChain, LangGraph, LlamaIndex, or similar frameworks.
  • Implement tool calling, context retrieval, and LLM orchestration.
  • Integrate AI solutions with APIs and cloud platforms.
  • Work with AWS/Azure/GCP, Docker, and CI/CD.

Required Skills:

  • Strong Python programming skills.
  • 3+ years of GenAI/Agentic AI experience.
  • RAG and LLM orchestration.
  • LangChain / LangGraph / LlamaIndex / AutoGen / CrewAI / Semantic Kernel.
  • MCP and A2A knowledge.
  • Cloud, APIs, Docker, and CI/CD experience.

Preferred Experience:

Hands-on experience building and deploying production-ready AI solutions.

Read more
VY SYSTEMS PRIVATE LIMITED
Bengaluru (Bangalore)
4 - 12 yrs
₹4L - ₹25L / yr
skill iconData Science
skill iconPython
Large Language Models (LLM) tuning
RAG
Langchain

Support with design and build to prove out agentic AI solution flow by working with other data 

scientists and engineers to build, train Large Language Model (LLM) architectures, RAG 

systems, and autonomous agentic workflows 

Key qualifications: 

  

>> AI solution design & Development: Design Agentic AI solutions using RAG (Retrieval-

Augmented Generation) and orchestration frameworks like LangGraph or LangChain. 

  

>> Model Fine-Tuning: Solid understanding and experience with Pre-train, fine-tune, and 

optimize open-source like BERT, LLama, and other proprietary foundation models for domain-

specific tasks 

  

>> Solid Stats and ML foundations and (vibe) coding skills with Python, PySpark 

  

>>  Implement validation frameworks and tracing practices (using tools like Arize) to monitor 

agent behavior, guard against model drift, and ensure compliance 

  

>> Collaborate with Engineering to deploy models securely on cloud and on-prem ecosystems 

 

Read more
Leadsquared
Leadsquared
Agency job
via Right Hire by Vrishali Mishra
Bengaluru (Bangalore)
2 - 4 yrs
₹25L - ₹45L / yr
Large Language Models (LLM) tuning

About LeadSquared

LeadSquared is a leading sales execution and marketing automation platform trusted by 2,000+ businesses globally, including healthcare, education, financial services, and real estate. Headquartered in Bengaluru with offices across the US, UK, UAE, and Southeast Asia, we empower sales teams to close faster, smarter, and at scale.

Our AI team is at the forefront of integrating cutting-edge large language model capabilities into enterprise workflows — building intelligent agents, copilots, and automation systems that redefine how businesses operate.

Role Overview

We are looking for a Senior AI Engineer with hands-on experience building LLM-powered agents and agentic AI systems. You will design, develop, and deploy autonomous AI pipelines that solve complex, multi-step business problems — from lead qualification and follow-up automation to intelligent CRM workflows and beyond.

This role is ideal for someone who is deeply excited about the frontier of AI, can move fast, and wants their work to directly impact millions of sales professionals worldwide.

Key Responsibilities

•

Design and build LLM-powered agentic systems using frameworks such as LangChain, LlamaIndex, AutoGen, or CrewAI to automate complex, multi-step workflows.

•

Develop and maintain Retrieval-Augmented Generation (RAG) pipelines with vector databases (Pinecone, Weaviate, Chroma, pgvector) for domain-specific knowledge grounding.

•

Build and integrate tool-use and function-calling capabilities into AI agents, enabling dynamic interaction with internal APIs, databases, and third-party services.

•

Implement prompt engineering strategies including chain-of-thought, few-shot prompting, and structured output parsing to ensure reliable agent behavior.

•

Design evaluation frameworks and observability pipelines (LangSmith, Helicone, custom metrics) to monitor agent performance, accuracy, and cost.

•

Collaborate with product, sales, and domain teams to translate business requirements into AI-driven solutions and features.

•

Optimize LLM inference for latency and cost using techniques like caching, model distillation, quantization, and batching.

•

Stay current with the rapidly evolving LLM ecosystem and proactively propose improvements and new approaches.

•

Contribute to internal best practices, documentation, and knowledge-sharing across the engineering org.

Required Qualifications

Experience

•

2–4 years of professional software engineering experience, with at least 1–2 years focused on LLM/AI systems.

•

Proven experience shipping LLM-based products or agentic AI systems into production environments.

Technical Skills

•

Strong proficiency in Python and familiarity with async programming patterns for AI pipelines.

•

Hands-on experience with LLM APIs: OpenAI (GPT-4o), Anthropic (Claude), Google (Gemini), or open-source models (Llama, Mistral).

•

Experience with agentic frameworks: LangChain, LangGraph, LlamaIndex, AutoGen, CrewAI, or similar.

•

Solid understanding of RAG architectures, embedding models, and semantic search.

•

Experience with vector databases and similarity search infrastructure.

•

Knowledge of REST APIs, microservices architecture, and containerization (Docker/Kubernetes).

Problem-Solving & Mindset

•

Strong ability to decompose ambiguous, open-ended problems into structured AI system designs.

•

Experience with prompt debugging, LLM evaluation, and iterative refinement workflows.

•

Ability to balance research exploration with engineering pragmatism to ship reliable systems.

Preferred Qualifications

•

Experience with multi-agent orchestration and agent memory systems (short-term and long-term).

•

Familiarity with fine-tuning or RLHF workflows for domain adaptation.

•

Background in NLP, information retrieval, or conversational AI.

•

Prior experience in B2B SaaS or CRM domain is a plus.

•

Contributions to open-source AI/ML projects or published research/blogs.

•

Experience with cloud platforms: AWS, GCP, or Azure — particularly AI/ML services

Read more
Wissen Technology
at Wissen Technology
4 recruiters
Shakthi M
Posted by Shakthi M
Mumbai
5 - 13 yrs
Best in industry
AIML
skill iconPython
Lang chain

Role Overview

We are looking for an experienced AI/ML Engineer with strong expertise in Python, Generative AI, LLMs, LangChain, and LangGraph. The candidate will be responsible for designing and developing AI-powered applications, intelligent agents, and scalable LLM-based solutions.

Key Responsibilities

  • Design and develop AI/ML and Generative AI applications using Python and modern LLM technologies.
  • Build LLM-based applications and AI agents using LangChain and LangGraph.
  • Develop agentic workflows involving tool calling, memory, reasoning, and multi-step orchestration.
  • Integrate LLMs such as OpenAI, Azure OpenAI, Anthropic, Gemini, or other foundation models.
  • Develop RAG (Retrieval-Augmented Generation) pipelines using vector databases.
  • Work with embeddings, prompt engineering, semantic search, and document processing.
  • Develop scalable APIs and backend services using Python, FastAPI, or Flask.
  • Build and integrate AI solutions with existing applications and enterprise systems.
  • Deploy and maintain AI/ML solutions on cloud platforms.
  • Collaborate with data scientists, software engineers, and product teams to develop business-focused AI solutions.

Required Skills

  • Strong hands-on experience in Python programming.
  • Strong experience in AI/ML and Generative AI.
  • Hands-on experience with LLMs and LLM-based application development.
  • Strong experience with LangChain and/or LangGraph.
  • Experience building AI Agents / Agentic AI workflows.
  • Strong understanding of RAG, embeddings, vector databases, and prompt engineering.
  • Experience with vector databases such as FAISS, Pinecone, Chroma, Weaviate, or Azure AI Search.
  • Experience developing REST APIs using FastAPI/Flask.
  • Good understanding of Machine Learning, NLP, and deep learning concepts.
  • Experience with Azure / AWS / GCP cloud platforms.

Good to Have

  • Experience with multi-agent systems and agent orchestration.
  • Knowledge of MLOps / LLMOps.
  • Experience with Docker, Kubernetes, and CI/CD.
  • Knowledge of LLM evaluation, monitoring, observability, and AI governance.
  • Experience with Azure OpenAI, Azure AI Foundry, or AWS Bedrock.


Read more
GYTWorkz Technologies Pvt Ltd
Hyderabad
5 - 10 yrs
₹10L - ₹40L / yr
Retrieval Augmented Generation (RAG)
LLM Evaluation Frameworks
Model Context Protocol (MCP)
Large Language Models (LLM) tuning
Fine-tuning LLMs
+6 more

Design and develop Agentic AI systems using LLMs, tools, memory,

workflows, and MCP.

Build production-grade RAG pipelines, including ingestion, chunking,

embeddings, retrieval, reranking, and evaluation.

Implement context engineering strategies for improving LLM accuracy,

relevance, and reliability.

Develop and integrate MCP-based tools and services for AI agents.

Work with LLMs, SLMs, quantized models, and model optimization

techniques for efficient inference.

Develop scalable backend services and APIs for AI applications.

Design databases and data models supporting AI/agentic applications.

Implement AI observability covering latency, token usage, cost, failures,

quality, and agent/tool execution.

Apply AI governance and responsible AI practices, including security,

access control, data privacy, and auditability.

Optimize AI systems for latency, scalability, cost, and reliability.

Collaborate with engineering and product teams to take AI solutions from

POC to production.

Strong hands-on experience with GenAI, LLMs, and Agentic AI.

Experience building RAG applications.

Strong understanding of Context Engineering and prompt/context

optimization.

Role Overview

We are looking for a hands-on AI/ML Engineer to design, develop, and deploy

production-ready GenAI and Agentic AI applications. The role involves building

intelligent agents, RAG pipelines, AI APIs, backend services, and scalable AI

infrastructure with a strong focus on context engineering, observability,

governance, and model optimisation.

Key Responsibilities

Required Skills

Practical experience with MCP (Model Context Protocol).

Experience with frameworks such as LangChain, LangGraph,

LlamaIndex, or equivalent.

Knowledge of LLM/SLM deployment and quantization techniques.

Strong Python backend development experience.

Experience developing REST APIs using FastAPI/Flask or equivalent.

Strong understanding of SQL/NoSQL databases and database design.

Experience with vector databases such as Qdrant, Pinecone, Weaviate,

ChromaDB, or FAISS.

Understanding of AI observability, evaluation, monitoring, and

governance.

Experience with cloud platforms and production deployment is preferred.

Strong understanding of software engineering principles, Git, testing, and

CI/CD.

Read more
Service Co
Service Co
Agency job
via Vikash Technologies by Rishika Teja
Pune, Mumbai
6 - 12 yrs
₹20L - ₹45L / yr
Artificial Intelligence (AI)
Generative AI
Large Language Models (LLM)
Retrieval Augmented Generation (RAG)

Hiring for AI Engineer


Exp: 6 - 12 yrs

Edu : BE/B.Tech/MCA

Work Location : Pune / Mumbai


Skill Set:


Total experience ranging from 5–10 years in software engineering/AI roles

Min 5 years strong programming experience in Python or Typescript is a MUST

Min 2.5 years hands-on experience in AI with LLMs, RAG pipelines, and AI frameworks

2+ years shipping LLM systems in production

Experience with cloud platforms (AWS/Azure/GCP)

Read more
Staffnixcom
Mayank Choudhary
Posted by Mayank Choudhary
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Bengaluru (Bangalore)
3 - 5 yrs
₹20L - ₹25L / yr
Artificial Intelligence (AI)

Strong AI/ML Engineer Profile

Mandatory (Experience) : Must have 3+ years of experience in software engineering with atleast 1+ years in GenAI application development and production deployment

Mandatory (GenAI Application Development): Must have proven experience building GenAI applications covering RAG pipelines, multi-agent systems, Text2SQL, and fine-tuning

Mandatory (Production GenAI Deployment): Must have expertise deploying production-grade GenAI applications including model evaluation, optimisation, and ownership of full production rollouts

Mandatory (ML & Data Science Tooling): Must have strong hands-on experience with core ML and data science tools including pandas, scikit-learn, and PyTorch

Mandatory (Cloud ML Infrastructure): Must have experience building and deploying production-grade ML workloads on at least one of AWS, Azure, or GCP

Mandatory (Communication): Must have strong English communication skills with the ability to work across time zones and collaborate cross-functionally with product, engineering, and business stakeholders

Mandatory (Note 1) : Role is Hybrid, WFH flexibility as well upto 6 days a month

Mandatory (Note 2) : CTC is inclusive of 10% variable

Mandatory (Note 3): Candidates should be available to join within May 31st or June first week max

Read more
KDK Software
Priyanka Khandelwal
Posted by Priyanka Khandelwal
Jaipur
3 - 8 yrs
₹10L - ₹12L / yr
Generative AI (GenAI)
Large Language Models (LLM)
Retrieval Augmented Generation (RAG)

Job Description – AI Engineer (End-to-End Development & Deployment)


Role Summary

We are looking for an AI Engineer with hands-on experience in designing, developing, deploying, and maintaining Generative/Agentic AI solutions in production. The ideal candidate should have end-to-end ownership of AI applications, from development to deployment, monitoring, and optimization.

Key Responsibilities

●        Design, build, and deploy Generative/Agentic AI solutions.

●        Develop applications using LLMs, RAG, AI agents, and vector databases.

●        Build scalable APIs and integrate AI solutions with enterprise applications.

●        Implement CI/CD pipelines, containerization, and MLOps best practices.

●        Monitor, optimize, and maintain production AI systems.

●        Collaborate with cross-functional teams to deliver business-driven AI solutions.

Required Skills

●       Strong programming skills in Python.

●       Experience with vector databases (e.g., Pinecone, FAISS, ChromaDB) and graph memory systems

●       Knowledge of atleast one agent development framework: Google ADK (preferred), LangChain/LangGraph/LlamaIndex, CrewAI

●       Experience with LLMs, RAG, GenAI, AgenticAI Agents

●       Hands-on experience with FastAPI, and REST APIs.

●       Knowledge of Docker, Kubernetes, Git, CI/CD.

●       Experience with AWS, Azure, or GCP. 

●       Experience with security compliance, monitoring and observability tools such as AWS CloudWatch, Azure Monitor, Google Cloud Monitoring.


Read more
Bengaluru (Bangalore), Hyderabad
5 - 8 yrs
₹2L - ₹15L / yr
skill iconData Science
skill iconPython
skill iconMachine Learning (ML)
Statistics/ML Fundamentals
GenAI/LLM
+8 more

Role Overview

We are looking for an experienced Data Scientist – Agentic AI with strong expertise in Python, Machine Learning, Generative AI, Large Language Models (LLMs), RAG and Agentic AI.

The ideal candidate should have hands-on experience in developing, fine-tuning, evaluating and deploying machine learning and GenAI solutions. The candidate should be comfortable working with open-source LLMs, LangChain/LangGraph, PySpark and AI observability/tracing frameworks.

The role involves building intelligent AI systems that can reason, use tools, retrieve information and execute multi-step tasks using Agentic AI architectures.

Mandatory Technical Skills

1. Data Science & Python

  • 5+ years of experience in Data Science / Machine Learning / AI.
  • Strong programming experience in Python.
  • Strong understanding of data analysis, feature engineering and statistical techniques.
  • Experience with Python ML and data science libraries such as:
  • NumPy
  • Pandas
  • Scikit-learn
  • Matplotlib / Seaborn
  • Good understanding of data preprocessing, exploratory data analysis and experimentation.

2. Machine Learning & Statistics

  • Strong understanding of Machine Learning fundamentals.
  • Experience with supervised and unsupervised learning techniques.
  • Knowledge of:
  • Regression
  • Classification
  • Clustering
  • Feature Engineering
  • Model Selection
  • Hyperparameter Tuning
  • Cross-validation
  • Strong understanding of Statistics / ML fundamentals.
  • Ability to interpret model performance and statistical results.

3. Generative AI / LLM

  • Strong hands-on experience with Generative AI and Large Language Models (LLMs).
  • Understanding of Transformer architecture and modern LLM-based applications.
  • Experience working with commercial or open-source LLMs.
  • Strong understanding of:
  • Prompt Engineering
  • Context Management
  • Embeddings
  • Tokenization
  • LLM inference
  • Hallucination mitigation

4. RAG – Retrieval Augmented Generation

  • Strong hands-on experience developing RAG applications.
  • Experience with:
  • Document ingestion
  • Chunking
  • Embeddings
  • Vector search
  • Semantic search
  • Retrieval pipelines
  • Context retrieval
  • Reranking
  • Ability to optimize RAG pipelines for relevance, accuracy and latency.
  • Experience integrating LLMs with enterprise knowledge sources.

5. Agentic AI

  • Hands-on experience building Agentic AI / AI Agent solutions.
  • Understanding of agent architecture and multi-step reasoning workflows.
  • Experience with:
  • AI Agents
  • Multi-Agent systems
  • Tool Calling
  • Function Calling
  • Agent orchestration
  • Planning and reasoning workflows
  • Memory
  • Workflow automation
  • Ability to build agents that can interact with tools, APIs, databases and external systems.

6. LangChain / LangGraph

  • Strong hands-on experience with LangChain and/or LangGraph.
  • Experience building LLM workflows and agent-based applications.
  • Understanding of:
  • Chains
  • Agents
  • Tools
  • State management
  • Graph-based workflows
  • Agent orchestration
  • Retrieval workflows
  • Experience designing scalable Agentic AI workflows.

7. LLM Fine-Tuning

  • Hands-on experience with LLM fine-tuning.
  • Understanding of techniques such as:
  • Supervised Fine-Tuning (SFT)
  • Parameter-Efficient Fine-Tuning
  • LoRA
  • QLoRA
  • Experience preparing datasets for fine-tuning.
  • Ability to evaluate fine-tuned models against baseline models.
  • Understanding of model optimization and inference considerations.

8. BERT / LLaMA / Open-Source LLMs

Experience working with one or more open-source / transformer-based models such as:

  • BERT
  • LLaMA / Llama
  • Mistral
  • Gemma
  • Qwen
  • Other open-source LLMs

Candidate should understand model loading, inference, fine-tuning and evaluation.

9. PySpark

  • Strong experience with PySpark for large-scale data processing.
  • Experience working with large datasets and distributed data processing.
  • Knowledge of:
  • Data transformations
  • Data cleaning
  • Aggregations
  • Joins
  • Spark SQL
  • Performance optimization
  • Ability to build scalable data processing pipelines.

10. Model Validation & Evaluation

  • Experience validating and evaluating ML and GenAI models.
  • Understanding of traditional ML evaluation metrics.
  • Experience evaluating LLM/RAG applications using relevant quality metrics.
  • Ability to compare model performance and identify areas for improvement.
  • Experience with:
  • Accuracy
  • Precision
  • Recall
  • F1 Score
  • ROC-AUC
  • Retrieval metrics
  • LLM response quality
  • Groundedness / relevance
  • Experience designing evaluation datasets and test cases is preferred.

11. AI Tracing / Observability

  • Experience with AI/LLM tracing and observability.
  • Ability to monitor AI applications in production.
  • Experience tracking:
  • LLM requests/responses
  • Latency
  • Token usage
  • Errors
  • Retrieval performance
  • Agent/tool execution
  • Model performance
  • Exposure to tools/frameworks such as LangSmith, OpenTelemetry, Arize Phoenix, MLflow or similar is preferred.

12. Model Deployment

  • Experience deploying ML/LLM/GenAI solutions into production.
  • Exposure to cloud and/or on-premise model deployment.
  • Experience with model serving, APIs and production inference.
  • Knowledge of deployment environments such as:
  • AWS
  • Azure
  • GCP
  • On-premise infrastructure
  • Experience with Docker, APIs and CI/CD is an advantage.

Key Responsibilities

  • Design, develop and deploy Data Science, Machine Learning and GenAI solutions.
  • Build production-ready RAG and Agentic AI applications.
  • Develop intelligent agents capable of tool calling, reasoning and multi-step task execution.
  • Build LLM-powered applications using LangChain/LangGraph.
  • Work with open-source LLMs including BERT, LLaMA and other transformer-based models.
  • Fine-tune LLMs for specific business use cases.
  • Develop scalable data processing pipelines using PySpark.
  • Perform data analysis, feature engineering and statistical modeling.
  • Develop and maintain model validation and evaluation frameworks.
  • Evaluate ML and LLM models using appropriate performance and quality metrics.
  • Implement AI tracing, monitoring and observability for production GenAI systems.
  • Deploy models and AI applications in cloud or on-premise environments.
  • Optimize model performance, response quality, latency and cost.
  • Troubleshoot issues related to model inference, retrieval, agents and LLM workflows.
  • Collaborate with Data Scientists, ML Engineers, Software Engineers and business stakeholders.
  • Convert business requirements into scalable AI/ML solutions.

Good to Have

  • Experience with Vector Databases such as:
  • FAISS
  • Pinecone
  • Weaviate
  • Milvus
  • Chroma
  • Azure AI Search
  • Experience with MLflow or similar ML lifecycle tools.
  • Experience with Docker/Kubernetes.
  • Experience with REST APIs / FastAPI.
  • Knowledge of cloud AI/ML services.
  • Experience with MLOps / LLMOps.
  • Experience with multi-agent frameworks other than LangChain/LangGraph.
  • Experience working with enterprise GenAI applications.

Ideal Candidate Profile

The ideal candidate should be a Data Scientist / ML Engineer with strong GenAI and Agentic AI experience, rather than a pure Python developer.

A strong candidate would typically have:

Data Science + Python + ML + Statistics + GenAI/LLM + RAG + Agentic AI + LangChain/LangGraph + LLM Fine-Tuning + Open-Source LLMs + PySpark + Model Evaluation + AI Observability + Model Deployment.

Core Mandatory Skills

Data Science, Python, Machine Learning, Statistics/ML Fundamentals, GenAI/LLM, RAG, Agentic AI, LangChain/LangGraph, LLM Fine-Tuning, BERT/LLaMA/Open-Source LLMs, PySpark, Model Validation/Evaluation, AI Tracing/Observability, Cloud/On-Prem Model Deployment.

Read more
Unico Connect Private Limited
Remote, Mumbai
2 - 4 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Generative AI
LangGraph
FastAPI
+7 more

AI Engineer

LLMs, Agents & AI Services

📍 Mumbai (On-site) | Full-time | 2-4 years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

AI is core to how we design, deliver, and scale software for our customers.

We are hiring an AI Engineer for a dedicated client engagement building a complex production AI platform, working on the AI capabilities and agentic features at the core of the product.

The mandatory requirement for this role is at least one AI feature personally shipped to production for real users, with operational ownership.

The role suits someone who thinks quickly on solutioning, can take an ambiguous problem to a working prototype in days, and has the discipline to carry it through to production with predictable economics.

You will work alongside the Senior AI Engineer and the wider pod, with ownership of parts of the AI surface area of the product.


Responsibilities:

Solutioning and POCs

Translate ambiguous customer problems into working POCs at speed.

Pick the right model, framework, and architecture, and demonstrate value early before scaling investment.


LLM Application Development

Build AI features and services using LLM APIs from OpenAI, Anthropic, Google, and self-hosted open-weight models (Llama, Qwen, Mistral).

Choose the right model per use case based on cost, latency, capability, and context-window trade-offs.


Agentic System Design

Design and implement agentic workflows using LangGraph, CrewAI, AutoGen, LlamaIndex Agents, or custom orchestration.

Cover tool use, planning, memory, and multi-step reasoning appropriate to the problem.


API and Service Development

Build production AI services and APIs using Python and FastAPI.

Handle streaming responses, async processing, structured outputs, retries, and graceful degradation when models or tools fail.


Retrieval and Tool Integration

Implement RAG pipelines with vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma), embeddings, chunking strategies, hybrid search, and reranking.

Integrate external tools, internal APIs, and document sources through tool-calling and MCP-style patterns.


Cost Analysis and Unit Economics

Model the per-request and per-user cost of every AI feature before it ships.

Track token usage, prompt caching, batching, and model-routing strategies.

Drive measurable improvements in unit economics.


Production Hardening

Add observability and tracing (LangSmith, Langfuse, OpenTelemetry), guardrails, content safety checks, prompt injection defences, and fallback behaviour.


Prompt Engineering and Evaluation

Design, test, and iterate prompts with measured outcomes.

Build evaluation harnesses for accuracy, hallucination, latency, and cost.

Run benchmarks across models and prompt variants before locking in a design.


Requirements:

AI Feature Shipped to Production (Mandatory)

Must have personally built and shipped at least one AI feature that runs in production for real users, with operational ownership.

POCs, internal demos, and one-off scripts do not qualify.


2 to 4 Years of Professional Software or AI Engineering Experience

With at least one production AI feature owned end to end.


Strong Python Proficiency and API Development with FastAPI

Comfort with type hints, async, packaging, testing, streaming responses, and authentication.

Production-grade Python, not notebook-only code.


Hands-on Depth Across the LLM and Agent Stack

Working experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or self-hosted open-weight models (vLLM, Ollama, Together, Replicate).

Working familiarity with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Working knowledge of RAG, embeddings, and vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma).


Solutioning Speed and POC Velocity

Demonstrated ability to move from a fuzzy problem to a working prototype in days.

Strong instinct for what to build first, what to defer, and what to throw away.


Cost Discipline for Production AI

Ability to calculate, monitor, and optimise the cost of LLM APIs, tokens, embeddings, vector store usage, and infrastructure.

Treats unit economics as a first-class concern.


AWS Familiarity

Working knowledge of EC2, S3, IAM, and at least one of Bedrock, SageMaker, or equivalent.


Comfortable in a Fast-Moving Environment

Self-directed, comfortable with ambiguity, takes ownership without being asked, and ships under shifting priorities.


Strong Written and Spoken English Communication

Able to explain trade-offs to non-AI engineers, designers, product managers, and clients in plain language.


Nice to Have

  • fine-tuning or LoRA, QLoRA, PEFT exposure
  • MCP server authoring
  • eval framework experience (LangSmith, Promptfoo, Ragas, DeepEval)
  • open-source AI contributions
  • multi-modal models (vision, audio)
Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos