Cutshort logo
For Employers
EMB Global logo
Lead AI Engineer
Lead AI Engineer

Lead AI Engineer at EMB Global · Gurugram · 7 - 12 years · ₹20L - ₹50L / yr · Raised funding · Posted 24 Sep 2026

EMB Global's logo

Lead AI Engineer

Rishu Dutta's profile picture
Posted by Rishu Dutta
7 - 12 yrs
₹20L - ₹50L / yr
Gurugram
Skills
Retrieval Augmented Generation (RAG)
Agentic AI
Multi-agent Systems

Role Overview 

We are looking for an AI Engineer to design, build, and ship production AI systems, including agentic AI applications, for enterprise clients. This is a hands-on engineering role: you will write production code, build and evaluate models and agents, and work closely with architects and product teams to take solutions from prototype to scale. 


Key Responsibilities 

Design and build agentic AI systems: agent workflows, tool/function-calling, memory, and human-in-the-loop patterns. Build and productionise RAG pipelines, prompt-based applications, and LLM integrations across providers. Develop and maintain data and ML pipelines: feature engineering, model training, evaluation, and monitoring. Integrate AI systems with enterprise applications (CRMs, ERPs, ITSM tools) via APIs, events, and MCP-based tool servers. Implement guardrails, prompt-injection defences, and evaluation frameworks to keep AI systems safe and reliable in production. 

Write clean, tested, production-grade code and participate actively in code and design reviews. 

Collaborate with architects, product managers, and delivery teams to translate requirements into working AI solutions. Troubleshoot and optimise AI systems for accuracy, latency, and cost in production. 


Required Qualifications 

8–12 years of hands-on software engineering experience, with a strong, unbroken technical track record. Hands-on experience building and shipping AI/ML systems in production, not just POCs. 

Practical experience with agentic AI systems and at least one major agent framework (LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, Bedrock Agents/Strands, or Semantic Kernel). 

Experience with LLM/GenAI systems: RAG pipelines, prompt engineering, structured outputs, and tool calling across providers. 

Strong Python skills (TypeScript/Node.js a plus), with production-grade testing, CI/CD, and API design practices. Working knowledge of ML fundamentals: model evaluation, feature engineering, and experimentation. Cloud-native experience on AWS and/or Azure: containers, serverless, event backbones, and vector databases. Understanding of LLM safety and reliability practices: guardrails, prompt-injection defences, and observability. 



Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About EMB Global

Founded :
2018
Type :
Services
Size :
100-1000
Stage :
Raised funding

About

EMB Global is an AI-powered technology platform founded in 2018 and based in Gurugram, Haryana, India. The company connects enterprises with a curated network of IT service providers through its full-stack, AI-driven "Operating System for outsourcing" called Aura. This platform combines proprietary artificial intelligence with specialized human expertise to create more efficient tech teams. EMB Global offers a wide range of customizable IT and digital services, including software development, web design, digital marketing, and emerging technologies like Web 3.0. Their proprietary AI agent, Aura, automates the delivery of complex tech projects, while their staff augmentation services help businesses quickly scale development teams. The company serves startups, enterprises, and agencies across 23 countries, focusing on markets in India, the MENA region, and the USA.

Read more

Company social profiles

instagramlinkedintwitterfacebook

Similar jobs (10)

company logo
Priyanka Khandelwal
Posted by Priyanka Khandelwal
Jaipur
3 - 8 yrs
₹10L - ₹12L / yr
Generative AI (GenAI)
Large Language Models (LLM)
Retrieval Augmented Generation (RAG)

Job Description – AI Engineer (End-to-End Development & Deployment)


Role Summary

We are looking for an AI Engineer with hands-on experience in designing, developing, deploying, and maintaining Generative/Agentic AI solutions in production. The ideal candidate should have end-to-end ownership of AI applications, from development to deployment, monitoring, and optimization.

Key Responsibilities

●        Design, build, and deploy Generative/Agentic AI solutions.

●        Develop applications using LLMs, RAG, AI agents, and vector databases.

●        Build scalable APIs and integrate AI solutions with enterprise applications.

●        Implement CI/CD pipelines, containerization, and MLOps best practices.

●        Monitor, optimize, and maintain production AI systems.

●        Collaborate with cross-functional teams to deliver business-driven AI solutions.

Required Skills

●       Strong programming skills in Python.

●       Experience with vector databases (e.g., Pinecone, FAISS, ChromaDB) and graph memory systems

●       Knowledge of atleast one agent development framework: Google ADK (preferred), LangChain/LangGraph/LlamaIndex, CrewAI

●       Experience with LLMs, RAG, GenAI, AgenticAI Agents

●       Hands-on experience with FastAPI, and REST APIs.

●       Knowledge of Docker, Kubernetes, Git, CI/CD.

●       Experience with AWS, Azure, or GCP. 

●       Experience with security compliance, monitoring and observability tools such as AWS CloudWatch, Azure Monitor, Google Cloud Monitoring.


Read more
company logo
Shruti mujbaile
Posted by Shruti mujbaile
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Gurugram, Pune
6 - 12 yrs
₹8L - ₹25L / yr
Generative AI
Agentic AI
skill iconPython
skill iconMachine Learning (ML)

Location: Pune / Gurgaon

Position: AI Engineer

work mode: WFO


  Job Description.

​

 Job responsibilities:

  • Responsibility for design, implementation and deployment of Generative AI, Agentic frameworks at scale
  • Strong in programming - Python a
  • Previous experience of working on Computer Vision projects and VLM /VLAM models.
  • In depth awareness of Transformer architectures and End to End Deep neural networks
  • Full stack AI / ML development experience
  • Design, build & maintain efficient and reliable Agentic / Generative AI code leveraging pipelines
  • Hosting and deployment knowledge in GCP or AWS or Azure along with advanced engineering concepts to build user friendly UI interface for easy adoption.


    Requirements:

 ·      4 to 8 years overall years of experience (Agentic AI, Generative AI, VLM, VLAM and LLM) with significant exposure in Development, Architecture design, scaling and hosting in cloud.


    Must Have –

 ·      Architecting and solutioning experience with Python and FAST API, Agentic Ai frameworks, VLMs, VLAMs, Open source LLM’s and Code based LLM models at scale with - Langchain /      Ollama, embeddings, Memory      Management etc.,

·      Practical experience in implementing Explainable and ethical AI models  Practical experience in implementing frameworks like RAG/ CAG/ Self-reflective RAG etc.,

·      Experience in cloud hosting either AWS or Azure or GCP.

·      Experience in ML-OPS - Implement a feedback mechanism to continually improve the model over time through feedback loop and monitoring KPI’s in production.

·      Experience with Quantization and Kubernetes or docker


    Good to have

·      gRPC implementation to expose the API’s on a server for easy usage and good user interface

·      Streamlit front end creation

·      Experience with SAFe framework deliveries.


Read more
company logo
Remote only
4 - 8 yrs
₹10L - ₹30L / yr
Generative AI (GenAI)
Large Language Models (LLM) tuning
Fine-tuning LLMs
Retrieval Augmented Generation (RAG)
skill iconPython
+2 more

Forward-deployed engineers (FDEs) are Mactores' services layer. You embed with the customer's team, own outcomes from discovery through the production cutover, and personally carry the delivery commitment.

The agent platform we deploy absorbs 60–70% of engagement work, discovery, assessment, design, and testing. You absorb the judgment: target architecture, refactoring trade-offs, model selection, cutover strategy, and the decisions an agent platform cannot make. The agent absorbs scale. You absorb judgment. 

This is not a staff-augmentation seat and not an advisory role. You ship.

 

What you will do?

  • Deliver production agentic AI systems and AWS modernization engagements on committed dates across three pillars: Data Platform Modernization, Application & Database Modernization, and AI Agents for Apps.
  • Build and productionize AI agents, orchestration, retrieval pipelines, evaluation harnesses, observability running against real customer data, not demo data.
  • Convert existing products into agents: expose product functionality as callable tools for agent-to-agent composition, or replace form-and-click UX with agent-native, intent-driven interfaces.
  • Convert existing Business processes into agents: expose process functionality as callable tools for agent-to-agent composition, or replace form-and-click UX with agent-native, intent-driven interfaces.
  • Embed directly with customer engineering teams. Run architecture sessions, defend design decisions, and align stakeholders from VP Engineering to CTO.
  • Make agent decisions traceable and defensible, validation runs in parallel with live workloads, and outputs hold up to internal audit and regulators (HIPAA, PCI-DSS, FSI-grade governance where the vertical demands it).
  • Feed field experience back into the platform and practice: your deployment patterns, integration playbooks, and edge cases shape how we deliver.


What are we looking for?

  • Excellent communication skills (English) — verbal and written. Non-negotiable. You will present architecture to customer CTOs, write documents that hold up in audit, and defend judgment calls in the room. If you can build but not explain, this role is not a fit.
  • You have shipped production agentic AI systems on AWS. Not POCs, not notebooks — systems running in production for real users. This is the primary qualification. Be prepared to walk through what you shipped, the decisions you made, and what broke.
  • Deep understanding of agentic architecture — you can design an agent system from first principles and explain why each component exists:
  • Agent design patterns: single-agent vs. multi-agent systems, supervisor/orchestrator patterns, hierarchical agent topologies, planner–executor separation, and when each applies.
  • Orchestration: building and operating orchestrator agents that decompose tasks, route work to specialist agents or tools, and manage state across multi-step workflows (LangGraph, Strands Agents, CrewAI, or equivalent).
  • Memory: short-term/working memory (context management, conversation state) and long-term memory (episodic and semantic stores, vector- and graph-backed retrieval), and the production trade-offs of each.
  • Reflection and self-correction: critique loops, self-evaluation, retry-with-feedback patterns, and evaluation harnesses that catch agent failures before customers do.
  • Tool use and function calling: schema design, tool-selection reliability, error handling, and agent-to-agent composition.
  • RAG and retrieval pipelines: chunking, embedding, hybrid retrieval, reranking, and grounding agent decisions in customer data.
  • Strong AWS production experience: Amazon Bedrock and AWS AI services, plus core platform services (Lambda, API Gateway, DynamoDB, RDS/Aurora, Glue, EMR, Redshift, Kinesis, or similar depending on specialization).
  • Solid software engineering fundamentals Python, TypeScript, CI/CD, infrastructure-as-code, testing-driven development discipline.
  • Experience with data or application modernization (database migration, legacy refactoring, data platform builds) is a strong plus, since agents run against these workloads.
  • Indicative experience: roughly 3–10 years in engineering roles, with agentic AI / GenAI as your current day job. We have demonstrated agent-native expertise over tenure — an engineer with 3–4 years of hands-on agentic AI work typically outperforms a 12-year generalist on this work.


You'll be preferred if you've:

  • US English verbal and written fluency 
  • Delivery experience in one or more of our verticals: Financial Services, Healthcare & Life Sciences, Internet & Software, Manufacturing, or Telco/Media/Entertainment/Gaming/Sports.
  • Model tuning and fine-tuning: systematic prompt engineering and optimization; parameter-efficient fine-tuning (LoRA/QLoRA or similar); instruction tuning; working knowledge of RLHF/DPO; sound judgment on when to fine-tune vs. prompt vs. RAG; and evaluation of tuned models against baselines. Fine-tuning experience on Amazon Bedrock or SageMaker is a plus.
  • Experience with compliance-sensitive AI systems (HIPAA, PCI-DSS, SOC 2, data residency).
  • Knowledge graph, code-analysis (AST), or CDC/streaming experience (Debezium, Kafka/MSK).
  • Solid software engineering fundamentals — Java, C++, Go Lang, .Net, Rust
  • Prior customer-facing consulting or forward-deployed experience.
  • AWS certifications (Solutions Architect Professional, Machine Learning Specialty, or Data Analytics).


Why This Role?

  • You own outcomes, not tickets. FDEs carry the delivery commitment personally — architecture, judgment, and cutover are yours.
  • You work agent-native from day one. Our delivery model would not function without agents. You build with the platform, not around it.
  • You ship. Engagements measured in weeks to production, legacy retired, outcomes named. No archived pilots.
  • You compound. Field delivery informs the Aedeon platform roadmap; the platform's growth expands what you can deliver. Few engineering roles sit in that loop.


Read more
company logo
HR  GYTWorkz
Posted by HR GYTWorkz
Hyderabad
5 - 10 yrs
₹10L - ₹40L / yr
Retrieval Augmented Generation (RAG)
LLM Evaluation Frameworks
Model Context Protocol (MCP)
Large Language Models (LLM) tuning
Fine-tuning LLMs
+6 more

Design and develop Agentic AI systems using LLMs, tools, memory,

workflows, and MCP.

Build production-grade RAG pipelines, including ingestion, chunking,

embeddings, retrieval, reranking, and evaluation.

Implement context engineering strategies for improving LLM accuracy,

relevance, and reliability.

Develop and integrate MCP-based tools and services for AI agents.

Work with LLMs, SLMs, quantized models, and model optimization

techniques for efficient inference.

Develop scalable backend services and APIs for AI applications.

Design databases and data models supporting AI/agentic applications.

Implement AI observability covering latency, token usage, cost, failures,

quality, and agent/tool execution.

Apply AI governance and responsible AI practices, including security,

access control, data privacy, and auditability.

Optimize AI systems for latency, scalability, cost, and reliability.

Collaborate with engineering and product teams to take AI solutions from

POC to production.

Strong hands-on experience with GenAI, LLMs, and Agentic AI.

Experience building RAG applications.

Strong understanding of Context Engineering and prompt/context

optimization.

Role Overview

We are looking for a hands-on AI/ML Engineer to design, develop, and deploy

production-ready GenAI and Agentic AI applications. The role involves building

intelligent agents, RAG pipelines, AI APIs, backend services, and scalable AI

infrastructure with a strong focus on context engineering, observability,

governance, and model optimisation.

Key Responsibilities

Required Skills

Practical experience with MCP (Model Context Protocol).

Experience with frameworks such as LangChain, LangGraph,

LlamaIndex, or equivalent.

Knowledge of LLM/SLM deployment and quantization techniques.

Strong Python backend development experience.

Experience developing REST APIs using FastAPI/Flask or equivalent.

Strong understanding of SQL/NoSQL databases and database design.

Experience with vector databases such as Qdrant, Pinecone, Weaviate,

ChromaDB, or FAISS.

Understanding of AI observability, evaluation, monitoring, and

governance.

Experience with cloud platforms and production deployment is preferred.

Strong understanding of software engineering principles, Git, testing, and

CI/CD.

Read more
company logo
Meenal Patil
Posted by Meenal Patil
Pune
3 - 4 yrs
₹5L - ₹15L / yr
Agent development
legacy migration
AI Copilot
Claude AI APP

Role: AI Developer

Experience: 3–4 Years

Employment Type: Full-Time

Location: Goregaon, Mumbai


About the Role

We are looking for an experienced AI Developer with 3–4 years of software development experience and strong hands-on exposure to Generative AI, AI Agents, Copilots, and AI-powered application development.

The candidate will be responsible for building production-ready AI solutions, developing agentic workflows, modernizing legacy applications, and integrating LLM capabilities into enterprise applications.


Key Responsibilities

  • Design, develop, and deploy AI Agents and agentic workflows for enterprise use cases.
  • Build AI Copilots and LLM-powered applications using modern AI frameworks and APIs.
  • Develop RAG-based applications using embeddings, vector databases, and enterprise data.
  • Work on legacy application migration and modernization, leveraging AI-assisted development and code transformation techniques.
  • Analyze legacy codebases and design strategies for AI-driven migration, refactoring, and modernization.
  • Integrate LLMs with enterprise applications, APIs, databases, and third-party systems.
  • Implement tool calling, function calling, multi-agent workflows, and workflow automation.
  • Perform prompt engineering, context optimization, model evaluation, and AI application testing.
  • Take ownership of AI solutions from POC and prototyping through production deployment.
  • Collaborate with product managers, architects, and engineering teams to convert business requirements into scalable AI solutions.
  • Stay updated with emerging technologies in Generative AI, Agentic AI, LLMs, and AI-assisted software development.


Required Skills

  • 3–4 years of professional software development experience.
  • Strong proficiency in Python and/or JavaScript/TypeScript.
  • Hands-on experience developing Generative AI / LLM-based applications.
  • Strong understanding of AI Agents, RAG, Prompt Engineering, LLM APIs, and embeddings.
  • Experience with frameworks such as LangChain, LangGraph, Semantic Kernel, AutoGen, or equivalent.
  • Experience working with REST APIs, databases, Git, and cloud environments.
  • Hands-on experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, or equivalent.
  • Good understanding of software architecture, debugging, testing, and deployment practices.


Good to Have

  • Experience with Microsoft Copilot / Copilot Studio.
  • Experience working with Claude, OpenAI, Gemini, Azure OpenAI, or open-source LLMs.
  • Experience in legacy application migration, modernization, or code conversion.
  • Knowledge of Azure AI / AWS / Google Cloud AI services.
  • Experience with MCP, multi-agent systems, tool calling, and AI orchestration.
  • Experience building enterprise-grade AI solutions with focus on security, scalability, and performance.


Read more
company logo
Deep Bhadja
Posted by Deep Bhadja
Remote, Ahmedabad
3 - 6 yrs
₹8L - ₹12L / yr
Artificial Intelligence (AI)
Build automation

Role Overview:

As an AI Executor/AI Automation Engineer, you will be responsible for designing and integrating AI capabilities into production systems using Python and key ML libraries. This role requires a strong backend development foundation and a proven track record of deploying AI use cases using tools like TensorFlow, Keras, or OpenAI APIs. You'll work cross-functionally to deliver scalable AI-driven solutions.

 

Key Responsibilities:

  • Design and develop backend solutions using Python, with a focus on AI-driven features.
  • Implement and integrate AI/ML models using tools like OpenAI, Hugging Face, or Lang Chain.
  • Use core Python libraries (NumPy, Pandas, TensorFlow, Keras) to process data, train, or implement models.
  • Translate business needs into AI use cases and deliver working solutions.
  • Collaborate with product, engineering, and data teams to define integration workflows.
  • Develop REST APIs and micro services to deploy AI components within applications.
  • Maintain and optimize AI systems for scalability, performance, and reliability.
  • Keep pace with advancements in the AI/ML landscape and evaluate tools for continuous improvement.

 

Required Skills & Qualifications:

  • 2+ years of professional experience as an AI/ML Engineer, including strong backend development expertise in Python.
  • Proficiency in libraries such as NumPy, Pandas, TensorFlow, and Keras
  • Practical exposure to AI platforms/APIs (e.g., OpenAI, LangChain, Hugging Face)
  • Solid understanding of REST APIs, micro services, and integration practices
  • Ability to work independently in a remote setup with strong communication and ownership
  • Excellent problem-solving and debugging capabilities
  • Experience with the MERN stack will be an added advantage.


Read more
company logo
Stuti Jain
Posted by Stuti Jain
Hyderabad
7 - 10 yrs
₹25L - ₹35L / yr
Retrieval Augmented Generation (RAG)
skill iconAmazon Web Services (AWS)

Location: Hyderabad, India. Based at the KnackLabs headquarters, with occasional travel to client locations for workshops and reviews. This role does not involve extended onsite deployments.

About the Role

You will work as an AI Architect who designs the systems behind our client engagements: AI agents, RAG systems, automation platforms, and the conventional backend systems around them.

This is a hands-on design role, not a slideware role. You will scope architectures with clients, make the hard technical decisions, defend them in review, and stay accountable for how the systems perform in production.


You will work directly with clients. Everyone at KnackLabs does. You will sit in design discussions with client engineering teams, present architecture decisions to technical and business stakeholders, and answer for the choices you make.


A full KnackLabs engineering team in Hyderabad builds with you. You own the technical design and the quality of what ships.

What you'll own

  1. Architecture - Design AI agents, RAG systems, integrations, and the scalable backend systems around them, for multiple client engagements.
  2. Technical scoping - Work directly with clients to turn a business problem into a system design, with clear trade-offs and clear reasons.
  3. Scale and reliability - Make sure what we build handles real load: data stores, queues, caching, horizontal scaling, and fault tolerance.
  4. Design reviews - Review designs and builds across engagements. Set the technical bar and hold it.
  5. Evaluation strategy - Define how we measure accuracy, safety, latency, and cost for the AI systems we ship.
  6. Guiding engineers - Raise the level of the engineers building with you, through reviews and direct pairing.
  7. Feedback to the platform - Feed what you learn across engagements back into our platform and internal tools.

What we are looking for

  1. Around 7 or more years of software engineering experience, including direct work with customers on design or delivery.
  2. Full-stack development experience with strength in backend technologies.
  3. Experience designing and building scalable applications. You understand how large-scale distributed systems work: data partitioning, queues, caching, horizontal scaling, and fault tolerance.
  4. At least 2 years of strong, hands-on AI experience with large language models in production.
  5. You build with AI coding tools like Claude Code or Codex as your default way of working. You understand Claude Skills, have written skills yourself, use them actively, and have contributed to them.
  6. Hands-on experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
  7. Hands-on experience building AI agents.
  8. Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
  9. Experience with at least one cloud platform (AWS, Azure, or GCP).
  10. Clear communication. You can explain an architecture decision to an engineer and to a business leader, and defend it under questioning.
  11. High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a design.

Nice to have

  1. Experience building evaluations to measure accuracy, safety, latency, and cost.
  2. Experience with observability and tracing tools such as LangSmith or Braintrust.
  3. Experience with on-premises or private cloud (VPC) deployments.
  4. Experience deploying AI systems in regulated industries such as insurance, banking, or the public sector.
  5. Experience with data engineering and pipelines.
  6. A history of side projects, open source contributions, or products you shipped end-to-end.
  7. Experience working at a consulting or professional services firm in a client-facing delivery role.

Stack and tools

  1. Languages: Python and TypeScript.
  2. Models: Claude and other frontier or open-source models, chosen to fit the customer.
  3. AI patterns: RAG, agents, prompt engineering, skills, and evaluations.
  4. Vector and retrieval: vector databases and retrieval pipelines.
  5. Cloud: AWS, Azure, or GCP, on public or private cloud.
  6. Integration: REST APIs and enterprise system connectors.


Read more
company logo
Anish N
Posted by Anish N
Bengaluru (Bangalore)
3 - 5 yrs
₹10L - ₹20L / yr
skill iconPython
Generative AI
Agentic AI
LangChain
LlamaIndex
+3 more

Job Description:

We are looking for a hands-on AI Engineer with experience in Generative AI and Agentic AI to build and deploy production-ready AI solutions.

Key Responsibilities:

  • Develop and deploy GenAI and Agentic AI applications.
  • Build RAG pipelines, LLM workflows, and AI agents.
  • Develop solutions using Python, LangChain, LangGraph, LlamaIndex, or similar frameworks.
  • Implement tool calling, context retrieval, and LLM orchestration.
  • Integrate AI solutions with APIs and cloud platforms.
  • Work with AWS/Azure/GCP, Docker, and CI/CD.

Required Skills:

  • Strong Python programming skills.
  • 3+ years of GenAI/Agentic AI experience.
  • RAG and LLM orchestration.
  • LangChain / LangGraph / LlamaIndex / AutoGen / CrewAI / Semantic Kernel.
  • MCP and A2A knowledge.
  • Cloud, APIs, Docker, and CI/CD experience.

Preferred Experience:

Hands-on experience building and deploying production-ready AI solutions.

Read more
company logo
Orenda Finserv
Posted by Orenda Finserv
Ahmedabad
3 - 5 yrs
₹7L - ₹11L / yr
skill iconMachine Learning (ML)
Model Serving
Vision Models
skill iconPython
RESTful APIs
+2 more

About the role

We are building AI systems that read, understand and act on real business documents, bank statements, financial reports, policy documents and forms and putting them into production where accuracy and cost both matters.

This is not a research role and it is not a prompt-writing role. You will own features end to end: pick and deploy open-source models, build the pipelines around them, measure whether they actually work on our documents, drive the cost per document down, and keep the whole thing running in production.

You will work closely with the engineering and product teams, and your work will be directly used by business users from day one.


What you will do

Deploy and evaluate open-source models

  • Select, deploy and benchmark open-source LLMs and vision-language models for specific, narrow use cases not general chat.
  • Build evaluation sets from real documents and define what "good" means numerically (field-level accuracy, extraction recall, hallucination rate) before shipping.
  • Run structured comparisons between models and approaches, and write up the trade-offs so the team can make a decision.
  • Apply quantization, batching and other optimizations to fit models into a sensible GPU budget.

Build and optimize AI orchestration

  • Design multi-step pipelines that combine deterministic code, ML models and LLM calls and know when not to use an LLM.
  • Optimize for latency, cost and reliability: caching, batching, request routing, fallback tiers, retries and graceful degradation.
  • Instrument pipelines so failures are visible and traceable rather than silent.

Ship to production

  • Package models and services with Docker, expose them behind clean APIs, and deploy them to our GPU and CPU infrastructure.
  • Handle the unglamorous production concerns: cold starts, timeouts, concurrency limits, versioning, rollback and monitoring.
  • Own on-call-style responsibility for the AI features you build, including cost tracking.


Must-have skills


Programming & engineering

  • Strong Python: type hints, async/await, dataclasses/Pydantic, clean module design, testing.
  • REST API development with FastAPI (or Flask/Django with a willingness to move to FastAPI).
  • Git, code review discipline, and the ability to write code someone else can maintain.
  • Comfortable in Linux and on the command line.

Machine learning fundamentals

  • Working knowledge of PyTorch and the Hugging Face ecosystem (transformers, tokenizers, accelerate).
  • Understanding of inference-time concepts: tokenization, context windows, batching, precision (FP16/BF16/INT8), memory footprint.
  • Ability to read a model card and a paper well enough to judge whether a model fits a use case.

Document processing

  • Hands-on experience with at least two of: pypdfium2, PyMuPDF, pdfplumber, pdfminer.six, Docling, Unstructured, Surya, DocTR, LayoutLM family.
  • Practical OCR experience (Tesseract, PaddleOCR, or a cloud OCR) and an understanding of when OCR is the wrong tool.
  • Experience extracting tables from PDFs and dealing with merged cells, multi-line rows, and inconsistent column layouts.


Strongly preferred

You will be a much stronger candidate with any of these. We do not expect all of them.

Model serving & optimization

  • vLLM, TGI, Ollama, llama.cpp, or Triton Inference Server.
  • Quantization formats and tooling: GGUF, AWQ, GPTQ, bitsandbytes, ONNX Runtime, INT8 export.
  • Serverless GPU platforms: Modal, RunPod, Replicate, Baseten including cold-start and container-lifecycle management.
  • LoRA / QLoRA fine-tuning with PEFT for narrow, task-specific improvements.

Vision-language models

  • Practical use of open VLMs: Qwen2.5-VL, InternVL, Granite Vision, Molmo, Phi-Vision, or similar.
  • Awareness of where VLMs hallucinate especially on numeric and financial content and patterns for constraining them (using the model for layout only, sourcing values from the text layer, constrained decoding).

Orchestration & pipelines

  • Workflow orchestration: Dagster, Airflow, Prefect, or Temporal.
  • Async job patterns: Celery, RQ, or platform-native spawn/poll patterns.
  • LLM orchestration frameworks (LangGraph, LlamaIndex, Haystack) with the judgement to know when plain Python is a better answer.
  • Structured output enforcement: Instructor, Outlines, XGrammar, JSON schema / tool-use modes.

Evaluation & observability

  • Building golden datasets and regression suites for extraction tasks.
  • Eval tooling: promptfoo, DeepEval, Ragas, or in-house harnesses.
  • LLM tracing and monitoring: Langfuse, Arize Phoenix, LangSmith, OpenTelemetry.

Nice extras

  • Rule engines and policy evaluation (Open Policy Agent / Rego, Drools, rule-engine).
  • Experience in fintech, lending, insurance or accounting documents.
  • Handling of PII and data-security practices in document pipelines.
  • Contributions to open-source ML or document-processing projects.


Why join us

  • Real production ownership from month one your work goes to actual users, not a demo.
  • Genuinely hard technical problems in document AI, not wrappers over an API.
  • Small team, short decision cycles, direct access to leadership.
  • Budget and freedom to evaluate and adopt new open-source models as they land.


To apply: send your CV along with a short note on one AI system you have taken to production what it did, what the accuracy was, and what broke.


Read more
company logo
Banu S
Posted by Banu S
Bengaluru (Bangalore)
4 - 12 yrs
₹4L - ₹25L / yr
skill iconData Science
skill iconPython
Large Language Models (LLM) tuning
RAG
Langchain

Support with design and build to prove out agentic AI solution flow by working with other data 

scientists and engineers to build, train Large Language Model (LLM) architectures, RAG 

systems, and autonomous agentic workflows 

Key qualifications: 

  

>> AI solution design & Development: Design Agentic AI solutions using RAG (Retrieval-

Augmented Generation) and orchestration frameworks like LangGraph or LangChain. 

  

>> Model Fine-Tuning: Solid understanding and experience with Pre-train, fine-tune, and 

optimize open-source like BERT, LLama, and other proprietary foundation models for domain-

specific tasks 

  

>> Solid Stats and ML foundations and (vibe) coding skills with Python, PySpark 

  

>>  Implement validation frameworks and tracing practices (using tools like Arize) to monitor 

agent behavior, guard against model drift, and ensure compliance 

  

>> Collaborate with Engineering to deploy models securely on cloud and on-prem ecosystems 

 

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos