Applied AI Engineer (AI/ML & LLM Deployment) at NeoGenCode Technologies Pvt Ltd · Remote only · 3 - 6 years · ₹8L - ₹15L / yr · Raised funding · Remote only · Posted 25 May 2026

Applied AI Engineer (AI/ML & LLM Deployment)
Job Title : Applied AI Engineer (AI/ML & LLM Deployment)
Experience : 4+ Years
Location : Remote
Job Summary :
We are looking for an Applied AI Engineer with strong expertise in Large Language Models (LLMs) to design, optimize, and deploy advanced AI solutions. The role focuses on LLM integration, fine-tuning, and performance optimization for scalable production use.
Mandatory Skills :
LLM Integration, Prompt Engineering, LangChain/LlamaIndex, Fine-tuning & LoRA, Model Quantization, Local Deployment, Tokenizers, Performance Optimization.
Key Responsibilities :
- Integrate GPT-4, Claude, and open-source LLMs into applications.
- Design and implement prompt engineering for context-specific, tone-appropriate outputs.
- Build AI pipelines using LangChain, LlamaIndex, or similar frameworks.
- Perform fine-tuning and LoRA (Low-Rank Adaptation) to customize models.
- Apply model quantization (GGML, GPTQ, bitsandbytes) for efficient deployment.
- Deploy models with Hugging Face Transformers, vLLM, llama.cpp.
- Optimize tokenization and inference performance for low latency.
Required Skills :
- Strong hands-on experience with LLMs (OpenAI GPT-4, Claude, open-source models).
- Proficiency in Prompt Engineering and tone-specific output design.
- Experience with LangChain, LlamaIndex, or other orchestration frameworks.
- Knowledge of fine-tuning, LoRA techniques, and quantization methods.
- Familiarity with deployment frameworks (Transformers, vLLM, llama.cpp).
- Strong understanding of tokenizers and efficient text preprocessing.
- Ability to optimize models for real-time, low-latency inference.
- Solid programming skills in Python and experience with PyTorch/TensorFlow.
Nice to Have :
- Exposure to MLOps tools (Docker, Kubernetes, CI/CD for AI).
- Experience with evaluation frameworks for LLMs.
- Knowledge of multi-modal models (text, vision, audio).

About NeoGenCode Technologies Pvt Ltd
About
Welcome to Neogencode Technologies, an IT services and consulting firm that provides innovative solutions to help businesses achieve their goals. Our team of experienced professionals is committed to providing tailored services to meet the specific needs of each client. Our comprehensive range of services includes software development, web design and development, mobile app development, cloud computing, cybersecurity, digital marketing, and skilled resource acquisition. We specialize in helping our clients find the right skilled resources to meet their unique business needs. At Neogencode Technologies, we prioritize communication and collaboration with our clients, striving to understand their unique challenges and provide customized solutions that exceed their expectations. We value long-term partnerships with our clients and are committed to delivering exceptional service at every stage of the engagement. Whether you are a small business looking to improve your processes or a large enterprise seeking to stay ahead of the competition, Neogencode Technologies has the expertise and experience to help you succeed. Contact us today to learn more about how we can support your business growth and provide skilled resources to meet your business needs.
Candid answers by the company
IT & Engineering Talent Staffing
- Provides full-time and contract-based hiring, delivering handpicked, pre‑screened developers across tech stacks—ranging from web, mobile, AI/ML, Web3/blockchain.
- Maintains a bench o vetted candidates, offering fast delivery of interview-ready profiles—often within 24 hours.
- Offers payroll management, handling compliance, tax, attendance, and documentation for both contractors and full-time employees.
2. End-to-End Project Delivery
- Delivers full-stack development solutions: web, mobile, cloud, AI/ML, Blockchain/Web3.
- Manages entire project lifecycle—requirements gathering, design (UI/UX), development, deployment, and ongoing support .
3. Additional Offerings
- Expands into cybersecurity consulting, digital marketing, and cloud platform services (like AWS, GCP, Azure) .
- Provides strategic IT consulting to align technology solutions with business objectives
Similar jobs (10)
Support with design and build to prove out agentic AI solution flow by working with other data
scientists and engineers to build, train Large Language Model (LLM) architectures, RAG
systems, and autonomous agentic workflows
Key qualifications:
>> AI solution design & Development: Design Agentic AI solutions using RAG (Retrieval-
Augmented Generation) and orchestration frameworks like LangGraph or LangChain.
>> Model Fine-Tuning: Solid understanding and experience with Pre-train, fine-tune, and
optimize open-source like BERT, LLama, and other proprietary foundation models for domain-
specific tasks
>> Solid Stats and ML foundations and (vibe) coding skills with Python, PySpark
>> Implement validation frameworks and tracing practices (using tools like Arize) to monitor
agent behavior, guard against model drift, and ensure compliance
>> Collaborate with Engineering to deploy models securely on cloud and on-prem ecosystems
We are hiring a Generative AI Engineer to build production LLM applications.
Responsibilities
- Build RAG pipelines with LangChain or LlamaIndex
- Design prompts and evaluate model outputs
- Manage embeddings in vector databases such as Pinecone, Weaviate or FAISS
- Deploy and monitor LLM features in production
Requirements
- 1+ years building LLM-powered applications
- Hands-on with LangChain or LlamaIndex and vector databases
- Experience with the OpenAI, Anthropic or open-source model APIs
The Role
You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.
What You Will Own
• Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.
• Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.
• Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.
• Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.
Observability
An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.
• Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.
• Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.
• Build dashboards and alerting so model degradation is caught before users feel it.
• Trace failures across a distributed, always-on system using metrics, logs, and traces.
• Close the loop. Production signals feed back into evaluation and model selection.
Evaluation
We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.
• Build and own ground-truth evaluation harnesses for every model in the stack.
• Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.
• Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.
• Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.
• Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.
What We Are Looking For
• 3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.
• Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.
• Strong software engineering. You write code that ships and survives contact with real users.
• Fluency in Python and the production ML ecosystem.
• Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.
• A working command of observability and evaluation. You measure first and trust metrics over intuition.
• First-principles reasoning and metric discipline.
Nice to Have
• Research background or publications. A strong signal, not a substitute for production work.
• Audio and speech ML experience (STT, diarization, voice).
• Experience self-hosting and optimizing open models.
• Experience with LLM gateway and agent orchestration patterns.
• Experience building eval harnesses or production model-monitoring systems.
Requirements
Agentic work is must. Audio is good to have
. Self hosting models is a must
Experience with LLM gateway and agent orchestration is a must have
Hiring for AI Engineer
Exp: 5 - 10 yrs
Edu : BE/B.Tech/MCA
Work Location : Pune / Mumbai
Skill Set:
Total experience ranging from 5–10 years in software engineering/AI roles
Min 5 years strong programming experience in Python is a MUST
Min 3.5 years hands-on experience in AI with LLMs, RAG pipelines, and AI frameworks
2+ years shipping LLM systems in production
Experience with cloud platforms (AWS/Azure/GCP)
About the role
We are seeking an AI Engineer to build and implement AI systems for content production at scale. You'll work at the intersection of engineering and content designing prompt pipelines, integrating generative models, and building the tooling that turns source material into finished creative output. The ideal candidate is technically strong but also has taste: someone who understands story and craft, and can tell the difference between output that's technically correct and output that's actually good.
Responsibilities
- Build and iterate on prompt pipelines and multi-agent workflow components
- Design and integrate agentic workflows orchestrate multi-step, tool-using agents that plan, call models, and hand off between stages in production
- Deploy and serve open-source models set up inference endpoints, manage GPU compute, and optimize for latency and cost
- Write evals compare outputs against references, quantify quality, and feed results back into the pipeline
- Work on data pipelines: structured extraction from messy source text, localization, similarity/dedup
- Debug and maintain pipeline stages in production
What you bring:
- (1+/3+) years of engineering experience, or a strong portfolio of shipped projects
- Solid Python fundamentals clean, working, readable code
- Hands-on experience with LLM APIs and prompt engineering (personal projects count)
- Comfort with Git, REST APIs, and working in a Linux environment
- A feel for content and narrative you can judge whether generated output is actually good, not just valid
- Curiosity and clear communication you ask good questions and don't stay stuck silently
Preferred
- Exposure to agent/orchestration frameworks (LangGraph, LangChain, CrewAI)
- Familiarity with vector databases, embeddings, or RAG (Qdrant, pgvector)
- Hands-on work with open-source generative media models Flux, LTX, Wan, or similar
- Experience deploying open-source models for inference (vLLM, ComfyUI, Replicate/Cog, Docker + GPU)
- Experience writing evals or LLM-as-judge scoring
- Node.js and Fastapi familiarity, or experience deploying on AWS
About the role
We are seeking an AI Engineer to build and implement AI systems for content production at scale. You'll work at the intersection of engineering and content designing prompt pipelines, integrating generative models, and building the tooling that turns source material into finished creative output. The ideal candidate is technically strong but also has taste: someone who understands story and craft, and can tell the difference between output that's technically correct and output that's actually good.
Responsibilities
- Build and iterate on prompt pipelines and multi-agent workflow components
- Design and integrate agentic workflows orchestrate multi-step, tool-using agents that plan, call models, and hand off between stages in production
- Deploy and serve open-source models set up inference endpoints, manage GPU compute, and optimize for latency and cost
- Write evals compare outputs against references, quantify quality, and feed results back into the pipeline
- Work on data pipelines: structured extraction from messy source text, localization, similarity/dedup
- Debug and maintain pipeline stages in production
What you bring:
- (1+/3+) years of engineering experience, or a strong portfolio of shipped projects
- Solid Python fundamentals clean, working, readable code
- Hands-on experience with LLM APIs and prompt engineering (personal projects count)
- Comfort with Git, REST APIs, and working in a Linux environment
- A feel for content and narrative you can judge whether generated output is actually good, not just valid
- Curiosity and clear communication you ask good questions and don't stay stuck silently
Preferred
- Exposure to agent/orchestration frameworks (LangGraph, LangChain, CrewAI)
- Familiarity with vector databases, embeddings, or RAG (Qdrant, pgvector)
- Hands-on work with open-source generative media models Flux, LTX, Wan, or similar
- Experience deploying open-source models for inference (vLLM, ComfyUI, Replicate/Cog, Docker + GPU)
- Experience writing evals or LLM-as-judge scoring
- Node.js and Fastapi familiarity, or experience deploying on AWS
Strong AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) – Must have minimum 3+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
3
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
4
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
5
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
6
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
7
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
8
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
9
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
10
Mandatory (Age) - Candidate's Age should be below 28 Years
AI Engineer
LLMs, Agents & AI Services
📍 Mumbai (On-site) | Full-time | 2-4 years
About the Role:
Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.
AI is core to how we design, deliver, and scale software for our customers.
We are hiring an AI Engineer for a dedicated client engagement building a complex production AI platform, working on the AI capabilities and agentic features at the core of the product.
The mandatory requirement for this role is at least one AI feature personally shipped to production for real users, with operational ownership.
The role suits someone who thinks quickly on solutioning, can take an ambiguous problem to a working prototype in days, and has the discipline to carry it through to production with predictable economics.
You will work alongside the Senior AI Engineer and the wider pod, with ownership of parts of the AI surface area of the product.
Responsibilities:
Solutioning and POCs
Translate ambiguous customer problems into working POCs at speed.
Pick the right model, framework, and architecture, and demonstrate value early before scaling investment.
LLM Application Development
Build AI features and services using LLM APIs from OpenAI, Anthropic, Google, and self-hosted open-weight models (Llama, Qwen, Mistral).
Choose the right model per use case based on cost, latency, capability, and context-window trade-offs.
Agentic System Design
Design and implement agentic workflows using LangGraph, CrewAI, AutoGen, LlamaIndex Agents, or custom orchestration.
Cover tool use, planning, memory, and multi-step reasoning appropriate to the problem.
API and Service Development
Build production AI services and APIs using Python and FastAPI.
Handle streaming responses, async processing, structured outputs, retries, and graceful degradation when models or tools fail.
Retrieval and Tool Integration
Implement RAG pipelines with vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma), embeddings, chunking strategies, hybrid search, and reranking.
Integrate external tools, internal APIs, and document sources through tool-calling and MCP-style patterns.
Cost Analysis and Unit Economics
Model the per-request and per-user cost of every AI feature before it ships.
Track token usage, prompt caching, batching, and model-routing strategies.
Drive measurable improvements in unit economics.
Production Hardening
Add observability and tracing (LangSmith, Langfuse, OpenTelemetry), guardrails, content safety checks, prompt injection defences, and fallback behaviour.
Prompt Engineering and Evaluation
Design, test, and iterate prompts with measured outcomes.
Build evaluation harnesses for accuracy, hallucination, latency, and cost.
Run benchmarks across models and prompt variants before locking in a design.
Requirements:
AI Feature Shipped to Production (Mandatory)
Must have personally built and shipped at least one AI feature that runs in production for real users, with operational ownership.
POCs, internal demos, and one-off scripts do not qualify.
2 to 4 Years of Professional Software or AI Engineering Experience
With at least one production AI feature owned end to end.
Strong Python Proficiency and API Development with FastAPI
Comfort with type hints, async, packaging, testing, streaming responses, and authentication.
Production-grade Python, not notebook-only code.
Hands-on Depth Across the LLM and Agent Stack
Working experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or self-hosted open-weight models (vLLM, Ollama, Together, Replicate).
Working familiarity with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.
Working knowledge of RAG, embeddings, and vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma).
Solutioning Speed and POC Velocity
Demonstrated ability to move from a fuzzy problem to a working prototype in days.
Strong instinct for what to build first, what to defer, and what to throw away.
Cost Discipline for Production AI
Ability to calculate, monitor, and optimise the cost of LLM APIs, tokens, embeddings, vector store usage, and infrastructure.
Treats unit economics as a first-class concern.
AWS Familiarity
Working knowledge of EC2, S3, IAM, and at least one of Bedrock, SageMaker, or equivalent.
Comfortable in a Fast-Moving Environment
Self-directed, comfortable with ambiguity, takes ownership without being asked, and ships under shifting priorities.
Strong Written and Spoken English Communication
Able to explain trade-offs to non-AI engineers, designers, product managers, and clients in plain language.
Nice to Have
- fine-tuning or LoRA, QLoRA, PEFT exposure
- MCP server authoring
- eval framework experience (LangSmith, Promptfoo, Ragas, DeepEval)
- open-source AI contributions
- multi-modal models (vision, audio)
Job Description:
We are looking for a hands-on AI Engineer with experience in Generative AI and Agentic AI to build and deploy production-ready AI solutions.
Key Responsibilities:
- Develop and deploy GenAI and Agentic AI applications.
- Build RAG pipelines, LLM workflows, and AI agents.
- Develop solutions using Python, LangChain, LangGraph, LlamaIndex, or similar frameworks.
- Implement tool calling, context retrieval, and LLM orchestration.
- Integrate AI solutions with APIs and cloud platforms.
- Work with AWS/Azure/GCP, Docker, and CI/CD.
Required Skills:
- Strong Python programming skills.
- 3+ years of GenAI/Agentic AI experience.
- RAG and LLM orchestration.
- LangChain / LangGraph / LlamaIndex / AutoGen / CrewAI / Semantic Kernel.
- MCP and A2A knowledge.
- Cloud, APIs, Docker, and CI/CD experience.
Preferred Experience:
Hands-on experience building and deploying production-ready AI solutions.
Kody Technolab Limited is seeking an experienced AI/ML Engineer to design, develop, and deploy cutting-edge Artificial Intelligence and Machine Learning solutions. The ideal candidate will have
strong expertise in Machine Learning, Deep Learning, Generative AI, LLMs, MLOps, and cloud-based AI deployments.
Key Responsibilities
• Design, develop, and deploy Machine Learning and Deep Learning models for classification, regression, recommendation systems, NLP, Computer Vision, and Generative AI applications.
• Build and maintain end-to-end ML pipelines including data preprocessing, feature engineering, model training, validation, evaluation, and deployment.
• Develop AI solutions using PyTorch, TensorFlow, Scikit-learn, Hugging Face, and related frameworks.
• Work with Large Language Models (LLMs) and foundation models such as GPT, BERT, Llama, Claude, and Stable Diffusion.
• Collaborate with product, engineering, and business teams to translate requirements into scalable AI solutions.
• Optimize model performance, scalability, and reliability for production environments.
• Implement MLOps best practices using tools such as MLflow, Docker, Kubernetes, and Kubeflow.
• Stay updated with emerging trends and research in AI, ML, Deep Learning, and Generative AI.
Required Qualifications
• Bachelor’s or Master’s degree in Computer Science, Data Science, Artificial Intelligence, Mathematics, or a related field.
• 7+ years of hands-on experience in AI/ML product development.
• Strong proficiency in Python and ML frameworks including Scikit-learn, TensorFlow, PyTorch, and Hugging Face.
• Experience with Generative AI, LLMs, GANs, VAEs, diffusion models, and prompt engineering.
• Strong understanding of the ML lifecycle including model training, tuning, deployment, monitoring, and optimization.
• Experience with MLOps tools such as MLflow, Docker, Kubeflow, and CI/CD pipelines.
• Experience with AWS, Azure, or GCP cloud platforms.
• Strong problem-solving and analytical skills.
Preferred Skills
• Fine-tuning and deployment of Large Language Models.
• Experience with RAG (Retrieval Augmented Generation) architectures.
• Contributions to open-source AI projects or research publications.
• Knowledge of model interpretability, data annotation, and feature engineering.
• C++ experience for high-performance AI applications.
Why Join Kody Technolab Limited?
Opportunity to work on innovative AI products, Generative AI solutions, robotics integrations,
and enterprise-scale applications while collaborating with a highly skilled technology team.
Visit the Website to know more about us.
Company Website - Kody Technolab | Deep Tech Company in Robotics & AI Solution
Kody Robots | Robotics Company in India for Autonomous Robots






