Machine Learning Lead at Process Nine Technologies · Gurugram · 4 - 8 years · Raised funding · Posted 1 Sep 2026
ML Leads JD
Key Responsibilities
- Model Training & Fine-Tuning: Build, fine-tune, and optimize state-of-the-art NLP, LLM, Speech, and Vision models for scheduled Indian languages, utilizing parameter-efficient methods (LoRA, QLoRA, PEFT).
- Indic Tokenization & Linguistics: Architect custom tokenizers and text-normalization pipelines to address the "fertility problem" in Devanagari, Dravidian, and other regional scripts, ensuring low-latency and cost-effective model inference.
- Multimodal System Design: Develop robust OCR engines capable of parsing complex script geometries (conjoint consonants, Shirorekha, vowel modifiers) and integrate them into document intelligence pipelines.
- Speech Engineering: Deploy and scale robust STT (Speech-to-Text) and TTS (Text-to-Speech) pipelines capable of handling heavy code-mixing (e.g., Hinglish, Tanglish), regional accents, and localized dialects.
- Vernacular Guardrails & Evaluation: Establish culturally contextual benchmark datasets and implement safety guardrails.
- Production Deployment (MLOps): Package and serve models using high-throughput frameworks (vLLM, Triton, ONNX) optimized for GPU environments, minimizing computational overhead for massive cross-lingual workloads.
- Vernacular Fraud & Anomaly Detection: Architect risk-scoring systems and anomaly detection models capable of identifying fraud patterns in native scripts and code-mixed formats.
Essential Qualifications & Technical Skills
- Education: Bachelor’s or Master's degree in Computer Science, Mathematics, Statistics, or a closely related quantitative field.
- Experience: 4+ years of professional experience building and deploying machine learning models in production environments, with a proven track record in Indian Language NLP, Speech, or Anomaly Detection.
- Programming: Expert-level proficiency in Python and standard ML frameworks (PyTorch, TensorFlow).
- Indic AI Stack: Direct, hands-on experience with specialized Indic frameworks and datasets (e.g., AI4Bharat's IndicTrans2/IndicWhisper, Bhashini API, Kathbath, Sarvam-105B, or Aksharantar).
- Fraud Stack: Proficiency in tabular/graph-based ML toolkits (XGBoost, LightGBM, PyTorch Geometric) and handling highly imbalanced target variables (SMOTE, class weights).
- NLP & LLMs: Deep understanding of Transformer architectures, sequence-to-sequence modeling, cross-lingual embeddings, vector databases (Milvus, Pinecone, Qdrant), and quantization tools (bitsandbytes, GPTQ).
- Speech & Vision Processing: Experience processing raw audio signals (grapheme-to-phoneme conversion, spectrogram analysis) or document structures using OCR networks (CRAFT, DBNet, LayoutLM).
- Handling Code-Mixing: Proven ability to build models that gracefully parse text or speech containing heavy code-switching (mixed Latin/regional scripts, multi-language grammar).

About Process Nine Technologies
About
Company video


Connect with the team
Similar jobs (10)
Strong AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) – Must have minimum 3+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
3
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
4
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
5
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
6
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
7
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
8
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
9
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
10
Mandatory (Age) - Candidate's Age should be below 28 Years
Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2
Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3
Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.
4
Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5
Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6
Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.
7
Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8
Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9
Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10
Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11
Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12
Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
The Role
You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.
What You Will Own
• Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.
• Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.
• Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.
• Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.
Observability
An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.
• Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.
• Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.
• Build dashboards and alerting so model degradation is caught before users feel it.
• Trace failures across a distributed, always-on system using metrics, logs, and traces.
• Close the loop. Production signals feed back into evaluation and model selection.
Evaluation
We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.
• Build and own ground-truth evaluation harnesses for every model in the stack.
• Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.
• Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.
• Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.
• Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.
What We Are Looking For
• 3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.
• Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.
• Strong software engineering. You write code that ships and survives contact with real users.
• Fluency in Python and the production ML ecosystem.
• Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.
• A working command of observability and evaluation. You measure first and trust metrics over intuition.
• First-principles reasoning and metric discipline.
Nice to Have
• Research background or publications. A strong signal, not a substitute for production work.
• Audio and speech ML experience (STT, diarization, voice).
• Experience self-hosting and optimizing open models.
• Experience with LLM gateway and agent orchestration patterns.
• Experience building eval harnesses or production model-monitoring systems.
Requirements
Agentic work is must. Audio is good to have
. Self hosting models is a must
Experience with LLM gateway and agent orchestration is a must have
Kody Technolab Limited is seeking an experienced AI/ML Engineer to design, develop, and deploy cutting-edge Artificial Intelligence and Machine Learning solutions. The ideal candidate will have
strong expertise in Machine Learning, Deep Learning, Generative AI, LLMs, MLOps, and cloud-based AI deployments.
Key Responsibilities
• Design, develop, and deploy Machine Learning and Deep Learning models for classification, regression, recommendation systems, NLP, Computer Vision, and Generative AI applications.
• Build and maintain end-to-end ML pipelines including data preprocessing, feature engineering, model training, validation, evaluation, and deployment.
• Develop AI solutions using PyTorch, TensorFlow, Scikit-learn, Hugging Face, and related frameworks.
• Work with Large Language Models (LLMs) and foundation models such as GPT, BERT, Llama, Claude, and Stable Diffusion.
• Collaborate with product, engineering, and business teams to translate requirements into scalable AI solutions.
• Optimize model performance, scalability, and reliability for production environments.
• Implement MLOps best practices using tools such as MLflow, Docker, Kubernetes, and Kubeflow.
• Stay updated with emerging trends and research in AI, ML, Deep Learning, and Generative AI.
Required Qualifications
• Bachelor’s or Master’s degree in Computer Science, Data Science, Artificial Intelligence, Mathematics, or a related field.
• 7+ years of hands-on experience in AI/ML product development.
• Strong proficiency in Python and ML frameworks including Scikit-learn, TensorFlow, PyTorch, and Hugging Face.
• Experience with Generative AI, LLMs, GANs, VAEs, diffusion models, and prompt engineering.
• Strong understanding of the ML lifecycle including model training, tuning, deployment, monitoring, and optimization.
• Experience with MLOps tools such as MLflow, Docker, Kubeflow, and CI/CD pipelines.
• Experience with AWS, Azure, or GCP cloud platforms.
• Strong problem-solving and analytical skills.
Preferred Skills
• Fine-tuning and deployment of Large Language Models.
• Experience with RAG (Retrieval Augmented Generation) architectures.
• Contributions to open-source AI projects or research publications.
• Knowledge of model interpretability, data annotation, and feature engineering.
• C++ experience for high-performance AI applications.
Why Join Kody Technolab Limited?
Opportunity to work on innovative AI products, Generative AI solutions, robotics integrations,
and enterprise-scale applications while collaborating with a highly skilled technology team.
Visit the Website to know more about us.
Company Website - Kody Technolab | Deep Tech Company in Robotics & AI Solution
Kody Robots | Robotics Company in India for Autonomous Robots
Must of Skills/Experience
• System Design
• Python
• TensorFlow
• Google ADK or Lang Graph
• Lang Chain , Lang Graph
• Spark
• Agentic AI Design
• ML Ops
• MCP (client and server)
• FastAPI
• Doc Factory
• RAG
• Golang
• LLMs – Gemini, Open AI
• NLP
• Dev Assistant - AI based code - generation
(Qwen or Claude or Copilot)
• CI/CD
• Good in oral and written communication,
collaboration and be a team player
Good to have skills
• DevOps with K8
• Scripting
• Java
• REST API
• UV
• ReACT
• DocFactory
• Unix
Location: Hyderabad, India (home base), deployed at client sites in India. Occasional Middle East exposure possible.
About the Role
You will work as a senior AI engineer who embeds inside a customer's business. Your job is to learn how the business makes money, find the highest value problem, and build a working system that solves it.
Four behaviors define this role:
- Go where the work happens. You work onsite with the customer, in the room where decisions are made.
- Show working software early. You build a prototype in days, not a document in weeks.
- One person owns the outcome. You are the single point of accountability for the result.
- Stay after go-live. You keep running and improving the system after launch.
You are the single point of accountability. You are not a solo builder. A full KnackLabs engineering team in Hyderabad builds and runs the production systems behind you.
This role involves extended onsite deployments at client locations in other cities, sometimes up to six months at a stretch. Please apply only if you are ready for this way of working.
What you'll own
- Discovery - Learn how the customer makes money. Find the highest value problem to solve first.
- The prototype - Build a working prototype fast, using real or sample data, to prove the idea.
- The roadmap - Decide what to build, in what order, and set clear success measures tied to business outcomes.
- The build - Design and ship the production system with the Hyderabad engineering team. This includes data integration, agents, retrieval, and evaluations.
- The client relationship - Be the trusted technical contact for the customer, from engineers to senior leaders.
- Go live and after - Deploy the system, watch how it performs, fix problems, and improve it over time.
- Feedback to the product - Share what you learn in the field so the vendor's product and our internal tools get better.
What we are looking for
- Around 4 or more years of software engineering experience, including customer-facing or client delivery work.
- Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
- A full-stack development experience with strength in backend technologies.
- Production experience with large language models, including prompt engineering and agent development.
- You build with AI coding tools like Claude Code or Codex as your default way of working, and you have shipped real apps or agents this way.
- Experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
- Experience building and deploying AI systems.
- Experience integrating with APIs and enterprise systems.
- Experience with at least one cloud platform (AWS, Azure, or GCP).
- Clear communication. You can explain a technical choice to an engineer and to a business leader.
- High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a plan.
- Willingness to work onsite at client locations in India for extended periods, and to travel as the work needs.
Nice to have
- Experience with on-premises or private cloud (VPC) deployments.
- Experience with observability and tracing tools such as LangSmith or Braintrust.
- Experience with data engineering and pipelines.
- A history of side projects, open source contributions, or products you shipped end-to-end.
- Experience in embedded or forward-deployed roles before.
- Experience working at a consulting or professional services firm in a client-facing delivery role.
Stack and tools
- Languages: Python and TypeScript.
- Models: Claude and other frontier or open-source models, chosen to fit the customer.
- AI patterns: RAG, agents, prompt engineering, and evaluations.
- Vector and retrieval: vector databases and retrieval pipelines.
- Cloud: AWS, Azure, or GCP, on public or private cloud.
- Integration: REST APIs and enterprise system connectors.
This is a remote position.
About Leegality:
Leegality works with large Indian businesses to digitally transform critical compliance processes in a fast, easy and secure way.
We have multiple products across 2 categories:
Document Infrastructure:
Products that help businesses build paperless processes at scale:
- Document Execution Workflow: A unified platform for businesses to digitally execute (eSign, eStamp, Template Pre-fill, Document Fraud Prevention etc.) agreements, forms and other documents in a compliant way. Currently in use by 2000+ Indian businesses from giants like HDFC and SBI Cards to high-growth disruptors like goDigit and Cars24.
- Contract Management: An AI-powered platform for businesses to quickly review, negotiate and take action on contract
- Signstation: A simple platform for businesses to digitally sign simple documents like invoices, policies and letters in a cost effective manner
Consent Infrastructure:
- Consentin: An end-to-end DPDP and Privacy compliance platform for Indian businesses
- Consentin Lens: A data discovery platform for businesses to identify the personal data they collect and store.
If you’re interested in building mission critical software that operates at population scale (75 million + Indians have signed at least one document through Leegality) then join Leegality.
Curious about our impact? Explore our customer success stories: leegality.com/case-studies
Our Culture
At Leegality, trust, ownership, transparency, and having fun while doing meaningful work are core to how we operate — not just values on paper. Our team rated us an incredible 97 eNPS for FY 2023–24 — the highest among 175+ startups surveyed.
We focus deeply on helping our people grow and stay motivated. Some of the perks you’ll enjoy:
- Flexible working hours
- Hybrid work setup
- Bi-annual performance appraisals
- A culture that rewards initiative, curiosity, and impact
If you're looking for a place where you can make a real difference while working with smart, driven, and genuinely nice people, welcome to Leegality.
Location: Hybrid
Job Brief:
- As a Machine Learning Engineer specializing in Computer Vision (CV) and Natural Language Processing (NLP), you will develop solutions to interesting technical problems, exploring exciting growth opportunities and having a real impact on our product, particularly focusing on document and content intelligence.
- To ensure success, you should demonstrate solid data science knowledge and experience in a related ML, CV, or NLP role. A first-class engineer will be someone whose expertise enhances our systems for document intelligence and content processing
Responsibilities:
- Designing machine learning systems, self-running artificial intelligence (AI) software, and specialized models for Computer Vision and Natural Language Processing applications.
- Transforming data science prototypes and applying appropriate deep learning algorithms and tools to text and image/document data.
- Solving complex CV and NLP problems with multi-layered data types, such as image/document classification, information extraction, semantic search, and object detection.
- Optimizing existing machine learning models, with a focus on high-performance model deployment for CV and NLP tasks.
- Developing ML algorithms (including large language models/LLMs and computer vision models) to analyze huge volumes of historical text, image, and document data to make predictions and automate workflows.
- Running tests, performing statistical analysis, and interpreting test results for CV/NLP model performance.
- Documenting machine learning processes, model architectures, and data pipelines.
- Keeping abreast of developments in machine learning, Computer Vision, and Natural Language Processing.
Requirements:
- 3+ years of relevant experience in Machine Learning Engineering, with a strong focus on Computer Vision and/or Natural Language Processing.
- Advanced proficiency with Python.
- Extensive knowledge of ML frameworks, libraries (e.g., PyTorch, Transformers), data structures, data modeling, and software architecture.
- Experience with building and maintaining scalable RESTful APIs (e.g., FastAPI).
- In-depth knowledge of mathematics, statistics, deep learning (CNNs, RNNs, Transformers), and algorithms.
- Superb analytical and problem-solving abilities, especially for unstructured data challenges.
- Great communication and collaboration skills.
- Excellent time management and organizational abilities.
- Experience with cloud platforms (e.g., AWS) for model deployment and MLOps.
Recruitment Process:
- Our hiring process combines AI-powered evaluations with structured interviews to ensure a fair and seamless experience.
- You will be contacted via email with the next steps upon being shortlisted.
- The process may include Assessments, AI-enabled interviews, and In-Person Interviews with our team.
- Final selection and CTC will be based on your overall performance and experience.
Apply directly through our career page: https://careers.leegality.com/jobs/Careers
For more information about us please visit our:
Our Company and Culture: https://bit.ly/3Iqm5SB
Our Website: www.leegality.com/
Our LinkedIn Page: www.linkedin.com/company/leegality/
Leegality's Privacy Notice: https://www.leegality.com/employee-privacy-notice
KEY RESPONSIBILITIES:
•Build agents with persistent context & memory
•Design self-learning feedback loops
•Implement RAG pipelines for domain knowledge
•Manage conversation state & orchestration
•Integrate with LLM APIs (OpenAI, Claude, open-source)
Iterate fast — ship daily, measure weekly
MUST-HAVE SKILLS
•Python / TypeScript proficiency
•LangChain, CrewAI, AutoGen or custom frameworks
•Experience with vector DBs (Pinecone, Weaviate, Qdrant)
•Prompt engineering & evaluation pipelines
•Understanding of agent architectures (ReAct, tool-use)
Git, CI/CD, containerization basics
🚨 Hiring – Data Scientist | Python + Agentic AI
💼 Experience: 5+ Years
Must Have:
• Strong Data Science experience
• Python
• Agentic AI / AI Agents
• Generative AI / LLMs
• RAG / Vector Databases
• LangChain / LangGraph or similar Agent Frameworks
• Machine Learning & NLP
Experience - 4 to 6 year
Location – Ahmedabad/Pune/Indore
- Additional Job Description
Additional Job Description
Required Skills and Experience:
- Strong proficiency in Python and experience with ML/AI libraries (scikit-learn, TensorFlow, PyTorch, Hugging Face ecosystem).
- Hands-on experience with LLMs, RAG, vector databases, and retrieval pipelines.
- Practical experience deploying agentic workflows and building multi-step, tool-enabled agents.
- Experience using Garak (or similar LLM red-teaming/vulnerability scanners) to identify model weaknesses and harden deployments.
- Demonstrated experience implementing content filtering / moderation systems.
- Solid skills working with structured and unstructured data and advanced feature engineering.
- Familiarity with cloud GenAI platforms and services (Azure AI Services preferred; AWS/GCP acceptable).
- Experience building APIs/microservices; containerization (Docker), orchestration (Kubernetes).
- Strong understanding of model evaluation, performance profiling, inference cost optimization, and observability.
- Good knowledge of security, data governance, and privacy best practices for AI systems.






