Data Scientist at Fintech lead, · Remote only · 3 - 7 years · ₹5L - ₹15L / yr · Remote only · Posted 20 Dec 2023

Job description
Who we are looking for
· A Natural Language Processing (NLP) expert with strong computer science fundamentals and experience in working with deep learning frameworks. You will be working at the cutting edge of NLP and Machine Learning.
Roles and Responsibilities
· Work as part of a distributed team to research, build and deploy Machine Learning models for NLP.
· Mentor and coach other team members
· Evaluate the performance of NLP models and ideate on how they can be improved
· Support internal and external NLP-facing APIs
· Keep up to date on current research around NLP, Machine Learning and Deep Learning
Mandatory Requirements
· Any graduation with at least 2 years of demonstrated experience as a Data Scientist.
Behavioural Skills
· Strong analytical and problem-solving capabilities.
· Proven ability to multi-task and deliver results within tight time frames
· Must have strong verbal and written communication skills
· Strong listening skills and eagerness to learn
· Strong attention to detail and the ability to work efficiently in a team as well as individually
Technical Skills
Hands-on experience with
· NLP
· Deep Learning
· Machine Learning
· Python
· Bert
Preferred Requirements
· Experience in Computer Vision is preferred

Similar jobs (4)
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Description:
We are seeking a highly skilled Machine Learning Engineer to join our team. The ideal candidate will have a strong background in Natural Language Processing (NLP), Large Language Models (LLMs), and Python programming.
You will work closely with data scientists, product managers, and data engineers to design, develop, and deploy high-performance AI/ML models and integrate generative AI solutions into existing workflows.
Your responsibilities will include:
- Collaborating with cross-functional teams to design and deliver high-performance AI models, including NLP, computer vision, semantics engines, linguistic analysis, risk management, and time-series prediction models. Integrating generative AI solutions into existing workflow systems.
- Developing and maintaining the ML Operations CI/CD pipeline for seamless deployment and monitoring. Training, tuning, and optimizing AI models and algorithms for enhanced performance.
- Implementing complex real-time data and AI/ML applications to capture knowledge and automate decision-making processes.
- Creating ML/AI models for business teams and establishing metrics to track their accuracy and performance. Overseeing the full lifecycle of algorithm development, from ideation to deployment and monitoring. Evaluating and ranking ML algorithms based on their potential success in solving specific problems.
- Serving as an internal resource for AI/ML needs, providing guidance and insights to stakeholders during strategic discussions.
Required Experience and Skills:
Machine Learning:
- Proficient in generative AI techniques, prompt engineering, and Retrieval-Augmented Generation (RAG) (3+ years).
- Experience with Large Language Models (LLMs) such as OpenAI, Gemini, LLAMA, and other state-of-the-art models (3+ years).
- Expertise in using ML/AI libraries such as Pandas, NumPy, PyTorch, TensorFlow, Keras, BERT, LayoutLM, and traditional ML algorithms (5+ years).
- Experience with distributed ML/AI training libraries/models: Koalas, Horovod, DDP.
Python Programming and Software Engineering:
- Expertise in Pythonic clean coding practices, including the use of decorators, generators, and descriptors (5+ years).
- Strong understanding of software design principles such as DRY, OAOO, YAGNI, KIS, EAFP/LBYL, and defensive programming (2+ years).
- Proficient in software design concepts focusing on cohesion and coupling (2+ years). Knowledge of SOLID principles (2+ years).
Education and Experience:
- Minimum Bachelor's degree or foreign equivalent in Computer Science, Electrical Engineering, or a closely related field.
- At least 5 years of experience as a software engineer and 5 years of ML-related programming.
Key Responsibilities
• Design, build, and deploy machine learning and AI models that power Transient.AI's core products (research
automation, document intelligence, investor matching, and workflow orchestration).
• Work on applied NLP/LLM systems, including retrieval-augmented generation, structured extraction from
unstructured financial documents, and model evaluation pipelines.
• Partner closely with product and founding engineers to translate capital markets workflows into scalable AI
systems.
• Own model performance, reliability, and cost — from experimentation through production deployment.
• Build and maintain data pipelines, feature stores, and evaluation frameworks to support rapid iteration.
• Ensure systems meet the compliance, auditability, and security standards required in regulated financial
environments.
What We're Looking For
• 5+ years of experience building and deploying machine learning or AI systems in production.• Strong hands-on experience with Python and modern ML/AI frameworks (PyTorch, TensorFlow, Hugging Face,
LangChain, or equivalent).
• Experience with LLMs — fine-tuning, prompt engineering, RAG architectures, or agentic systems — is highly
valued.
• Solid grounding in data structures, distributed systems, and MLOps practices (model serving, monitoring,
versioning).
• Prior experience at a strong product company, high-growth startup, or a top-tier engineering background
• Comfort operating in an early-stage, high-ownership environment with limited process and high ambiguity.
• Exposure to fintech, capital markets, or other regulated industries is a plus, though not mandatory
Sr Engineer – Artificial Intelligence
Job Summary
As an AI Engineer at Emerson, you will be responsible for analysing complex data sets to
identify trends, develop predictive models, and provide actionable insights. You will work closely
with cross-functional teams to understand business needs and deliver data-driven solutions that
enhance decision-making and drive business growth.
In This Role, Your Responsibilities Will Be:
Analyze large, complex data sets using statistical methods and machine learning
techniques to extract meaningful insights.
Develop and implement predictive models and algorithms to solve business problems
and improve processes.
Create visualizations and dashboards to effectively communicate findings and insights to
stakeholders.
Work with data engineers, product managers, and other team members to understand
business requirements and deliver solutions.
Clean and preprocess data to ensure accuracy and completeness for analysis.
Prepare and present reports on data analysis, model performance, and key metrics to
stakeholders and management.
Participate in regular Scrum events such as Sprint Planning, Sprint Review, and Sprint
Retrospective
Stay updated with the latest industry trends and advancements in data science and
machine learning techniques.
For This Role, You Will Need:
Bachelor’s degree in computer science, Data Science, Statistics, or a related field or a
master's degree or higher is preferred.
Total 5-7 years of industry experience
More than 3 years of experience in a data science or analytics role, with a strong track
record of building and deploying models.
Proficiency in programming languages such as Python or R, and experience with data
manipulation libraries (e.g., pandas, NumPy).
Excellent understanding of Agentic Frameworks like Microsoft Agent Framework.
Experience with NLP, NLG, and Large Language Models Open Source as well as Cloud
based models.
Experience with SQL and NoSQL databases such as MongoDB, Cassandra, Vector
databases
Experience with Dockers, Asynchronous Data Orchestrators, environments etc.
Strong analytical and problem-solving skills, with the ability to work with complex data
sets and extract actionable insights.
Excellent verbal and written communication skills, with the ability to present complex
technical information to non-technical stakeholders.
Preferred Qualifications that Set You Apart:
Prior experience in engineering domain would be nice to have
Prior experience in working with teams in Scaled Agile Framework (SAFe) is nice to
have
Possession of relevant certification/s in data science from reputed universities
specializing in AI.
Familiarity with cloud platforms, Microsoft Azure is preferred
Ability to work in a fast-paced environment and manage multiple projects simultaneously.
Strong analytical and troubleshooting skills, with the ability to resolve issues related to
model performance and infrastructure.
The Role
You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.
What You Will Own
• Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.
• Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.
• Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.
• Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.
Observability
An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.
• Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.
• Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.
• Build dashboards and alerting so model degradation is caught before users feel it.
• Trace failures across a distributed, always-on system using metrics, logs, and traces.
• Close the loop. Production signals feed back into evaluation and model selection.
Evaluation
We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.
• Build and own ground-truth evaluation harnesses for every model in the stack.
• Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.
• Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.
• Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.
• Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.
What We Are Looking For
• 3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.
• Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.
• Strong software engineering. You write code that ships and survives contact with real users.
• Fluency in Python and the production ML ecosystem.
• Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.
• A working command of observability and evaluation. You measure first and trust metrics over intuition.
• First-principles reasoning and metric discipline.
Nice to Have
• Research background or publications. A strong signal, not a substitute for production work.
• Audio and speech ML experience (STT, diarization, voice).
• Experience self-hosting and optimizing open models.
• Experience with LLM gateway and agent orchestration patterns.
• Experience building eval harnesses or production model-monitoring systems.
Requirements
Agentic work is must. Audio is good to have
. Self hosting models is a must
Experience with LLM gateway and agent orchestration is a must have






