Audio Machine Learning Engineer (Voice AI) at AI Product Company · Hyderabad · 3 - 8 years · ₹30L - ₹50L / yr · Posted 7 Aug 2026

Job Description
Qualified applicants with experience in the following role or comparable job titles are encouraged to apply:
- Audio AI Engineer
- Speech AI Engineer
- Audio Software Engineer
- Speech Processing Engineer
- Audio Machine Learning Engineer
- Machine Learning Engineer – Speech
- AI Engineer – Audio
- Speech Processing Engineer
- Research Engineer – Speech AI
- Research Scientist – Speech
- Speech Recognition Engineer
- ASR Engineer
- Automatic Speech Recognition Engineer
- Voice AI Engineer
- Speech Scientist
- Audio Research Engineer
- Audio Systems Engineer
- Audio Software Engineer
- Audio R&D Engineer
- Speech Research Engineer
- Audio DSP Engineer
- Voice Processing Engineer
- Acoustic AI Engineer
- Research Engineer – Speech AI
- Research Scientist – Speech
- AI Engineer – Speech
- AI Engineer – Audio
Requirement:
a) Working on Design, Development, testing and deployment of different speech enhancement products using frameworks like PyTorch and TensorFlow.
b) Working on enhancing existing speech enhancement products w.r.t
performance, low latency and less computation.
c) Work on improvement of adapting the model to multi channels.
d) Should be self-motivated to learn and explore new areas and able to work independently and contribute
Experience
a) Good understanding of signal processing and machine learning.
b) Hands-on experience with Deep Learning techniques like CNNs, RNNs, LSTMs, Transformers, etc., for speech processing is essential.
c) Good understanding of: Linear algebra, Optimization techniques, Statistics and pattern recognition
d) Minimum 3 years of work experience in the Audio/Speech domain
e) Good programming skills in C/C++, Python, AI/ML frameworks.
Qualification
Bachelor's / Master’s Degree from AI/ML, CSE, ECE, or PhD

Similar jobs (10)
Strong AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) – Must have minimum 3+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
3
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
4
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
5
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
6
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
7
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
8
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
9
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
10
Mandatory (Age) - Candidate's Age should be below 30 Years
11
Preferred (Experience 1) – Experience with MLFlow, Kubeflow, Airflow, Prefect, Feature Stores, Model Registry, or MLOps/LLMOps frameworks.
12
Preferred (Experience 2) – Experience working with Vector Databases, Spark, PySpark, distributed ML pipelines, large-scale data processing, or real-time ML systems..
13
Preferred (Experience 3) – Familiarity with Docker, Kubernetes, Azure, AWS, GCP, cloud-native AI deployments, and scalable ML architecture.
14
Preferred (Company) – Candidates from AI-first startups, Fintech, Banking, Lending, Fraud Analytics, Risk Analytics, Product Companies, SaaS organizations, or data-driven technology companies
15
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.
Strong AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) – Must have minimum 5+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
3
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
4
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
5
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
6
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
7
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
8
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
9
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
10
Mandatory (Age) - Candidate's Age should be below 30 Years
11
Preferred (Experience 1) – Experience with MLFlow, Kubeflow, Airflow, Prefect, Feature Stores, Model Registry, or MLOps/LLMOps frameworks.
12
Preferred (Experience 2) – Experience working with Vector Databases, Spark, PySpark, distributed ML pipelines, large-scale data processing, or real-time ML systems..
13
Preferred (Experience 3) – Familiarity with Docker, Kubernetes, Azure, AWS, GCP, cloud-native AI deployments, and scalable ML architecture.
14
Preferred (Company) – Candidates from AI-first startups, Fintech, Banking, Lending, Fraud Analytics, Risk Analytics, Product Companies, SaaS organizations, or data-driven technology companies
15
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.
Strong AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) – Must have minimum 3+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
3
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
4
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
5
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
6
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
7
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
8
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
9
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
10
Mandatory (Age) - Candidate's Age should be below 28 Years
AI based systems design and development, entire pipeline from image/ video ingest, metadata ingest, processing, encoding, transmitting.
Implementation and testing of advanced computer vision algorithms.
Dataset search, preparation, annotation, training, testing, fine tuning of vision CNN models. Multimodal AI, LLMs, hardware deployment, explainability.
Detailed analysis of results. Documentation, version control, client support, upgrades.
About Us
Invorto is our Voice AI product, bringing intelligent voice agents to real-world customer and operational use cases. Our voice pipeline is built in Python, running an STT → LLM → TTS architecture on top of the Pipecat framework.
This is a chance to work on hard problems in voice AI — latency, accuracy, naturalness, and reliability — building zero-to-one, owning your area end-to-end, and shipping to production at scale.
Note: This is a customer-facing role, and strong communication skills are essential.
About the Role
We're looking for a Voice AI Research Engineer to join the Invorto team and help build and continuously improve the voice AI systems that power our intelligent voice agents. This role is focused on the specialized craft of voice AI — designing evaluation and automation frameworks that ensure our STT, LLM, and TTS pipeline performs reliably in real-world, production conditions.
What You'll Do
- Design and build automated testing and quality frameworks for our STT → LLM → TTS voice pipeline, built on Pipecat
- Evaluate and benchmark STT, LLM, and TTS/ASR components on accuracy, latency, naturalness, and robustness across accents, languages, and real-world audio conditions
- Work hands-on with STT, TTS, and ASR models — fine-tuning, evaluating, and improving them for production use cases
- Identify failure modes and edge cases across the pipeline (background noise, accents, interruptions, turn-taking, latency, pipeline-stage handoffs) and build systems to catch them before production
- Collaborate closely with engineering to integrate quality checks and automation into the voice agent development lifecycle within the Pipecat-based architecture
- Research and stay current with advances in voice AI, and bring in new techniques, models, and tools to improve pipeline performance
- Work directly with customers to understand real-world voice use cases and translate them into evaluation criteria and quality benchmarks
- Partner with product and engineering to define what "production-grade quality" means for voice agents and drive the team toward it
What We're Looking For
- 4–6 years of experience, with a specialization in voice AI systems and automated quality evaluation
- Hands-on experience with STT (Speech-to-Text), TTS (Text-to-Speech), and ASR (Automatic Speech Recognition) models
- Experience designing and building automated testing/evaluation frameworks for voice or speech systems
- Strong understanding of what drives voice AI quality — accuracy, latency, naturalness, and robustness to real-world variability
- Strong programming skills in Python; familiarity with Pipecat or similar voice pipeline/orchestration frameworks is a plus
- Understanding of STT → LLM → TTS pipeline architectures and the trade-offs involved at each stage
- Research mindset — comfortable exploring new models, techniques, and tools and translating them into practical improvements
- Excellent communication skills — this is a customer-facing role, and you'll regularly engage directly with customers to understand needs and validate quality expectations
Strong AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) – Must have minimum 3+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
3
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
4
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
5
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
6
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
7
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
8
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
9
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
10
Mandatory (Age) - Candidate's Age should be below 28 Years
11
Preferred (Experience 1) – Experience with MLFlow, Kubeflow, Airflow, Prefect, Feature Stores, Model Registry, or MLOps/LLMOps frameworks.
12
Preferred (Experience 2) – Experience working with Vector Databases, Spark, PySpark, distributed ML pipelines, large-scale data processing, or real-time ML systems..
13
Preferred (Experience 3) – Familiarity with Docker, Kubernetes, Azure, AWS, GCP, cloud-native AI deployments, and scalable ML architecture.
14
Preferred (Company) – Candidates from AI-first startups, Fintech, Banking, Lending, Fraud Analytics, Risk Analytics, Product Companies, SaaS organizations, or data-driven technology companies
15
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.
Must of Skills/Experience
• System Design
• Python
• TensorFlow
• Google ADK or Lang Graph
• Lang Chain , Lang Graph
• Spark
• Agentic AI Design
• ML Ops
• MCP (client and server)
• FastAPI
• Doc Factory
• RAG
• Golang
• LLMs – Gemini, Open AI
• NLP
• Dev Assistant - AI based code - generation
(Qwen or Claude or Copilot)
• CI/CD
• Good in oral and written communication,
collaboration and be a team player
Good to have skills
• DevOps with K8
• Scripting
• Java
• REST API
• UV
• ReACT
• DocFactory
• Unix
AuxoAI is hiring a Senior Applied AI Engineer to design and deploy production-grade computer vision systems that operate reliably in real-world environments.
This role focuses on building end-to-end visual intelligence systems, combining deep learning, classical computer vision techniques, and multimodal models. It is not limited to model training and requires strong ownership of system design, deployment, and real-world performance.
You will work on systems that perform perception, understanding, and reasoning over visual data, and integrate these capabilities into larger AI platforms and agent-based workflows.
You will also work on problems where existing approaches may not be sufficient, and will be expected to combine deep learning, geometric methods, and multimodal reasoning to build robust, production-grade systems.
Location – Mumbai / Bangalore / Hyderabad / Gurgaon (Hybrid – 3 days per week in office)
Responsibilities:
- Design and deploy computer vision systems for tasks such as:
- Object detection, segmentation, and tracking
- Scene understanding and structured perception
- Video understanding and temporal reasoning
- Build and optimize models using architectures such as:
- CNNs (ResNet, EfficientNet)
- Vision Transformers (ViT, Swin, DeiT)
- Detection/segmentation models (YOLO, DETR, Mask R-CNN)
- Develop multimodal systems combining vision and language:
- CLIP-style models
- Vision-language models (VLMs)
- Visual grounding and captioning systems
- Implement algorithms for:
- Multi-object tracking (SORT, DeepSORT, ByteTrack)
- Feature matching and representation learning
- Temporal modeling (RNNs, Transformers for video)
- Apply geometric and classical computer vision methods where relevant:
- Camera calibration
- Epipolar geometry
- Pose estimation
- 3D reconstruction or depth estimation
- Optimize systems for:
- Low-latency, real-time inference
- Throughput and scalability
- Edge and distributed deployment
- Design and build data pipelines for:
- Annotation workflows
- Dataset curation
- Synthetic data generation
- Integrate vision systems into:
- Multimodal AI pipelines
- Agent-based systems
- Decision-making workflows
Requirements:
- 5+ years of experience building computer vision systems in production environments
- Strong experience with deep learning frameworks (PyTorch / TensorFlow)
- Hands-on experience with:
- Detection, segmentation, or tracking systems
- Model training, fine-tuning, and evaluation
- Strong understanding of:
- Representation learning
- Loss functions (contrastive loss, focal loss, etc.)
- Evaluation metrics (mAP, IoU, precision/recall)
- Experience building and deploying end-to-end vision systems, not just training models
Candidates whose primary experience is limited to academic projects or model experimentation without real-world deployment may not be a fit for this role.
Nice to Have:
- Experience with multimodal systems (vision + language)
- Familiarity with models such as:
- CLIP, BLIP, Flamingo, or similar
- Experience with 3D vision:
- NeRFs
- SLAM
- Point clouds
- Experience with video understanding:
- Action recognition
- Event detection
- Experience building data engines:
- Active learning
- Hard negative mining
- Experience working with large-scale datasets and distributed training pipelines
Procedure is hiring for WorkHero.
WorkHero is building the AI-powered back office for the skilled trades, starting with the $50B+ HVAC industry. Small contractors are great at their trade but lose 20+ hours a week to invoicing, permits, scheduling, and paperwork. WorkHero combines expert office managers with automation and AI tooling, enabling a small team to take real ownership of that back-office work
We’re hiring a senior engineer to own our real-time voice stack end to end—AI agents operating on live phone calls—and the data platform that turns those calls into insight: call → transcript → events → warehouse → dashboards. You’ll own meaningful systems end to end alongside a small, senior team with deep experience in AI, product, and the trades.
What you’ll build
- New product screens and flows (jobs, customers, invoices, scheduling) in React and React Native, especially AI chat UI (chat & tool result rendering, streaming responses, human review and feedback loops)
- AI workflows in production: tool-using agents, RAG/search, classification/extraction, and human-in-the-loop flows
- Automations: Contribute new features and improvements to our AI-powered business automation platform
In addition, you’ll own our first investments into a realtime voice stack and the call-data platform behind it. For example:
- Realtime voice agents on live phone calls: telephony/WebRTC integration, streaming speech-to-text and text-to-speech, turn-taking, interruption handling, and latency optimization
- Voice pipeline reliability: backpressure, failover, graceful degradation, and monitoring for live calls
- Call-data pipeline: transcripts, events, and structured extraction flowing from every call into the warehouse
- Analytics & dashboards: data modeling and conversation-intelligence features on top of call data
- Evals & monitoring for voice agents: quality metrics, drift detection, and cost/latency tracking
- Cloud infrastructure: scaling our platform with infrastructure as code, queues and orchestration, and CI/CD
Responsibilities
- Analyze requirements and propose innovative AI-native solutions to technical problems
- Write clean scalable code
- Own the voice and data stack end-to-end: design, build, test, deploy, and operate
- Optimize the performance, latency, and cost of our real-time AI systems
- Respond to critical system issues and ensure continuous system reliability
- Mentor team members and collaborate across teams, especially with product and subject matter experts
- Work to understand the needs of our users and think creatively about how to solve design challenges in your work
- This is a Remote role. We expect a minimum 4 hours overlap with the WorkHero team (11 AM - 3 PM ET).
Qualifications
- Senior-level backend experience (typically 5+ years) shipping production systems that you've owned
- Hands-on experience with realtime voice or streaming systems: telephony (SIP/Twilio), WebRTC, streaming STT/TTS, or frameworks like LiveKit or Pipecat — or comparable experience with demanding realtime/streaming infrastructure
- Data engineering fundamentals: event pipelines, data modeling, warehousing, and analytics on production data
- Strong proficiency in a typed backend language (TypeScript preferred; comparable experience welcome)
- Hands-on experience with LLM-powered features (usage, prompting, optimization, etc) and AI architectures
- The ability to work with infrastructure as code (terraform), cloud, and CI/CD systems at scale. We're a small team, so we own the whole stack!
- Excitement to leverage AI coding tools to their maximum benefit. We love Claude Code and Cursor and are constantly looking for better ways to leverage our time to build fast and build for scale.
Nice to have
- experience with voice-AI platforms (Vapi, Retell, Bland, Deepgram, LiveKit) or conversation-intelligence products (e.g. Gong-style analytics)
- experience scaling cloud infrastructure, especially AWS, and how to get the most out of key AWS services
- experience with workflow automation tools like n8n or Lindy
- experience with React for building internal dashboards
- experience with HVAC or back-office business workflows
WorkHero is committed to building a diverse team. We encourage candidates from all backgrounds to apply.
ML Leads JD
Key Responsibilities
- Model Training & Fine-Tuning: Build, fine-tune, and optimize state-of-the-art NLP, LLM, Speech, and Vision models for scheduled Indian languages, utilizing parameter-efficient methods (LoRA, QLoRA, PEFT).
- Indic Tokenization & Linguistics: Architect custom tokenizers and text-normalization pipelines to address the "fertility problem" in Devanagari, Dravidian, and other regional scripts, ensuring low-latency and cost-effective model inference.
- Multimodal System Design: Develop robust OCR engines capable of parsing complex script geometries (conjoint consonants, Shirorekha, vowel modifiers) and integrate them into document intelligence pipelines.
- Speech Engineering: Deploy and scale robust STT (Speech-to-Text) and TTS (Text-to-Speech) pipelines capable of handling heavy code-mixing (e.g., Hinglish, Tanglish), regional accents, and localized dialects.
- Vernacular Guardrails & Evaluation: Establish culturally contextual benchmark datasets and implement safety guardrails.
- Production Deployment (MLOps): Package and serve models using high-throughput frameworks (vLLM, Triton, ONNX) optimized for GPU environments, minimizing computational overhead for massive cross-lingual workloads.
- Vernacular Fraud & Anomaly Detection: Architect risk-scoring systems and anomaly detection models capable of identifying fraud patterns in native scripts and code-mixed formats.
Essential Qualifications & Technical Skills
- Education: Bachelor’s or Master's degree in Computer Science, Mathematics, Statistics, or a closely related quantitative field.
- Experience: 4+ years of professional experience building and deploying machine learning models in production environments, with a proven track record in Indian Language NLP, Speech, or Anomaly Detection.
- Programming: Expert-level proficiency in Python and standard ML frameworks (PyTorch, TensorFlow).
- Indic AI Stack: Direct, hands-on experience with specialized Indic frameworks and datasets (e.g., AI4Bharat's IndicTrans2/IndicWhisper, Bhashini API, Kathbath, Sarvam-105B, or Aksharantar).
- Fraud Stack: Proficiency in tabular/graph-based ML toolkits (XGBoost, LightGBM, PyTorch Geometric) and handling highly imbalanced target variables (SMOTE, class weights).
- NLP & LLMs: Deep understanding of Transformer architectures, sequence-to-sequence modeling, cross-lingual embeddings, vector databases (Milvus, Pinecone, Qdrant), and quantization tools (bitsandbytes, GPTQ).
- Speech & Vision Processing: Experience processing raw audio signals (grapheme-to-phoneme conversion, spectrogram analysis) or document structures using OCR networks (CRAFT, DBNet, LayoutLM).
- Handling Code-Mixing: Proven ability to build models that gracefully parse text or speech containing heavy code-switching (mixed Latin/regional scripts, multi-language grammar).







