Digital Audio Specialist at Digital Product Engineering company · Remote only · 5 - 7 years · ₹25L - ₹28L / yr · Remote only · Posted 1 Jun 2026

Staff Engineer, Digital Audio Specialist
Location: Remote
Employment Type: Full-time
Experience: 5.5-7 years
Job description
REQUIREMENTS:
- Strong experience in Digital Audio Engineering, Audio Technology Platforms, and Media Infrastructure
- Deep hands-on expertise in Digital Audio Processing, Encoding, Streaming, and Audio Delivery Technologies (Must Have)
- Strong experience with Audio Content Management Systems and Digital Media Workflows (Must Have)
- Hands-on experience with audio editing, mixing, mastering, and broadcast technologies
- Strong understanding of audio codecs, compression standards, and streaming protocols (Must Have)
- Experience with cloud-based media platforms and audio content distribution networks
- Strong experience with automation and workflow optimization in digital media environments
- Experience managing large-scale audio content libraries and metadata systems
- Hands-on experience with Digital Asset Management (DAM) platforms
- Strong understanding of audio quality assurance, testing, and troubleshooting
- Experience integrating audio platforms with third-party applications and APIs
- Knowledge of podcasting, internet radio, OTT audio, and streaming ecosystems
- Experience with monitoring, analytics, and performance optimization of audio platforms
- Strong understanding of security, compliance, and content rights management
- Experience working in Agile/Scrum environments
- Strong troubleshooting, analytical, and problem-solving skills
- Excellent communication and stakeholder management skills
RESPONSIBILITIES:
- Design, implement, and manage digital audio platforms and content delivery systems
- Support audio processing, encoding, streaming, and distribution workflows
- Manage audio content lifecycle including ingestion, storage, metadata, and publishing
- Ensure high-quality audio delivery across web, mobile, and streaming platforms
- Develop and maintain automation solutions for audio operations and media workflows
- Troubleshoot audio quality issues, streaming failures, and platform performance problems
- Collaborate with content, engineering, and product teams to deliver audio solutions
- Manage integrations between audio platforms, content management systems, and third-party services
- Monitor platform health, streaming performance, and user experience metrics
- Implement best practices for audio quality, scalability, reliability, and security
- Support digital broadcasting, podcasting, and on-demand audio initiatives
- Perform root cause analysis and resolution for production incidents
- Define standards, processes, and governance for audio platform operations
- Optimize infrastructure and workflows to improve operational efficiency
- Mentor junior engineers and provide technical leadership on digital audio initiatives
- Participate in Agile ceremonies and continuous improvement programs
- Work closely with stakeholders to understand business requirements and deliver optimal solutions
- Ensure compliance with content rights, licensing, and regulatory requirements
Qualifications
Bachelor’s or master’s degree in computer science, Information Technology, or a related fields

Similar jobs (4)
Procedure is hiring for WorkHero.
WorkHero is building the AI-powered back office for the skilled trades, starting with the $50B+ HVAC industry. Small contractors are great at their trade but lose 20+ hours a week to invoicing, permits, scheduling, and paperwork. WorkHero combines expert office managers with automation and AI tooling, enabling a small team to take real ownership of that back-office work
We’re hiring a senior engineer to own our real-time voice stack end to end—AI agents operating on live phone calls—and the data platform that turns those calls into insight: call → transcript → events → warehouse → dashboards. You’ll own meaningful systems end to end alongside a small, senior team with deep experience in AI, product, and the trades.
What you’ll build
- New product screens and flows (jobs, customers, invoices, scheduling) in React and React Native, especially AI chat UI (chat & tool result rendering, streaming responses, human review and feedback loops)
- AI workflows in production: tool-using agents, RAG/search, classification/extraction, and human-in-the-loop flows
- Automations: Contribute new features and improvements to our AI-powered business automation platform
In addition, you’ll own our first investments into a realtime voice stack and the call-data platform behind it. For example:
- Realtime voice agents on live phone calls: telephony/WebRTC integration, streaming speech-to-text and text-to-speech, turn-taking, interruption handling, and latency optimization
- Voice pipeline reliability: backpressure, failover, graceful degradation, and monitoring for live calls
- Call-data pipeline: transcripts, events, and structured extraction flowing from every call into the warehouse
- Analytics & dashboards: data modeling and conversation-intelligence features on top of call data
- Evals & monitoring for voice agents: quality metrics, drift detection, and cost/latency tracking
- Cloud infrastructure: scaling our platform with infrastructure as code, queues and orchestration, and CI/CD
Responsibilities
- Analyze requirements and propose innovative AI-native solutions to technical problems
- Write clean scalable code
- Own the voice and data stack end-to-end: design, build, test, deploy, and operate
- Optimize the performance, latency, and cost of our real-time AI systems
- Respond to critical system issues and ensure continuous system reliability
- Mentor team members and collaborate across teams, especially with product and subject matter experts
- Work to understand the needs of our users and think creatively about how to solve design challenges in your work
- This is a Remote role. We expect a minimum 4 hours overlap with the WorkHero team (11 AM - 3 PM ET).
Qualifications
- Senior-level backend experience (typically 5+ years) shipping production systems that you've owned
- Hands-on experience with realtime voice or streaming systems: telephony (SIP/Twilio), WebRTC, streaming STT/TTS, or frameworks like LiveKit or Pipecat — or comparable experience with demanding realtime/streaming infrastructure
- Data engineering fundamentals: event pipelines, data modeling, warehousing, and analytics on production data
- Strong proficiency in a typed backend language (TypeScript preferred; comparable experience welcome)
- Hands-on experience with LLM-powered features (usage, prompting, optimization, etc) and AI architectures
- The ability to work with infrastructure as code (terraform), cloud, and CI/CD systems at scale. We're a small team, so we own the whole stack!
- Excitement to leverage AI coding tools to their maximum benefit. We love Claude Code and Cursor and are constantly looking for better ways to leverage our time to build fast and build for scale.
Nice to have
- experience with voice-AI platforms (Vapi, Retell, Bland, Deepgram, LiveKit) or conversation-intelligence products (e.g. Gong-style analytics)
- experience scaling cloud infrastructure, especially AWS, and how to get the most out of key AWS services
- experience with workflow automation tools like n8n or Lindy
- experience with React for building internal dashboards
- experience with HVAC or back-office business workflows
WorkHero is committed to building a diverse team. We encourage candidates from all backgrounds to apply.
About Jinn
Jinn is a Voice AI Tech Company. It helps businesses get better RoI specially by helping sales processes with the use of Voice AI Tech. We are adding a business line which includes an audio device that can reliably capture audio in various environments.
About role:
- This Role in a B2B SaaS startup in AI space led by 2X entrepreneurs from IIT, IIMs. Fast paced with a lot of learning and growth.
- Responsibility: Helping engineer/assemble/bring together an IoT/ hardware device that can accomplish product goals with required constraints. Great high stakes exposure for fresh grads
- Duration: 3-6 months internship || Converts to Full Time based on performance
- Compensation (Stipend): 20-25k per month || Full time 4.5lpa - 6LPA
Ideal Profile: Interested in building a career in IoT/Tech, good communication, good discipline, solid understand of tech (AI)
Skill Sets
Look for someone who:
• Has built at least 1 IoT project end-to-end
• Knows Arduino + one of ESP32 / nRF52
• Has touched audio input (even basic)
• Is comfortable debugging hardware (this is key)
1. Embedded Systems Programming (Must-have)
• C/C++ (Arduino framework or ESP-IDF)
• Working with:
* ESP32 OR
* Seeed Studio XIAO BLE nRF52840 Sense
• Skills:
* GPIO, I2S (for mic input)
* Power modes (deep sleep, wake triggers)
* Memory constraints (huge in audio use cases)
👉 This is the backbone. If they can’t do this well, project stalls.
2. Audio Handling + Signal Basics
• Understanding:
* Sampling rate (16kHz vs 44.1kHz)
* PCM audio buffers
* Latency vs quality trade-offs
• Practical skills:
* Using I2S microphones (INMP441, etc.)
* Basic noise filtering
* Voice Activity Detection (VAD)
👉 Without this, you’ll just get noisy unusable recordings.
3. Power & Hardware Basics (Often underestimated)
• LiPo battery handling
• Charging IC (TP4056 type)
• Power optimization:
* Sleep modes
* Sampling intervals
Many prototypes fail here (battery drains in 1 hour 😅)
4. Connectivity (BLE / WiFi)
• BLE (for XIAO nRF52840):
* Data chunking (BLE MTU limits!)
* Pairing + mobile relay model
• WiFi (for ESP32):
* HTTP / WebSocket streaming
* Retry + buffering
Trade
About the Company
The client is revolutionising the way businesses operate through cutting-edge technological solutions. Their focus is on developing intelligent agents and agentic workflows that automate processes and eliminate the need for human effort wherever possible. By leveraging
advanced AI and machine learning, they create systems that enhance productivity and drive efficiency.
Their expertise extends to the fintech, healthcare and medical technology sectors, where they develop innovative solutions that improve patient outcomes and streamline medical operations.
From medical devices to healthcare platforms, their work sits at the intersection of technology and medicine, pushing the boundaries of what's possible. The team is dedicated to continuous learning and growth, ensuring the team members are always at the forefront of the tech landscape.
About the Role
This is a senior, hands-on engineering role at the heart of our product team. You will be one of the most technical people in the room — setting the architecture for our real-time voice AI
agents and building the hardest parts of it yourself. From the systems that power live conversations to the interfaces our clients rely on, you will own how the product is engineered end to end.
We are looking for a genuine lead full-stack engineer with the depth to make architecture decisions that hold up as we scale, and the appetite to still be in the code every day. You should be as comfortable designing the backend services behind a live voice agent as you are shaping a clean interface on top of them — and comfortable being the person others turn to when something is hard.
You will work directly with the founder and product leadership on a fast-moving product, with real influence over technical direction. This is a role for someone who wants ownership at the level of "how the whole thing is built," not just individual features — and who raises the bar for
everyone around them.
What You'll Own
Set the technical direction
- Own the architecture of our core systems — the real-time voice agents, backend
- services, data and APIs — making the decisions that keep the product fast, reliable and scalable as it grows.
- Lead the hardest engineering problems and solve them personally.
- Establish engineering standards — code quality, review practices, testing and technical patterns that the team builds to.
- Drive technical strategy with the founder and product leadership — shaping the roadmap, flagging risk early, and turning product ambition into a sound technical plan.
Build the product end to end
- Design, build and ship features across the stack — backend services, APIs and front-ends — owning them from idea to production.
- Build the client-facing surfaces — dashboards, review tools and configuration interfaces that let our clients run and trust the product.
- Design and evolve the data models and APIs that hold up as we scale across clients.
Make it reliable and fast
- Own production quality — put the monitoring and alerting in place so issues are caught before clients feel them, and performance stays within target.
- Care about performance — find and fix bottlenecks across the stack.
- Build for correctness — put the testing and evaluation in place that keeps the product behaving predictably as it changes.
Lead through the team
- Mentor and grow engineers — through code review, pairing, and setting a technical example others learn from.
- Multiply the team's output — unblock others and lift the overall quality of the codebase.
- Take features from ambiguity to done — turn a rough product goal into a shipped, working capability with minimal hand-holding, and help others do the same.
What We're Looking For
- 8+ years of professional software engineering experience, with significant depth across backend and a track record of owning systems, not just features.
- Strong backend engineering, ideally in Python — building and scaling production services and APIs..
- Proven architecture and system-design ability — you have designed systems that scaled, and can reason clearly about trade-offs.
- Solid fundamentals across APIs, databases and cloud infrastructure.
- Experience building real-time and/or AI-powered products — or clear, demonstrable ability to lead in this area.
- A history of technical leadership — setting standards, mentoring engineers, and being trusted with the hardest problems — while remaining hands-on.
- Excellent communication and a genuine ownership mindset — someone who can be handed an ambiguous, high-stakes problem and be trusted to see it through.
Nice to Have
- Experience working with AI / large language models in production.
- Experience with voice or other real-time products.
- Exposure to healthcare, fintech, or other regulated / high-stakes domains.
- Experience as an early or senior engineer in a startup, where you set direction and wore many hats.

Role overview
The client is building a multimodal AI platform that processes multi-hour video, audio and text to generate structured insights, narratives and highlight workflows for broadcasters and media organisations.
We are seeking a Backend / Platform Engineer to design and build high-throughput media pipelines, robust APIs, and model-serving infrastructure that connect our AI engine (video perception + multimodal reasoning) to real products and customer environments.
This is not a CRUD‑only backend role.
You will work on:
- long‑running jobs
- distributed processing
- GPU inference orchestration
- storage for embeddings and metadata
- integration with AI models
- reliability and observability at scale
Key responsibilities
Media ingestion & processing pipelines
- Design and implement ingestion pipelines for multi‑hour video and audio content.
- Build microservices for frame extraction, audio processing, transcription integration and metadata generation.
- Handle long‑running, asynchronous jobs using queues, workers and robust retry strategies.
- Integrate with FFmpeg or similar tools for transcoding, segmenting and preparing media for AI models.
API & platform architecture
- Design and implement REST/gRPC APIs that expose AI model outputs (perception, multimodal alignment, narratives) to frontend and external systems.
- Define clear contracts for internal services and external integrations.
- Implement authentication, authorisation and rate‑limiting for platform endpoints.
- Ensure backward‑compatible API evolution as the product matures.
Model‑serving & AI integration
- Integrate with AI inference services (video models, multimodal models, LLM/VLM) running on GPUs or specialised infrastructure.
- Design request/response flows that handle large payloads, streaming outputs and structured results.
- Optimise throughput and latency for inference pipelines, including batching, caching and concurrency control.
- Collaborate closely with AI engineers to productionise models and debug end‑to‑end behaviour.
Storage, data models & performance
- Design data models to store embeddings, timelines, metadata, scene/shot boundaries, and narrative units.
- Work with appropriate storage technologies (SQL/NoSQL, object storage, search indices) based on access patterns.
- Implement indexing and query strategies for fast retrieval of segments, highlights and multimodal insights.
- Optimise performance for large datasets and high‑volume workloads.
Reliability, observability & operations
- Implement logging, metrics and tracing across services for debugging and monitoring.
- Set up health checks, circuit breakers and graceful degradation for critical services.
- Work with CI/CD pipelines to ensure safe, repeatable deployments.
- Collaborate on Kubernetes‑based deployments (or equivalent orchestration) for scaling services.
Requirements (must‑have)
Experience:
- 4–8 years in backend or platform engineering.
- At least 3 years working on distributed systems, high‑throughput services or complex pipelines (not just simple CRUD apps).
Languages & frameworks:
- Strong proficiency in Python or Node.js (one primary, both are a plus).
- Experience with at least one modern backend framework (FastAPI, Flask, Express, NestJS, etc.).
Distributed systems & pipelines:
- Hands‑on experience with queues and workers (e.g. Celery, RabbitMQ, Kafka, SQS, etc.).
- Experience building asynchronous, long‑running job pipelines.
- Understanding of idempotency, retries, backoff, and failure handling.
APIs & integration:
- Strong experience designing and implementing REST APIs (gRPC is a plus).
- Experience integrating with external services and handling network‑level failures.
Cloud & infrastructure:
- Experience deploying services on AWS, GCP or Azure (EC2/Compute Engine, S3/GCS, IAM, networking basics).
- Experience with Docker; exposure to Kubernetes is a strong plus.
Data & storage:
- Experience with SQL and at least one NoSQL store.
- Ability to design schemas and data models for performance and maintainability.
Engineering quality:
- Strong debugging skills across services and environments.
- Experience with unit/integration tests for backend systems.
- Clear, structured communication in English.
Nice‑to‑have
- Experience with media/video processing (FFmpeg, transcoding, segmenting).
- Experience with AI/ML model integration (serving models, handling inference requests).
- Experience with search/retrieval systems (e.g. Elasticsearch, vector databases).
- Experience with observability stacks (Prometheus, Grafana, OpenTelemetry).
- Experience working with remote teams across time zones.
What we are explicitly NOT looking for
To reduce noise and mismatches, we are not looking for:
- Pure CRUD‑only backend developers with no pipeline or distributed systems experience.
- Engineers who have only worked on small, single‑service apps without scale or complexity.
- Candidates who cannot explain trade‑offs in architecture, data modelling and reliability.
- Candidates who are uncomfortable with ownership of subsystems end‑to‑end.
Why join us
- Work on real, complex problems at the intersection of media, AI and distributed systems.
- Collaborate with senior AI engineers working on perception, multimodal fusion and narrative reasoning.
- Build the core platform that turns AI models into a usable product for broadcasters and media organisations.
- Operate with high ownership, clear expectations and direct access to the CTO.





