

EDM NETWORK
https://edmleadnetwork.comJobs at EDM NETWORK
Key Responsibilities
Platform Monitoring and Reliability
- Monitor the health and performance of EDM’s advertiser, publisher, broker, and internal platforms.
- Maintain monitoring, alerting, logging, and system-health dashboards.
- Investigate platform outages, degraded performance, failed transactions, delayed data, and integration errors.
- Respond to production incidents and coordinate resolutions with the appropriate engineers and vendors.
- Perform root-cause analysis and document corrective and preventive actions.
- Help maintain defined uptime, response-time, recovery-time, and system-reliability targets.
- Identify recurring problems and recommend permanent solutions.
Cloud Infrastructure and Systems Operations
- Maintain and support cloud infrastructure, servers, databases, networks, storage, and production environments.
- Support development, staging, and production environments.
- Assist with infrastructure scaling, system upgrades, patching, backups, and disaster recovery.
- Monitor cloud usage and help control infrastructure and technology costs.
- Maintain access controls, service accounts, certificates, domain configurations, and environment variables.
- Ensure production systems are properly documented and recoverable.
Deployment and Release Support
- Support safe and consistent application deployments.
- Maintain or improve continuous integration and deployment workflows.
- Coordinate release schedules, deployment validation, rollback procedures, and post-release monitoring.
- Help engineering teams identify configuration or infrastructure problems before releases reach production.
- Maintain deployment documentation, technical checklists, and change logs.
- Reduce manual deployment work through automation.
API and Integration Support
- Monitor and troubleshoot third-party APIs, webhooks, postbacks, dialer connections, CRM integrations, payment systems, tracking platforms, and compliance services.
- Investigate failed lead deliveries, missing postbacks, duplicate records, delayed reporting, and authentication problems.
- Support ping-post, real-time bidding, call-routing, SIP, and data-transfer workflows.
- Work with advertisers, publishers, and vendors to diagnose technical integration problems.
- Create clear documentation for common integration methods and troubleshooting procedures.
- Develop alerts that identify integration failures before clients report them.
Call and Lead Operations
- Monitor call-routing, tracking, recording, attribution, and disposition systems.
- Investigate calls that fail to route, connect, record, track, or report correctly.
- Support number provisioning, routing rules, caps, schedules, geographic restrictions, buyer availability, and failover logic.
- Validate that leads, calls, and transactions are properly attributed to the correct advertiser, publisher, campaign, and payout.
- Assist with discrepancies involving call duration, billable events, conversions, payouts, and reporting.
- Help protect revenue by identifying technical leakage and delivery failures.
Data and Reporting Support
- Monitor data pipelines, scheduled jobs, reporting processes, and database performance.
- Investigate discrepancies between platform reporting, billing records, payment records, and third-party systems.
- Write and maintain database queries for troubleshooting, validation, and operational reporting.
- Assist with data corrections using controlled and documented procedures.
- Support dashboards and operational alerts for revenue, margin, consumption, conversion, and platform activity.
- Maintain appropriate controls around production data access and modification.
Security and Access Management
- Support role-based access controls, multifactor authentication, audit logging, encryption, and secure system configuration.
- Provision and remove employee, contractor, client, and vendor access.
- Monitor suspicious activity and report potential security incidents.
- Assist with vulnerability remediation, security reviews, access audits, and incident-response procedures.
- Protect consumer, advertiser, publisher, employee, and company information.
- Follow company policies for credentials, production access, sensitive data, and change management.
Automation and Process Improvement
- Automate repetitive operational tasks using scripts, workflows, APIs, and infrastructure tools.
- Reduce manual work associated with monitoring, deployments, reporting, reconciliation, onboarding, and support.
- Build internal tools that improve visibility and response times.
- Identify operational bottlenecks that affect revenue, margin, client satisfaction, or employee productivity.
- Maintain clear runbooks and standard operating procedures for recurring technical tasks.
Technical Support and Documentation
- Serve as an escalation point for complex platform and integration issues.
- Translate technical problems into clear explanations for nontechnical teams.
- Create and maintain architecture diagrams, system inventories, runbooks, troubleshooting guides, and incident reports.
- Track incidents and technical requests through completion.
- Document known issues, temporary workarounds, permanent resolutions, and system dependencies.
- Participate in an on-call rotation for urgent production incidents.
First 90-Day Priorities
The successful candidate will be expected to:
- Learn EDM’s platforms, infrastructure, integrations, call-routing systems, reporting processes, and revenue workflows.
- Document critical systems, dependencies, credentials ownership, vendor contacts, and escalation procedures.
- Review existing monitoring, alerting, backups, access controls, and deployment procedures.
- Establish baseline metrics for uptime, incident volume, response time, recovery time, deployment success, and integration failures.
- Resolve high-priority recurring production and integration issues.
- Improve alerting for call-routing failures, API errors, delayed data, failed jobs, and reporting discrepancies.
- Create runbooks for the company’s most common and highest-risk technical incidents.
- Identify at least three meaningful automation or cost-saving opportunities.
- Participate in production support and demonstrate ownership of incidents through resolution.
Performance Expectations
Success will be measured by:
- Platform uptime and reliability
- Mean time to acknowledge and resolve incidents
- Reduction in recurring production problems
- Deployment success and rollback rates
- API, postback, webhook, and call-routing reliability
- Reporting and data accuracy
- Backup and recovery readiness
- Quality of technical documentation
- Security and access-control compliance
- Reduction in manual operational work
- Infrastructure costs relative to platform volume
- Responsiveness to internal teams, clients, and technical partners
Required Qualifications
- Three or more years of experience in operations engineering, DevOps, site reliability engineering, cloud infrastructure, systems administration, or production support.
- Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Experience supporting Linux-based production environments.
- Working knowledge of networking, DNS, SSL certificates, firewalls, load balancing, and application security.
- Experience with relational databases and SQL.
- Experience troubleshooting REST APIs, webhooks, authentication, and third-party integrations.
- Familiarity with monitoring, logging, alerting, and incident-management tools.
- Experience with scripting languages such as Python, Bash, JavaScript, or PowerShell.
- Understanding of source control, deployment pipelines, and release management.
- Strong troubleshooting, documentation, and communication skills.
- Ability to prioritize incidents based on business and revenue impact.
- Availability to participate in an on-call rotation.
Preferred Qualifications
- Experience in ad-tech, mar-tech, affiliate marketing, lead generation, pay-per-call, telecommunications, or SaaS.
- Familiarity with SIP, VoIP, dialers, call-tracking platforms, routing systems, and phone-number provisioning.
- Experience with containers, infrastructure as code, and automated deployment tools.
- Experience with Docker, Kubernetes, Terraform, GitHub Actions, or similar technologies.
- Familiarity with payment processing, usage-based billing, reconciliation, and commission systems.
- Experience working with real-time bidding, ping-post, lead distribution, or high-volume transactional systems.
- Understanding of TCPA-related controls, consent records, DNC suppression, data privacy, or regulated marketing environments.
- Experience with security audits, disaster-recovery testing, and compliance documentation.
Ideal Candidate
The ideal candidate:
- Takes ownership instead of waiting for someone else to fix the problem.
- Remains calm and methodical during high-impact incidents.
- Understands the difference between applying a temporary fix and eliminating a root cause.
- Communicates technical problems clearly and directly.
- Recognizes that production reliability is a business and revenue responsibility.
- Automates repetitive work whenever practical.
- Documents systems so the company is not dependent on one person’s memory.
- Protects security and stability without creating unnecessary bureaucracy.
- Is comfortable working in a fast-moving entrepreneurial environment.
- Can manage competing priorities while maintaining attention to detail.
Similar companies
About the company
Deep Tech Startup Focusing on Autonomy and Intelligence for Unmanned Systems. Guidance and Navigation, AI-ML, Computer Vision, Information Fusion, LLMs, Generative AI, Remote Sensing
Jobs
4
About the company
AI that writes Credit Reports on its own !!
AbleCredit is a friendly, supportive credit assistant. It generates Credit Appraisal Memos based on your policies, without any human intervention.
Jobs
12
About the company
Jobs
185
About the company
OIP Insurtech streamlines insurance operations and optimizes workflows by combining deep industry knowledge with advanced technology. Established in 2012, OIP InsurTech partners with carriers, MGAs, program managers, and TPAs in the US, Canada, and Europe, especially the UK.
With 1,200 professionals serving over 100 clients, we deliver insurance process automation, custom software development, high-quality underwriting services, and skilled tech staff to augment our clients.
While saving time and money is the immediate win, the real game-changer is giving our clients the freedom to grow their books, run their businesses, and focus on what they love. We’re proud to support them on this journey and make a positive impact on the industry!
Jobs
7
About the company
Automate Accounts is a technology-driven company dedicated to building intelligent automation solutions that streamline business operations and boost efficiency. We leverage modern platforms and tools to help businesses transform their workflows with cutting-edge solutions.
Jobs
5
About the company
Jobs
5
About the company
Stairio is a digital infrastructure company building scalable online systems for modern businesses.
We help service-driven brands establish strong digital foundations through high-performance websites, booking systems, management dashboards, and integrated payment solutions. Our goal is to give businesses ownership, control, and long-term digital assets that generate measurable revenue.
Jobs
2
About the company
Jobs
13
About the company
Sisuni Technology Private Limited is a Bengaluru-based technology and product development company focused on building its own e-commerce, SaaS, HRM and cloud-based technology products.
Jobs
6




