Cutshort logo
For Employers
Acceldata logo
Site Reliability Engineer Level 3
Site Reliability Engineer Level 3

Site Reliability Engineer Level 3 at Acceldata · Bengaluru (Bangalore) · 6 - 10 years · Raised funding · Posted 12 May 2022

Acceldata's logo

Site Reliability Engineer Level 3

Richa  Kukar's profile picture
Posted by Richa Kukar
6 - 10 yrs
Best in industry
Bengaluru (Bangalore)
Skills
SRE
Reliability engineering
Site reliability
Hadoop
HDFS
Apache HBase

Senior SRE - Acceldata (IC3 Level)


About the Job


You will join a team of highly skilled engineers who are responsible for delivering Acceldata’s support services. Our Site Reliability Engineers are trained to be active listeners and demonstrate empathy when customers encounter product issues. In our fun and collaborative environment  Site Reliability Engineers develop strong business, interpersonal and technical skills to deliver high-quality service to our valued customers.


When you arrive for your first day, we’ll want you to have:

  • Solid skills in troubleshooting to repair failed products or processes on a machine or a system using a logical, systematic search for the source of a problem in order to solve it, and make the product or process operational again
  • A strong ability to understand the feelings of our customers as we empathize with them on the issue at hand
  • A strong desire to increase your product and technology skillset; increase- your confidence supporting our products so you can help our customers succeed

In this position you will…

  • Provide Support Services to our Gold & Enterprise customers using our flagship Acceldata Pulse,Flow & Torch Product suits. This may include assistance provided during the engineering and operations of distributed systems as well as responses for mission-critical systems and production customers.
  • Demonstrate the ability to actively listen to customers and show empathy to the customer’s business impact when they experience issues with our products
  • Participate in the queue management and coordination process by owning customer escalations, managing the unassigned queue.
  • Be involved with and work on other support related activities - Performing POC & assisting Onboarding deployments of Acceldata & Hadoop distribution products.
  • Triage, diagnose and escalate customer inquiries when applicable during their engineering and operations efforts.
  • Collaborate and share solutions with both customers and the Internal team.
  • Investigate product related issues both for particular customers and for common trends that may arise
  • Study and understand critical system components and large cluster operations
  • Differentiate between issues that arise in operations, user code, or product
  • Coordinate enhancement and feature requests with product management and Acceldata engineering team.
  • Flexible in working in Shifts.
  • Participate in a Rotational weekend on-call roster for critical support needs.
  • Participate as a designated or dedicated engineer for specific customers. Aspects of this engagement translates to building long term successful relationships with customers, leading weekly status calls, and occasional visits to customer sites

In this position, you should have…

  • A strong desire and aptitude to become a well-rounded support professional. Acceldata Support considers the service we deliver as our core product.
  • A positive attitude towards feedback and continual improvement
  • A willingness to give direct feedback to and partner with management to improve team operations
  • A tenacity to bring calm and order to the often stressful situations of customer cases
  • A mental capability to multi-task across many customer situations simultaneously
  • Bachelor degree in Computer Science or Engineering or equivalent experience. Master’s degree is a plus
  • At least 2+ years of experience with at least one of the following cloud platforms: Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform (GCP), experience with managing and supporting a cloud infrastructure on any of the 3 platforms. Also knowledge on Kubernetes, Docker is a must.
  • Strong troubleshooting skills (in example, TCP/IP, DNS, File system, Load balancing, database, Java)
  • Excellent communication skills in English (written and verbal)
  • Prior enterprise support experience in a technical environment strongly preferred

Strong Hands-on Experience Working With Or Supporting The Following

  • 8-12 years of Experience with a highly-scalable, distributed, multi-node environment (50+ nodes)
  • Hadoop operation including Zookeeper, HDFS, YARN, Hive, and related components like the Hive metastore, Cloudera Manager/Ambari, etc
  • Authentication and security configuration and tuning (KNOX, LDAP, Kerberos, SSL/TLS, second priority: SSO/OAuth/OIDC, Ranger/Sentry)
  • Java troubleshooting, e.g., collection and evaluation of jstacks, heap dumps

You might also have…

  • Linux, NFS, Windows, including application installation, scripting, basic command line
  • Docker and Kubernetes configuration and troubleshooting, including Helm charts, storage options, logging, and basic kubectl CLI
  • Experience working with scripting languages (Bash, PowerShell, Python)
  • Working knowledge of application, server, and network security management concepts
  • Familiarity with virtual machine technologies
  • Knowledge of databases like MySQL and PostgreSQL,
  • Certification on any of the leading Cloud providers (AWS, Azure, GCP ) and/or Kubernetes is a big plus

The right person in this role has an opportunity to make a huge impact at Acceldata and add value to our future decisions. If this position has piqued your interest and you have what we described - we invite you to apply! An adventure in data awaits.

Learn more at https://www.acceldata.io/about-us">https://www.acceldata.io/about-us



Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Acceldata

Founded :
2018
Type :
Product
Size :
100-500
Stage :
Raised funding

About

Acceldata is the company that built the leading Multidimensional Data Observability Cloud. This cloud was designed to help data-driven organizations achieve agility in innovation, operational excellence, and enhanced returns on data investment. Embedded analytics and artificial intelligence technologies are becoming more reliant on contemporary organizations to fuel their business operations and choices.


The data observability technologies offered by Acceldata improve the performance of embedded artificial intelligence and analytics workloads by providing purpose-built monitoring and analytics. The first Data Observability Cloud is presently being developed by Acceldata for cloud data warehouses and hybrid data lakes. Acceldata makes it easy for businesses to expand their pipelines to meet the requirements of modern business, regardless of whether they are operating in a platform or cloud environment. Data Observability Cloud by Acceldata provides on-demand operational information to support analytics data workloads and embedded artificial intelligence.

Read more

Connect with the team

Profile picture
Richa Kukar
Profile picture
Abhishek Bharadawaj
Profile picture
Abhishek N
Profile picture
Swapna Chanamala
Profile picture
Akash Sampat

Company social profiles

linkedin

Similar jobs (7)

MNC
Bengaluru (Bangalore), Hyderabad
8 - 12 yrs
₹2L - ₹22L / yr
Production support
Reliability engineering
Linux/Unix
Monitoring
skill icongrafana

Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting

WFO-Immediate

8 to 12 Yrs

Bangalore/Hyderabad

Read more
EDM NETWORK
AHLOUCHE AHLOUCHE
Posted by AHLOUCHE AHLOUCHE
Remote only
1 - 10 yrs
$10K - $20K / yr (ESOP available)
User Experience (UX) Design

Key Responsibilities

Platform Monitoring and Reliability

  • Monitor the health and performance of EDM’s advertiser, publisher, broker, and internal platforms.
  • Maintain monitoring, alerting, logging, and system-health dashboards.
  • Investigate platform outages, degraded performance, failed transactions, delayed data, and integration errors.
  • Respond to production incidents and coordinate resolutions with the appropriate engineers and vendors.
  • Perform root-cause analysis and document corrective and preventive actions.
  • Help maintain defined uptime, response-time, recovery-time, and system-reliability targets.
  • Identify recurring problems and recommend permanent solutions.

Cloud Infrastructure and Systems Operations

  • Maintain and support cloud infrastructure, servers, databases, networks, storage, and production environments.
  • Support development, staging, and production environments.
  • Assist with infrastructure scaling, system upgrades, patching, backups, and disaster recovery.
  • Monitor cloud usage and help control infrastructure and technology costs.
  • Maintain access controls, service accounts, certificates, domain configurations, and environment variables.
  • Ensure production systems are properly documented and recoverable.

Deployment and Release Support

  • Support safe and consistent application deployments.
  • Maintain or improve continuous integration and deployment workflows.
  • Coordinate release schedules, deployment validation, rollback procedures, and post-release monitoring.
  • Help engineering teams identify configuration or infrastructure problems before releases reach production.
  • Maintain deployment documentation, technical checklists, and change logs.
  • Reduce manual deployment work through automation.

API and Integration Support

  • Monitor and troubleshoot third-party APIs, webhooks, postbacks, dialer connections, CRM integrations, payment systems, tracking platforms, and compliance services.
  • Investigate failed lead deliveries, missing postbacks, duplicate records, delayed reporting, and authentication problems.
  • Support ping-post, real-time bidding, call-routing, SIP, and data-transfer workflows.
  • Work with advertisers, publishers, and vendors to diagnose technical integration problems.
  • Create clear documentation for common integration methods and troubleshooting procedures.
  • Develop alerts that identify integration failures before clients report them.

Call and Lead Operations

  • Monitor call-routing, tracking, recording, attribution, and disposition systems.
  • Investigate calls that fail to route, connect, record, track, or report correctly.
  • Support number provisioning, routing rules, caps, schedules, geographic restrictions, buyer availability, and failover logic.
  • Validate that leads, calls, and transactions are properly attributed to the correct advertiser, publisher, campaign, and payout.
  • Assist with discrepancies involving call duration, billable events, conversions, payouts, and reporting.
  • Help protect revenue by identifying technical leakage and delivery failures.

Data and Reporting Support

  • Monitor data pipelines, scheduled jobs, reporting processes, and database performance.
  • Investigate discrepancies between platform reporting, billing records, payment records, and third-party systems.
  • Write and maintain database queries for troubleshooting, validation, and operational reporting.
  • Assist with data corrections using controlled and documented procedures.
  • Support dashboards and operational alerts for revenue, margin, consumption, conversion, and platform activity.
  • Maintain appropriate controls around production data access and modification.

Security and Access Management

  • Support role-based access controls, multifactor authentication, audit logging, encryption, and secure system configuration.
  • Provision and remove employee, contractor, client, and vendor access.
  • Monitor suspicious activity and report potential security incidents.
  • Assist with vulnerability remediation, security reviews, access audits, and incident-response procedures.
  • Protect consumer, advertiser, publisher, employee, and company information.
  • Follow company policies for credentials, production access, sensitive data, and change management.

Automation and Process Improvement

  • Automate repetitive operational tasks using scripts, workflows, APIs, and infrastructure tools.
  • Reduce manual work associated with monitoring, deployments, reporting, reconciliation, onboarding, and support.
  • Build internal tools that improve visibility and response times.
  • Identify operational bottlenecks that affect revenue, margin, client satisfaction, or employee productivity.
  • Maintain clear runbooks and standard operating procedures for recurring technical tasks.

Technical Support and Documentation

  • Serve as an escalation point for complex platform and integration issues.
  • Translate technical problems into clear explanations for nontechnical teams.
  • Create and maintain architecture diagrams, system inventories, runbooks, troubleshooting guides, and incident reports.
  • Track incidents and technical requests through completion.
  • Document known issues, temporary workarounds, permanent resolutions, and system dependencies.
  • Participate in an on-call rotation for urgent production incidents.

First 90-Day Priorities

The successful candidate will be expected to:

  • Learn EDM’s platforms, infrastructure, integrations, call-routing systems, reporting processes, and revenue workflows.
  • Document critical systems, dependencies, credentials ownership, vendor contacts, and escalation procedures.
  • Review existing monitoring, alerting, backups, access controls, and deployment procedures.
  • Establish baseline metrics for uptime, incident volume, response time, recovery time, deployment success, and integration failures.
  • Resolve high-priority recurring production and integration issues.
  • Improve alerting for call-routing failures, API errors, delayed data, failed jobs, and reporting discrepancies.
  • Create runbooks for the company’s most common and highest-risk technical incidents.
  • Identify at least three meaningful automation or cost-saving opportunities.
  • Participate in production support and demonstrate ownership of incidents through resolution.

Performance Expectations

Success will be measured by:

  • Platform uptime and reliability
  • Mean time to acknowledge and resolve incidents
  • Reduction in recurring production problems
  • Deployment success and rollback rates
  • API, postback, webhook, and call-routing reliability
  • Reporting and data accuracy
  • Backup and recovery readiness
  • Quality of technical documentation
  • Security and access-control compliance
  • Reduction in manual operational work
  • Infrastructure costs relative to platform volume
  • Responsiveness to internal teams, clients, and technical partners

Required Qualifications

  • Three or more years of experience in operations engineering, DevOps, site reliability engineering, cloud infrastructure, systems administration, or production support.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Experience supporting Linux-based production environments.
  • Working knowledge of networking, DNS, SSL certificates, firewalls, load balancing, and application security.
  • Experience with relational databases and SQL.
  • Experience troubleshooting REST APIs, webhooks, authentication, and third-party integrations.
  • Familiarity with monitoring, logging, alerting, and incident-management tools.
  • Experience with scripting languages such as Python, Bash, JavaScript, or PowerShell.
  • Understanding of source control, deployment pipelines, and release management.
  • Strong troubleshooting, documentation, and communication skills.
  • Ability to prioritize incidents based on business and revenue impact.
  • Availability to participate in an on-call rotation.

Preferred Qualifications

  • Experience in ad-tech, mar-tech, affiliate marketing, lead generation, pay-per-call, telecommunications, or SaaS.
  • Familiarity with SIP, VoIP, dialers, call-tracking platforms, routing systems, and phone-number provisioning.
  • Experience with containers, infrastructure as code, and automated deployment tools.
  • Experience with Docker, Kubernetes, Terraform, GitHub Actions, or similar technologies.
  • Familiarity with payment processing, usage-based billing, reconciliation, and commission systems.
  • Experience working with real-time bidding, ping-post, lead distribution, or high-volume transactional systems.
  • Understanding of TCPA-related controls, consent records, DNC suppression, data privacy, or regulated marketing environments.
  • Experience with security audits, disaster-recovery testing, and compliance documentation.

Ideal Candidate

The ideal candidate:

  • Takes ownership instead of waiting for someone else to fix the problem.
  • Remains calm and methodical during high-impact incidents.
  • Understands the difference between applying a temporary fix and eliminating a root cause.
  • Communicates technical problems clearly and directly.
  • Recognizes that production reliability is a business and revenue responsibility.
  • Automates repetitive work whenever practical.
  • Documents systems so the company is not dependent on one person’s memory.
  • Protects security and stability without creating unnecessary bureaucracy.
  • Is comfortable working in a fast-moving entrepreneurial environment.
  • Can manage competing priorities while maintaining attention to detail.


Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Farook sharief
Hyderabad
7 - 11 yrs
₹2L - ₹15L / yr
Network
Reliability engineering
skill iconAmazon Web Services (AWS)
Windows Azure

Job Summary

We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in networking, DNS, load balancing, and hybrid cloud environments. The candidate will be responsible for maintaining the reliability, availability, performance, and scalability of production infrastructure and services.

The ideal candidate should have strong troubleshooting skills and experience working across network, cloud, infrastructure, and application environments.

Key Responsibilities

  • Monitor and maintain the availability and reliability of production systems and services.
  • Troubleshoot complex network, infrastructure, and application connectivity issues.
  • Manage and troubleshoot DNS services, DNS resolution, records, and configuration issues.
  • Configure, manage, and troubleshoot Load Balancers and traffic routing.
  • Work with Layer 4 and Layer 7 networking and understand TCP/IP, HTTP/HTTPS, routing, and network connectivity.
  • Support hybrid cloud environments involving on-premises infrastructure and public cloud platforms.
  • Troubleshoot connectivity between on-premises data centers and cloud environments.
  • Participate in production incidents, troubleshooting, root cause analysis (RCA), and problem management.
  • Develop automation scripts and tools to reduce manual operational activities.
  • Configure and maintain monitoring, alerting, and observability solutions.
  • Work closely with Network, Cloud, DevOps, Security, and Application teams.
  • Participate in on-call support and resolve production issues within defined SLAs.
  • Document infrastructure, troubleshooting procedures, incident reports, and operational processes.
  • Identify opportunities to improve system reliability, performance, and scalability.
Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Akilandeswari Panneerselvam
Hyderabad
5 - 10 yrs
₹6L - ₹10L / yr
SRE
Production support
Cloud Computing
skill iconKubernetes
Linux/Unix
+1 more

Site Reliability Engineer (SRE) / Production Support Engineer

Experience: 5–10 Years

Location: Hyderabad

Work Mode: Face-to-Face Drive

Shift: Rotational Shifts

Job Description

Looking for an experienced SRE / Production Support Engineer with strong experience in application and production support, incident management, monitoring, troubleshooting, and cloud operations.

Key Skills

Production Support, Incident Management, Splunk, APM, SLI/SLO, Cloud, Kubernetes, Docker, Terraform, Linux/Windows Administration, Shell Scripting and Python.

Good understanding of production deployments, batch monitoring, network/load balancing, and troubleshooting is required.

Candidates from SRE, Production Support, Application Support, Cloud Operations, or DevOps backgrounds are preferred.

Read more
Hevo Data
at Hevo Data
1 video
7 recruiters
Anushka Sarkar
Posted by Anushka Sarkar
Bengaluru (Bangalore)
2 - 6 yrs
₹12L - ₹17L / yr (ESOP available)
Data Warehouse (DWH)
ETL
Snowflake
skill iconPostgreSQL
BigQuery
+5 more

As a Solutions Engineer/Sales Engineer, you will act as a trusted technical advisor to prospective customers in the US market. You will own the technical evaluation journey end-to-end, from the first discovery call to closing the deal, by demonstrating how our platform solves real-world data integration challenges. This role involves working closely with global customers (primarily North America ), designing scalable architectures, running tailored product demos, and building proofs-of-concept that showcase value.


Responsibilities:

  • Run technical discovery calls with Heads of Data, Data Engineering Managers, and CTOs to understand their stack, pain points, and constraints. These are global conversations, primarily with customers in North America and Europe, and you will own them end-to-end.
  • Design and present reference architectures showing how Hevo fits into their existing data stack (e. g., Salesforce, Postgres, GA4 Snowflake, dbt).
  • Run live product demos tailored to the customer's stack. These are not scripted walkthroughs. You will need to read the room, adapt on the fly, and handle unexpected technical questions with confidence.
  • Own the technical evaluation from the first conversation to the closed deal by addressing deep technical questions and earning trust with engineering teams.
  • Build proofs-of-concept by configuring real pipelines from source systems (Salesforce, Outreach, Postgres, MySQL) into destinations like Snowflake or BigQuery.
  • Develop deep expertise in the data integration landscape, including native solutions from AWS, GCP, and Azure (AWS Glue, Azure Data Factory, Google Cloud Dataflow) as well as other data pipeline and integration platforms, and articulate Hevo's technical advantages in a way that lands with both technical and non-technical buyers.
  • Partner with Account Executives to drive technical wins and with Customer Success to ensure smooth handoffs after a deal closes.
  • Share customer feedback and patterns with Product and Engineering to influence the roadmap.


Requirements:

  • 2 to 6 years of hands-on experience as a Data engineer, Analytics engineer, ETL/ELT developer, or Data warehouse engineer.
  • Experience building and operating real data pipelines in production, including handling source schema changes, failures, and reliability issues.
  • Strong fundamentals in SQL and at least one cloud data warehouse (Snowflake, BigQuery, or Redshift).
  • Experience with at least one orchestration or transformation tool (Airflow, dbt, or similar).
  • Exceptional communication and presentation skills.
  • This is non-negotiable. You will be the face of Hevo's technical capability to a global customer base. That means:
  • Your spoken English is clear, confident, and easy to follow for an international audience.
  • You are comfortable presenting to groups of 5 to 15 people, including senior technical leaders, over video calls.
  • You can structure a technical narrative, not just answer questions but guide a conversation toward the right conclusion.
  • You can hold your own when a customer pushes back or challenges your recommendation.
  • You have presented or demoed something technical to an audience before, in any context: internal reviews, client calls, meetups, or conferences.
  • Genuine interest in customer-facing work, even if you have not done it formally before.


Good to have:

  • Experience with CDC tools or data integration platforms, whether cloud-native (AWS Glue, Azure Data Factory, or Google Cloud Dataflow) or third-party pipeline and integration tools.
  • Exposure to multiple source system types such as CRMs, ad platforms, SaaS APIs, and OLTP databases.
  • Public technical writing, conference talks, or open-source contributions.
  • Prior experience in customer conversations, even informally. For example, presenting to internal stakeholders or analytics teams or working with international teams across time zones.
Read more
Deltek
Remote only
3 - 5 yrs
Best in industry
CI/CD
skill iconPostgreSQL
skill iconPython
skill iconAmazon Web Services (AWS)
Artificial Intelligence (AI)
+2 more

SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.


Core responsibilities include:

  • Production monitoring and debugging of live systems.
  • Incident investigation, troubleshooting, and problem resolution.
  • AWS cloud infrastructure support and maintenance.
  • Deployment and operational support activities.
  • Supporting a 24x7 production environment.
  • Working with GitHub-based development workflows.
  • Technical debt remediation and platform improvements.
  • Customer issue investigation and support.
  • Security and compliance-related work, including FedRAMP initiatives.


Preferred Skills:

AWS (especially S3 and EC2)

Strong debugging and troubleshooting skills

Site Reliability Engineering (SRE) experience

GitHub experience

Basic software development skills

TypeScript/JavaScript knowledge

C# preferred

AI experience is a plus.


Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen
Hyderabad
7 - 11 yrs
₹2L - ₹15L / yr
Reliability engineering
skill iconAmazon Web Services (AWS)
Windows Azure
Network
DNS

SRE – Network / DNS / Load Balancer

Experience: 7–11 Years

Location: Hyderabad

Work Mode: WFO

Availability: Immediate Joiner

Job Description:

  • Strong experience in Site Reliability Engineering (SRE) with focus on infrastructure and application reliability.
  • Hands-on experience with Network, DNS and Load Balancer troubleshooting and administration.
  • Monitor system performance, availability, latency and infrastructure health.
  • Troubleshoot network connectivity, DNS resolution, routing and load-balancing issues.
  • Experience with load balancers such as F5, BIG-IP, HAProxy or similar technologies.
  • Good understanding of TCP/IP, HTTP/HTTPS, LAN/WAN, SSL/TLS and networking concepts.
  • Experience with DNS technologies such as BIND, Infoblox or equivalent.
  • Work on incident management, root cause analysis and problem resolution.
  • Collaborate with application, network, cloud and infrastructure teams to resolve production issues.
  • Experience with monitoring and alerting tools such as Splunk, Grafana, Prometheus, AppDynamics or similar tools.
  • Strong troubleshooting, production support and communication skills.
  • Willingness to work from Hyderabad office (WFO) and join immediately. 


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos