Site reliability Engineer at Acceldata · Bengaluru (Bangalore) · 3 - 8 years · ₹20L - ₹40L / yr (ESOP available) · Raised funding · Posted 12 Jan 2022
Responsibilities
- Our Site reliability engineers work on improving the availability, scalability, performance, and reliability of enterprise production services for our products as well as our customer’s data lake environments.
- You will use your expertise to improve the reliability and performance of Hadoop Data lake clusters and data management services. Just as our products, our SRE are expected to be platform and vendor-agnostic when it comes to implementing, stabilizing, and tuning Hadoop ecosystems.
- You’d be required to provide implementation guidance, best practices framework, and technical thought leadership to our customers for their Hadoop Data lake implementation and migration initiatives.
- You need to be 100% hand-on and as a required test, monitor, administer, and operate multiple Data lake clusters across data centers.
- Troubleshoot issues across the entire stack - hardware, software, application, and network.
- Dive into problems with an eye to both immediate remediations as well as the follow-through changes and automation that will prevent future occurrences.
- Must demonstrate exceptional troubleshooting and strong architectural skills and clearly and effectively describe this in both a verbal and written format.
Requirements
- Customer-focused, Self-driven, and Motivated with a strong work ethic and a passion for problem-solving.
- 4+ years of designing, implementing, tuning, and managing services in a distributed, enterprise-scale on-premise and public/private cloud environment.
- Familiarity with infrastructure management and operations lifecycle concepts and ecosystem.
- Hadoop cluster design, Implementation, management and performance tuning experience with HDFS, YARN,
- HIVE/IMPALA, SPARK, Kerberos and related Hadoop technologies are a must.
- Must have strong SQL/HQL query troubleshooting and tuning skills on Hive/HBase.
- Must have a strong capacity planning experience for Hadoop ecosystems/data lakes.
- Good to have hands-on experience with – KAFKA, RANGER/SENTRY, NiFi, Ambari, Cloudera Manager, and HBASE.
- Good to have data modeling, data engineering, and data security experience within the Hadoop ecosystem.Good to have deep JVM/Java debugging and tuning skills.

About Acceldata
About
Acceldata is the company that built the leading Multidimensional Data Observability Cloud. This cloud was designed to help data-driven organizations achieve agility in innovation, operational excellence, and enhanced returns on data investment. Embedded analytics and artificial intelligence technologies are becoming more reliant on contemporary organizations to fuel their business operations and choices.
The data observability technologies offered by Acceldata improve the performance of embedded artificial intelligence and analytics workloads by providing purpose-built monitoring and analytics. The first Data Observability Cloud is presently being developed by Acceldata for cloud data warehouses and hybrid data lakes. Acceldata makes it easy for businesses to expand their pipelines to meet the requirements of modern business, regardless of whether they are operating in a platform or cloud environment. Data Observability Cloud by Acceldata provides on-demand operational information to support analytics data workloads and embedded artificial intelligence.
Connect with the team
Similar jobs (1)
Job Summary:
- We are seeking an experienced Hadoop Engineer with strong hands-on expertise in MAPR and Hortonworks platform engineering and administration. The successful candidate will join our Hadoop Platform Engineering & Automation team to support our critical enterprise data lake platform. This role emphasizes platform support, patching, vulnerability remediation, L3 issue resolution, and automation initiatives utilizing Linux, Shell Scripting, and Ansible.
Responsibilities:
- Platform Engineering & Administration: Administer, maintain, and optimize MAPR and Hortonworks Hadoop clusters.
- Provide L3 support for complex technical issues related to cluster functioning, performance, and reliability.
- Monitor cluster health, capacity, performance, and security configurations.
- Patching & Vulnerability Remediation: Execute OS-level and platform-level patching across large-scale Hadoop clusters.
- Implement remediation for platform vulnerabilities in accordance with organizational InfoSec policies.
- Collaborate with other support teams to ensure compliance with standards and mitigation timelines.
- Automation Development: Develop, enhance, and maintain automation workflows using Linux, Shell Scripting, and Ansible.
- Automate recurring operational tasks including cluster patch deployment, configuration management, monitoring and ing integrations, and system health checks.
- Operational Support: Troubleshoot node failures, service crashes, cluster imbalance, and distributed computing issues.
- Perform root cause analysis for high-severity incidents.
- Ensure high availability and optimal performance of Hadoop platform services.
- Work closely with engineering teams to support consistency in deployment and configuration processes.
Mandatory Skills:
- Hands-on experience in Hadoop engineering and administration.
- Strong proficiency in MAPR and Hortonworks Administration.
- Deep understanding of Hadoop ecosystem components including HDFS, YARN, MapReduce, Hive, HBase, Spark, and Zookeeper.
- Experience with Linux system administration.
- Strong expertise in Shell Scripting and Ansible Automation.
- Experience with patching and security remediation for large-scale distributed systems.
- Understanding of configuration management, service orchestration, and cluster operations.
Preferred Skills:
- Exposure to DevOps tools such as Git.
- Experience with monitoring tools like Grafana, Prometheus, and Ambari.
- Understanding of ITIL processes including Incident, Change, and Problem management.
Qualifications:
- Strong analytical and troubleshooting skills for L3 support.
- Ability to work independently and in cross-functional teams.
- Excellent communication and documentation skills.
- Strong ownership mindset toward reliability and stability of platforms.






