Data Engineer at Talent500 · Bengaluru (Bangalore) · 1 - 10 years · ₹5L - ₹30L / yr · Posted 5 Jan 2023
A proficient, independent contributor that assists in technical design, development, implementation, and support of data pipelines; beginning to invest in less-experienced engineers.
Responsibilities:
- Design, Create and maintain on premise and cloud based data integration pipelines.
- Assemble large, complex data sets that meet functional/non functional business requirements.
- Identify, design, and implement internal process improvements: automating manual processes, optimizing data delivery, re-designing infrastructure for greater scalability, etc.
- Build the infrastructure required for optimal extraction, transformation, and loading of data from a wide variety of data sources.
- Build analytics tools that utilize the data pipeline to provide actionable insights into key business performance metrics.
- Work with stakeholders including the Executive, Product, Data and Design teams to assist with data-related technical issues and support their data infrastructure needs.
- Create data pipelines to enable BI, Analytics and Data Science teams that assist them in building and optimizing their systems
- Assists in the onboarding, training and development of team members.
- Reviews code changes and pull requests for standardization and best practices
- Evolve existing development to be automated, scalable, resilient, self-serve platforms
- Assist the team in the design and requirements gathering for technical and non technical work to drive the direction of projects
Technical & Business Expertise:
-Hands on integration experience in SSIS/Mulesoft
- Hands on experience Azure Synapse
- Proven advanced level of writing database experience in SQL Server
- Proven advanced level of understanding about Data Lake
- Proven intermediate level of writing Python or similar programming language
- Intermediate understanding of Cloud Platforms (GCP)
- Intermediate understanding of Data Warehousing
- Advanced Understanding of Source Control (Github)

Similar jobs (10)
Job Description
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Develop data processing solutions using Python.
- Write complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data ingestion and integration workflows.
- Implement data quality, validation, monitoring, and error-handling processes.
- Develop and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
- Collaborate with data analysts, data scientists, software engineers, and business teams.
- Optimize data pipelines for performance, reliability, and scalability.
- Troubleshoot production data issues and ensure timely resolution.
- Follow best practices for version control, code quality, testing, and deployment.
Mandatory Skills
- Python
- ETL
- SQL
- CI/CD
- DevOps
- Git / Version Control
- Strong problem-solving and debugging skills
Job Summary
Role Overview
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, Advanced SQL, CI/CD, DevOps, and Data Analytics. The ideal candidate should have hands-on experience designing and developing scalable data pipelines, transforming large datasets, and supporting data-driven applications.
Experience with Google Cloud Platform (GCP) will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Develop complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain reliable data integration workflows across multiple data sources.
- Perform data cleansing, validation, transformation, and quality checks.
- Analyze data and provide insights to support business and technical requirements.
- Implement and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps practices and tools to automate deployments, monitoring, and infrastructure processes.
- Troubleshoot data pipeline failures, performance issues, and production incidents.
- Optimize data processing workflows for performance, scalability, and reliability.
- Collaborate with Data Analysts, Data Scientists, Developers, and other stakeholders.
- Follow best practices for version control, testing, documentation, and deployment.
- Contribute to cloud-based data engineering initiatives, preferably on GCP.
Required Skills
- 5–7 years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Strong expertise in Advanced SQL and database concepts.
- Hands-on experience with ETL/ELT processes and data pipelines.
- Good understanding of Data Warehousing and Data Modeling concepts.
- Experience with CI/CD practices and tools.
- Strong understanding of DevOps principles, automation, and deployment processes.
- Strong data analytics and problem-solving skills.
- Experience working with large datasets and performance optimization.
- Good understanding of Git/version control and software development best practices.
Good to Have
- Hands-on experience with Google Cloud Platform (GCP).
- Exposure to GCP data services such as BigQuery, Cloud Storage, Dataflow, Composer, or Pub/Sub.
- Experience with containerization/orchestration technologies such as Docker/Kubernetes.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of cloud-based data architecture and distributed data processing.
Preferred Candidate Profile
- Strong analytical and problem-solving abilities.
- Good communication and stakeholder management skills.
- Ability to work independently as well as in a collaborative team environment.
- Strong ownership of data pipelines and production systems.
- Candidates who can join at short notice are preferred.
Mandatory Skills
Data Engineer, Python , ETL, GCP, Advanced SQL, Strong Data Analytics skills, CICD, Devops
Data Engineer Short Hiring Post
🚨 Hiring: Data Engineer
🔹 Experience: 5–9 Years
🔹 Location: Bangalore / Hyderabad
🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling
🔹 Process: L1 Virtual → L2 F2F Karat Test
🔹 F2F: Bangalore / Hyderabad Location
🔹 Positions: Immediate requirement
⚠️ Note: Candidates must be available for F2F Karat immediately after L1.
#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners
Role Summary
We are hiring a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the Enterprise AI platform:
What You'll Own
- Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON).
- AWS-based data architecture: S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed.
- Data validation frameworks: GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting.
- Agent run logging & observability: designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback).
- AI Factory monitoring dashboards: operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent.
- ML data pipeline support: dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues.
- APIs: designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools.
- Data quality & testing discipline: idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production.
Key Skills — Non-Negotiable (Must-Have, Strong Level)
- Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting.
- SQL — strong hands-on ability, including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries).
- AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale.
- Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions."
- ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built.
- Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV).
Key Skills — Good to Have
- Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable.
- FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs.
- ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself.
- Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats).
- Large-scale metadata querying — experience making file discovery fast across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
About AuxoAI:
AuxoAI is a global platform-based services firm. We help companies—turn their strategies into practical digital and AI solutions. By understanding how our clients make decisions, we use digital and Artificial Intelligence (AI) technologies to drive growth, enhance their operations, improve customer experiences, and provide clear, actionable insights from their data. What We Do We work across various industries such as healthcare, high-tech, consumer packaged goods (CPG), finance etc., and in sales, marketing, and customer support functions.
We help our clients with accelerating their digital and AI journeys through:
• AI Application Development
• Data, Digital and Cloud acceleration using AI
• AI Native Product Engineering
We are seeking a skilled and experienced Data Engineer to join our dynamic team. The ideal candidate will have 6+ years of prior experience in data engineering, with a strong background in AWS (Amazon Web Services) technologies. This role offers an exciting opportunity to work on diverse projects, collaborating with cross-functional teams to design, build, and optimize data pipelines and infrastructure.
Responsibilities:
* Design, develop, and maintain scalable data pipelines and ETL processes leveraging AWS services such as S3, Glue, EMR, Lambda, and Redshift.
* Collaborate with data scientists and analysts to understand data requirements and implement solutions that support analytics and machine learning initiatives.
* Optimize data storage and retrieval mechanisms to ensure performance, reliability, and cost-effectiveness.
* Implement data governance and security best practices to ensure compliance and data integrity.
* Troubleshoot and debug data pipeline issues, providing timely resolution and proactive monitoring.
* Stay abreast of emerging technologies and industry trends, recommending innovative solutions to enhance data engineering capabilities.
Requirements :
* Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
* 6+ years of prior experience in data engineering, with a focus on designing and building data pipelines.
* Proficiency in AWS services, particularly S3, Glue, EMR, Lambda, and Redshift.
* Strong programming skills in languages such as Python, Java, or Scala.
* Experience with SQL and NoSQL databases, data warehousing concepts, and big data technologies.
* Familiarity with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (e.g., Apache Airflow) is a plus.
Position title: Business Intelligence and Data Analyst
Summary -
• Operate and continuously improve the CARS/BI backend within the global Enterprise Data and Business Intelligence platform, ensuring system stability, continuous data loading, and Warehouse the timely delivery of business-critical reports.
• Own the underlying data foundation — SQL pipelines, relational data models, and data warehouse structures — that internal reporting teams and global Data Management initiatives depend on, and act as 1st and 2nd level support for internal customers such as SCM and Sales.
• The role is needed to secure daily BI availability in the Asian time zone, reduce the risk of late or incorrect reporting, and add dedicated database and ETL capacity to the global BI team, including support to the Data Centre team on infrastructure-related issues.
Responsibilities:
• Operate, monitor, and maintain the CARS/BI data integration and data processing workflows, ensuring continuous data loading and stable daily business operations.
• Provide 1st and 2nd level support for CARS/BI, including user support, incident handling, and root-cause analysis of data issues.
• Develop, maintain, and optimise SQL-based data pipelines and transformations within the Enterprise Data Warehouse.
• Design and maintain relational data models and data warehouse structures for CARS BI.
• Integrate data from ERP and other enterprise systems into the BI backend.
• Ensure data consistency, integrity, and performance at the database level, including query tuning and database efficiency improvements.
• Administer the BI technical environment and contribute to its continuous enhancement as part of a global team.
• Support reporting teams by providing structured, reliable datasets and safeguarding the timely delivery of business-critical reports.
• Contribute to global Data Management initiatives, with a focus on Master Data Management and the development of a Common Data Model within the Business Integration Platform.
• Collaborate with the Data Centre team to resolve infrastructure-related issues affecting BI availability.
Education and Experience:
Bachelor's degree in Engineering (BE/B.Tech) or equivalent in Computer Science, Information Technology, or a comparable technical field.
Minimum 3 years of relevant professional experience in Business Intelligence, data warehousing, or database development, including hands-on SQL and ETL work in an enterprise environment.
Experience supporting business users in a global or multi-time-zone IT organisation is an advantage.
Competencies:
• Knowledge about either BI platforms (e.g. IBM Cognos BI suite) or ETL tools (e.g. Informatica PowerCenter)
• Database, data modeling and SQL (preferred Oracle PLSQL)
• Good analytical and communication skills
• Teamplayer
• English
• Strong hands-on experience with SQL (advanced level)
• Solid understanding of relational databases and data warehouse concepts
• Experience with ETL tools and database technologies (e.g., Oracle, SQL Server, SAP BW,Informatica or similar)
• Performance tuning and query optimization skills
• Structured, detail-oriented, and quality-focused working style
• Good understanding of enterprise data flows (SAP/CARS is a plus)
Key Interfaces and Stakeholders:
Only internal customers with different topics: e.g. SCM, Sales etc.
Geography to cover and Travel requirements:
Asian Time Zone, Sometimes travel is required
Behavioral Characteristics
Reliability and accountability — the role safeguards daily reporting availability, so dependable ownership of monitoring and issue follow-up is expected. Structured, analytical problem solving with the patience for thorough root-cause analysis. Service orientation and clear communication towards internal customers such as SCM and Sales. Initiative taking and self-reliance, given largely independent work in the Asian time zone. Cooperation and team spirit within a globally distributed BI team across cultures and time zones. Flexibility and resilience under time pressure, including occasional off-hours support during critical data loads. Integrity and discretion when handling confidential business data. Quality focus and attention to detail, with a continuous improvement mindset.
Interview process
2 rounds - Virtual interview and 1 round Face to Face
Any other Criteria
- Notice Period: Below 60 or 90 Days . No Negotiation on Notice Period
- Gender: Female and Male; Female preferred
- Qualification:Bachelor's degree in Engineering (BE/B.Tech) or equivalent in Computer Science, Information Technology,
Roles & Responsibilities
- Design, develop, and deliver scalable end-to-end data pipelines using Azure Data Factory, ensuring robust integration
of enterprise-wide data from diverse sources
• Build and optimize data engineering workflows using Databricks and PySpark
• Write efficient, high-performance SQL for data transformation and analysis
• Work with the Azure Cloud platform and associated services, applying strong understanding of data warehousing,
data models, and pipelines
• Provide technical leadership to a team of developers, including code reviews and enforcing best practices across the
development lifecycle
• Oversee CI/CD implementation using Azure DevOps, managing deployments across development, QA, and production
environments with proper change control processes
• Collaborate with cross-functional teams to translate business requirements into scalable data solutions
• Ensure data quality, reliability, and performance across all pipelines and platforms
Ideal Candidate
1Strong Azure Databricks Engineer / Senior Data Engineer Profile
2Mandatory (Experience 1) – Must have minimum 8+ years of overall experience in Data Engineering, Data Development, or related data technology roles, with strong hands-on experience in enterprise data pipeline development.
3Mandatory (Experience 2) – Must have strong hands-on experience with Azure Databricks, including development and optimization of scalable data engineering workflows using Databricks and PySpark.
4Mandatory (Experience 3) – Must have strong hands-on proficiency in PySpark/Python and SQL, with proven experience developing complex data transformations, processing workflows, and performance-optimized queries.
5Mandatory (Experience 4) – Must have hands-on experience with Azure Data Factory (ADF) for designing, developing, and orchestrating end-to-end data pipelines and integrating data from multiple sources.
6Mandatory (Experience 5) – Must have strong experience working on the Azure Cloud platform and associated data services, with solid understanding of data warehousing, data modeling, pipeline architecture, and enterprise data solutions.
7Mandatory (Experience 6) – Must have hands-on experience implementing CI/CD using Azure DevOps, including deployment and release management across development, QA, and production environments.
8Mandatory (Experience 7) – Must have proven technical leadership experience, including code reviews, enforcing development best practices, mentoring developers, and providing technical guidance to a data engineering team.
9Mandatory (Notice Period) – Immediate joiners or candidates who can join within 15 days.
10Mandatory (Note) - The position is open across all Cognizant offices pan India. Candidates must be willing to attend the F2F interview at the nearest Cognizant office location.
At Mitratech, we are a team of technocrats focused on building world-class products that simplify operations in the Legal, Risk, Compliance, and HR functions. We are a close-knit, globally dispersed team that thrives in an ecosystem that supports individual excellence and takes pride in its diverse and inclusive work culture centered around great people practices, learning opportunities, and having fun! Our culture is the ideal blend of entrepreneurial spirit and enterprise investment, enabling the chance to move at a rapid pace with some of the most complex, leading-edge technologies available.
For over 35 years, the experts at Mitratech have been focused on solving the complex needs. Today, we serve 20,000 client companies of all sizes globally, representing 30% of the Fortune 500 and over 500,000 users in over 160 countries.
As we continue to grow, we’re always looking for resourceful, enthusiastic, and fresh perspectives. Join our global team and see what makes Mitratech a truly exceptional place to work!
Job Overview
Principal Data Engineer
About Engineering at Mitratech Legal Solutions
Mitratech's engineering organization is a collaborative and dynamic environment where engineers are empowered to drive technical direction and innovation. Our engineers are passionate about delivering high-quality products and solutions that meet the evolving needs of our customers, and we're committed to fostering a culture of continuous learning and growth.
About the Role
Mitratech is a fast-paced and dynamic environment, and this role requires someone who is adaptable, resilient, and able to thrive in a rapidly changing landscape. If you’re a seasoned engineer with a passion for technical leadership, innovation, and collaboration — including building the data foundations that power trusted reporting and agentic AI-driven products — we’d love to hear from you.
What You Will Do
• Drive technical direction for a significant product domain or platform capability, ensuring alignment with business objectives and customer needs
• Design and maintain data pipelines and reporting models that power trusted business metrics and increasingly feed agentic AI systems (e.g., RAG ingestion, embeddings, vector stores, AI agent workflows)
• Use AI-assisted and agentic engineering tools (e.g., Claude Code, Copilot, Cursor, AI agents) as part of your own workflow, and help other engineers adopt agentic development practices effectively
• Reduce systemic complexity by identifying and leading architectural debt remediation, and developing strategies for ongoing technical debt management
• Partner with Product and Engineering leadership to inform multi-quarter roadmap feasibility, and provide technical guidance and oversight to ensure successful implementation
• Elevate engineering craft across multiple teams through RFCs, mentorship, and knowledge sharing, and develop training programs to improve engineering skills and knowledge
• Represent Mitratech’s technical capabilities externally, including speaking at conferences, contributing to open-source projects, and engaging with industry peers and thought leaders
What We Are Looking For
To be successful in this role, you will need:
• 10+ years of experience in software engineering, with a focus on technical leadership and architecture
• Deep understanding of data engineering principles, including data modeling, data warehousing, reporting, and data governance
• Strong technical expertise in SQL, PostgreSQL, ETL/ELT pipelines, BI tools, and analytics platforms
• Practical experience with AI/LLM-adjacent and agentic AI data work — e.g., RAG ingestion pipelines, embedding generation, vector store management, or building/operating AI agent workflows over data — using AI coding assistants (Claude Code, Copilot, Cursor, or similar) as a regular part of the engineering workflow
• Working knowledge of modern cloud platforms such as AWS
• Experience with BI, reporting, dashboards, and customer-facing analytics
• Experience leading cross-functional initiatives with product, engineering, analytics, and business teams
Nice to Have
• Working knowledge of Ruby on Rails and React
• Experience with a semantic or metrics layer (e.g., dbt Semantic Layer, headless BI)
• Understanding of CI/CD, Git-based workflows, and infrastructure-as-code
The Stack Context
• Modern data stack: Fivetran, Airbyte, dbt, Snowflake, GitHub, Terraform, or similar tools
• Application context (nice to have): Ruby on Rails, React, or similar backend/frontend frameworks
• Data modeling: SQL, analytics models, documentation, testing, naming standards, and version control
• Infrastructure: cloud-based data infrastructure, infrastructure-as-code, CI/CD, monitoring, and cloud storage
• Data workflows: ingestion, transformation, orchestration, reporting, deployment, and change management
• Reporting focus: trusted metrics, scalable reporting models, dashboards, exports, and data quality
• AI surface: data pipelines and quality practices supporting AI/LLM and agentic AI use cases (RAG, embeddings, vector stores, AI agents) alongside traditional BI
Why This Role
This role offers a unique opportunity to drive technical direction and innovation at a rapidly growing company, while also mentoring and coaching engineers to improve their craft. As a Principal Data Engineer at Mitratech, you will have the chance to work on complex and challenging problems spanning trusted reporting and agentic AI systems, collaborate with cross-functional teams, and represent the company's technical capabilities externally. If you're looking for a role that offers a mix of technical leadership, data and reporting depth, agentic AI innovation, and collaboration, this could be the perfect fit for you.
We are an equal-opportunity employer that values diversity at all levels. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity, disability, or veteran status.
Job Summary
We are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and infrastructure. The ideal candidate should have strong expertise in SQL, Python, Linux, and modern data engineering practices to support data integration, transformation, and analytics.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop automation scripts using Python for data processing and workflow optimization.
- Work with Linux environments for deployment, monitoring, and troubleshooting.
- Ensure data quality, integrity, and reliability across data platforms.
- Collaborate with data analysts, software engineers, and business stakeholders to deliver data solutions.
- Monitor, troubleshoot, and optimize data pipelines for performance and scalability.
- Implement best practices for data security, governance, and documentation.
Required Skills
- Strong experience in Data Engineering concepts and ETL/ELT processes.
- Proficiency in SQL, including query optimization and database design.
- Strong programming skills in Python.
- Hands-on experience with Linux commands, shell scripting, and system administration basics.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.
- Familiarity with Git/version control.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms (AWS, Azure, or GCP).
- Knowledge of Apache Spark, Airflow, Kafka, or similar data engineering tools.
- Experience with data warehousing solutions and big data technologies.
- Understanding of CI/CD pipelines and containerization (Docker/Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications in cloud or data engineering are an added advantage.
About Us:
The QX Impact was launched with a mission to make A.I accessible and affordable and deliver AI Products/Solutions at scale for the enterprises by bringing the power of Data, AI, and Engineering to drive digital transformation. We believe without insights; businesses will continue to face challenges to better understand their customers and even lose them. Secondly, without insights businesses won't’ be able to deliver differentiated products/services; and finally, without insights, businesses can’t achieve a new level of “Operational Excellence” is crucial to remain competitive, meeting rising customer expectations, expanding markets, and digitalization.
Job Summary:
We are looking for a Senior Data Engineer who is creative, collaborative, and adaptable to join our agile team of data scientists, engineers, and UX developers. The role focuses on building and maintaining robust data pipelines to support advanced analytics, data science, and BI solutions.
As a Senior Data Engineer, you will work with internal and external data, collaborate with data scientists, and contribute to the design, development, and deployment of innovative solutions.
Key Responsibilities:
- Design, develop, test, and maintain optimal data pipeline and ETL architectures.
- Map out data systems and define/design required integrations, ETL, BI, and AI systems/processes.
- Prepare and optimize data for predictive and prescriptive modeling.
- Collaborate with teams to integrate ERP data into the enterprise data lake, ensuring seamless flow and quality.
- Enhance cloud data infrastructure on AWS or Azure for scalability and performance.
- Utilize big data tools and frameworks to optimize data acquisition and preparation.
- Build architectures to move data to/from data lakes and data warehouses for advanced analytics.
- Develop and curate data models for analytics, dashboards, and reports.
- Conduct code reviews, maintain production-level code, and implement testing approaches.
- Monitor, troubleshoot, and resolve data ingestion workflows to maintain reliability and uptime.
- Drive innovation and implement efficient new approaches to data engineering tasks.
Must-Have Skills:
- Bachelor’s degree in Computer Science, Mathematics, Engineering, or a related field.
- 5+ years of experience working with enterprise data platforms, including building and managing data lakes.
- 3–5 years of experience designing and implementing data warehouse solutions.
- Expertise in SQL, including developing stored procedures (SP) and applying advanced data design concepts.
- Proficiency in Spark (Python/Scala) and Spark Streaming for real-time data pipelines.
- Experience with AWS or Azure services (e.g., AWS Glue, Azure Data Factory, Redshift, Snowflake).
- Familiarity with big data tools such as Apache Kafka, Apache Spark, or Flink.
- Hands-on experience with orchestration tools (e.g., Apache Airflow, Prefect).
- Knowledge of CI/CD processes, version control (e.g., Git, Jenkins), and deployment automation.
- Strong problem-solving, communication, and collaboration skills.
Good-to-Have Skills:
- Experience in integrating ERP data into data lakes.
- Experience with traditional ETL tools (e.g., Talend, Pentaho).
Competencies:
- Tech Savvy - Anticipating and adopting innovations in business-building digital and technology applications.
- Self-Development - Actively seeking new ways to grow and be challenged using both formal and informal development channels.
- Action Oriented - Taking on new opportunities and tough challenges with a sense of urgency, high energy, and enthusiasm.
- Customer Focus - Building strong customer relationships and delivering customer-centric solutions.
- Optimize Work Processes - Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement.
Why Join Us?
- Be part of a collaborative and agile team driving cutting-edge AI and data engineering solutions.
- Work on impactful projects that make a difference across industries.
- Opportunities for professional growth and continuous learning.
- Competitive salary and benefits package.
Application Details
Ready to make an impact? Apply today and become part of the QX Impact team!







