About the Role
We are looking for an analytical and detail-oriented Data Scientist with deep expertise in graph analytics, network analysis, and scalable data engineering. This is not a conventional machine learning role — the focus is on understanding complex relationships, structures, and patterns within large-scale datasets using graph-based methods and statistical modeling. You will work in a collaborative, fast-paced environment and be expected to contribute across the full data lifecycle — from raw data exploration through to production-grade analytical systems.
This is a full-time, in-office role based out of our Pune office (5 days a week).
Key Responsibilities
Graph Analytics & Network Analysis
• Design and implement graph-based models to identify patterns, clusters, communities, and relationships within complex datasets.
• Apply network analysis techniques using metrics such as clustering coefficient, degree assortativity, density, Gini index, and small-world index.
• Perform multi-hop network traversals and community detection using algorithms such as Louvain partitioning and similar graph clustering approaches.
• Build and query graph databases (Neo4j, ArangoDB) using Cypher Query Language to extract structural insights from connected data.
• Leverage GPU-accelerated graph libraries (cuGraph) to scale graph computations across large datasets efficiently.
Data Engineering & Pipeline Development
• Conduct thorough Exploratory Data Analysis (EDA) on large-scale structured and semi-structured datasets to surface quality issues, distributions, and key features.
• Build, optimize, and maintain scalable data pipelines and stored procedures across cloud data platforms such as BigQuery, PostgreSQL, or Hive.
• Automate data workflows using orchestration tools such as Apache Airflow or Kubeflow Pipelines.
• Apply GPU-accelerated computing (CUDA, CuPy, cuDF) to optimize processing performance on high-volume data workloads.
• Ensure data integrity, reproducibility, and documentation across all analytical workflows.
Statistical Modeling & Insight Generation
• Apply statistical modeling techniques to detect behavioral anomalies, trends, and patterns within datasets.
• Use time series analysis to identify temporal patterns and changes in data over time.
• Translate analytical findings into clear, actionable insights for both technical and non-technical stakeholders.
• Build dashboards and reports using visualization tools (e.g., Trino Superset) to communicate results effectively.
Collaboration & Documentation
• Work closely with product, engineering, and business teams to understand requirements and deliver relevant analytical solutions.
• Contribute to internal knowledge sharing through workshops, documentation, and peer reviews.
• Maintain well-documented codebases and analytical frameworks for long-term maintainability.
Required Skills & Qualifications
Must-Have
• 3+ years of experience in a Data Science, Data Analytics, or Graph Analytics role.
• Strong proficiency in Python with hands-on experience using NumPy, Pandas, CuPy, and cuDF.
• Practical experience with graph analytics libraries — NetworkX, cuGraph, Neo4j, or ArangoDB.
• Solid understanding of graph theory concepts: community detection, network metrics, graph traversal, and clustering algorithms.
• Proficiency in SQL; experience with BigQuery, PostgreSQL, MS SQL, Hive, or similar databases.
• Experience with GPU-based computing using CUDA for performance-critical data tasks.
• Strong analytical thinking and ability to work independently on ambiguous, open-ended problems.
Good to Have
• Experience with Cypher Query Language for querying graph databases (Neo4j / ArangoDB).
• Familiarity with graph ML frameworks such as PyTorch Geometric.
• Exposure to workflow orchestration tools — Apache Airflow or Kubeflow Pipelines.
• Knowledge of cloud data tools such as Trino, MinIO, or IBM Datastage.
• Basic scripting skills in Bash or C++ for automation or performance tasks.
• Experience presenting data insights to senior stakeholders or cross-functional teams.
• Research publications, patents, or open-source contributions in graph analytics or data science are a strong plus.
What We Offer
• Opportunity to solve high-impact, large-scale data problems using modern graph and analytics tools.
• Exposure to cutting-edge technologies including GPU-accelerated computing and graph ML.
• Collaborative, intellectually stimulating work environment with a strong engineering culture.
• Competitive compensation with performance-based incentives.
• Learning & development support — certifications, courses, and conference participation.
• Centrally located Pune office with a structured, in-person team culture.