Web Scraper at B2B SaaS platform For BFSI · Mumbai · 1 - 5 years · ₹10L - ₹11L / yr · Posted 2 Mar 2022

Our client focuses on providing solutions in terms of data, analytics, decisioning and automation. They focus on providing solutions to the lending lifecycle of financial institutions and their products are designed to focus on systemic fraud prevention, risk management, compliance etc.
Our client is a one stop solution provider, catering to the authentication, verification and diligence needs of various industries including but not limited to, banking, insurance, payments etc.
Headquartered in Mumbai, our client was founded in 2015 by a team of three veteran entrepreneurs, two of whom are chartered accountants and one is a graduate from IIT, Kharagpur. They have been funded by tier 1 investors and have raised $1.1M in funding.
What you will do:
- Developing a deep understanding of our vast data sources on the web and knowing exactly how, when, and which data to scrap, parse and store
- Working closely with Database Administrators to store data in SQL and NoSQL databases
- Developing frameworks for automating and maintaining constant flow of data from multiple sources
- Working independently with little supervision to research and test innovative solutions skills
Desired Candidate Profile
What you need to have:- Bachelor/ Master’s degree in Computer science/ Computer Engineering/ Information Technology
- 1 - 5 years of relevant experience
- Strong coding experience in Python (knowledge of Java, JavaScript is a plus)
- Experience with SQL and NoSQL databases
- Experience with multi-processing, multi-threading, and AWS/Azure
- Strong knowledge of scraping frameworks such as Python (Request, Beautiful Soup), Web Harvest and others
- In depth knowledge of algorithms and data structures & previous experience with web crawling is a must

Similar jobs (10)
Must-Have Skills
- Minimum 3 years of experience in Data Engineering / Analytics Engineering / Fintech Data roles
- Must have worked on SMS Parsing, intelligent platform, converting RAW customer SMS data into structured actionable financial signals and enabling downstream usage of SMS derived variables
- Must have established a continuous learning cycle to expand parser coverage
- Experience in Lending / NBFC / Fintech domain
- Experience working with Bureau, SMS, Device, or Banking data
- Strong Python and SQL (production level)
- Experience handling unstructured data (SMS, logs, JSON, APIs)
- Experience building data pipelines, schedulers, and cron jobs
- Strong database design and data modelling skills
- Ability to work in a startup environment with high ownership
- Familiarity with modern platforms like AWS, Snowflake, Google BigQuery, Redshift
Good to Have
- Experience in STPL, especially less than 25K ticket size
- Experience with streaming (Kafka/Kinesis) and orchestration (Airflow or Step Functions)
- Experience with feature stores and risk analytics datasets
- Knowledge of regex, NLP basics for SMS parsing
- Experience supporting real-time decision engines/underwriting systems
Role Summary
This role will be responsible for owning the end-to-end data-structuring layer across the organisation. The individual will transform large volumes of raw, unstructured, and semi-structured data (such as SMS, device, bureau, and app data) into clean, standardised, and analysis-ready datasets. These structured datasets will directly power risk analytics, fraud detection, marketing insights, collections strategy, and policy decisioning.
Key Objective of the Role
Ensure all raw lending data (SMS, Bureau, Device, AA, App logs) is captured, parsed, structured, and stored in a clean analytics-ready format inside databases (PostgreSQL, DynamoDB, AWS stack) so that the Risk and Data Science team can directly use it for feature creation, policy building, and portfolio monitoring.
Core Responsibilities
- End-to-End Data Ownership
- Design, build, and maintain end-to-end data pipelines (batch + streaming) using AWS native services (Glue, Lambda, Step Functions, Kinesis, S3, Athena, Redshift, EMR/Spark, etc.): ingestion
→ parsing → structuring → storage
- Work closely with Tech, Product, and Data Science to define what data should be captured
- Maintain data documentation, data dictionaries, and schema governance
- Ensure data quality, consistency, and version control
- Unstructured Data Processing (Highest Priority)
- Parse raw SMS dumps and categorise into salary, EMI, loan apps, collections, credits, debits, OTP, etc.
- Process device fingerprint, behavioural logs, and vendor data (FinBox, AA, Bureau APIs)
- Convert JSON, logs, and raw API responses into structured feature tables
- Build regex/keyword-based parsers for financial SMS classification
- Feature Implementation (From Risk & Data Science Team)
- Implement feature creation logic provided by Risk/Data Science team
- Translate business and policy logic into SQL/Python pipelines
- Create reusable feature layers for underwriting, fraud, collections, and monitoring
- Maintain a feature store for consistent model and policy usage
- Lending Data Understanding (Domain-Specific Requirement)
- Work with Bureau data
- Structure SMS-derived financial variables (income, stress, EMI signals)
- Work with Account Aggregator and bank transaction datasets
- Understand fintech alternate data used in underwriting and fraud detection
- Data Pipelines & Automation
- Build and maintain ETL/ELT pipelines using Python & SQL
- Create cron jobs for automated data ingestion and feature refresh
- Automate vendor data pulls (Bureau, SMS SDK, AA, device data)
- Ensure low-latency pipelines for real-time underwriting use cases
- Database Structuring & Storage Architecture
- Structure clean datasets in PostgreSQL (analytics layer)
- Manage raw data storage in DynamoDB / S3 data lake
- Design normalized and denormalised tables for risk analytics
- Optimise database performance for large-scale query workloads
- Dashboards & Readable Data Layer
- Create analytics-ready datasets, implement & write Metabase queries and convert into dashboards (Metabase / Power BI)
- Enable self-serve data access for Risk, Business, and Founders
- Support ad-hoc analysis requirements from leadership
- Cross-Functional Collaboration (Very Important)
- The role requires close collaboration with data science, tech, product, and business teams to ensure reliable data pipelines, well-defined schemas, API integrations, logging architecture and high data quality, enabling faster and more accurate decision-making across lending workflows.
Tech Stack (Current Environment)
- AWS Services
- PostgreSQL (Primary analytics DB)
- DynamoDB (Raw/NoSQL storage)
- Python (Pandas, NumPy, ETL frameworks)
- Advanced SQL
- APIs, JSON, and Log Data Handling
About Us
Corporate Web Solutions works on technology-driven digital solutions involving data, automation, web technologies, and artificial intelligence. Our internship programs focus on practical learning and real-world project exposure.
Role Overview
As a Data Science Intern, you'll work with datasets to perform analysis, visualization, and machine learning tasks while learning modern AI-assisted workflows.
Key Responsibilities
- Collect, clean, and analyze datasets.
- Perform exploratory data analysis.
- Create data visualizations and reports.
- Assist in developing machine learning models.
- Work with Python-based data science tools.
- Explore AI tools for data analysis and productivity.
Requirements
- Basic knowledge of Python.
- Understanding of data analysis fundamentals.
- Familiarity with Pandas and NumPy is a plus.
- Basic understanding of statistics.
- Analytical and problem-solving skills.
Perks
- Certificate of Internship
- Flexible work hours
- Mentorship and real project exposure
- Potential for PPO
- Letter of Recommendation
- Performance-Based Stipend available up to ₹18,000/month
About the Role
We are seeking motivated Data Engineering Interns to join our team remotely for a 3-month internship. This role is designed for students or recent graduates interested in working with data pipelines, ETL processes, and big data tools. You will gain practical experience in building scalable data solutions. While this is an unpaid internship, interns who successfully complete the program will receive a Completion Certificate and a Letter of Recommendation.
Responsibilities
- Assist in designing and building data pipelines for structured and unstructured data.
- Support ETL (Extract, Transform, Load) processes to prepare data for analytics.
- Work with databases (SQL/NoSQL) for data storage and retrieval.
- Help optimize data workflows for performance and scalability.
- Collaborate with data scientists and analysts to ensure data quality and consistency.
- Document workflows, schemas, and technical processes.
Requirements
- Strong interest in data engineering, databases, and big data systems.
- Basic knowledge of SQL and relational database concepts.
- Familiarity with Python, Java, or Scala for data processing.
- Understanding of ETL concepts and data pipelines.
- Exposure to cloud platforms (AWS, Azure, or GCP) is a plus.
- Familiarity with big data frameworks (Hadoop, Spark, Kafka) is an advantage.
- Good problem-solving skills and ability to work independently in a remote setup.
What You’ll Gain
- Hands-on experience in data engineering and ETL pipelines.
- Exposure to real-world data workflows.
- Mentorship and guidance from experienced engineers.
- Completion Certificate upon successful completion.
- Letter of Recommendation based on performance.
Internship Details
- Duration: 3 months
- Location: Remote (Work from Home)
- Stipend: Unpaid
- Perks: Completion Certificate + Letter of Recommendation
About Nexora Group
Nexora Group is a forward-thinking technology and innovation company focused on leveraging Artificial Intelligence, Data Science, and emerging technologies to solve real-world business challenges. We provide opportunities for aspiring professionals to gain hands-on experience, work on impactful projects, and develop industry-relevant skills in a collaborative environment.
Internship Overview
We are looking for enthusiastic and motivated Data Science with AI Interns to join our growing team. This internship is designed for students and recent graduates who are passionate about data analytics, machine learning, artificial intelligence, and data-driven decision-making.
The selected candidates will work alongside experienced professionals on real-world datasets, AI models, and business intelligence projects while gaining practical exposure to industry-standard tools and technologies.
Key Responsibilities
- Collect, clean, and preprocess structured and unstructured datasets.
- Perform exploratory data analysis (EDA) and generate actionable insights.
- Assist in developing and deploying machine learning and AI models.
- Work with Python, SQL, and data visualization tools.
- Create dashboards, reports, and data-driven presentations.
- Support predictive analytics and model evaluation activities.
- Collaborate with cross-functional teams on AI-driven projects.
- Research emerging trends in Data Science, Machine Learning, and Generative AI.
- Document project findings and maintain technical reports.
Required Skills
- Basic understanding of Data Science and Machine Learning concepts.
- Knowledge of Python and data analysis libraries (Pandas, NumPy, Matplotlib, Scikit-learn).
- Familiarity with SQL and database concepts.
- Understanding of AI, Generative AI, and Large Language Models (LLMs) is a plus.
- Strong analytical and problem-solving skills.
- Good communication and teamwork abilities.
- Eagerness to learn and adapt to new technologies.
Eligibility
- Undergraduate or postgraduate students pursuing Computer Science, Data Science, AI, IT, Statistics, Mathematics, or related fields.
- Recent graduates looking to gain practical industry experience.
- Candidates with personal projects, certifications, or relevant coursework will be preferred.
What You'll Gain
- Hands-on experience with real-world AI and Data Science projects.
- Mentorship from industry professionals.
- Exposure to modern AI tools and technologies.
- Internship Certificate upon successful completion.
- Letter of Recommendation (based on performance).
- Opportunity for a Pre-Placement Offer (PPO) for outstanding performers.
- Professional networking and career development opportunities.
Django + Scraper Developer – Job Description
Position: Django + Scraper Developer
Job Summary
We are looking for a skilled Django + Web Scraping Developer with 3+ years of experience in developing scalable web applications using Django and building robust web scrapers. The ideal candidate should have strong Python knowledge, experience with scraping frameworks and APIs, and the ability to work with databases and third-party integrations.
Key Responsibilities
- Develop, maintain, and enhance web applications using Python and Django.
- Design and implement reliable web scraping solutions for extracting structured data from websites.
- Develop scrapers using tools such as Scrapy, Selenium, BeautifulSoup, Requests, or similar technologies.
- Handle dynamic websites, pagination, authentication, proxies, CAPTCHA-related challenges, and anti-bot mechanisms where legally and technically appropriate.
- Clean, validate, transform, and store scraped data in databases.
- Develop and integrate REST APIs and third-party APIs.
- Work with databases such as MySQL, PostgreSQL, or MongoDB.
- Optimize scraper performance, reliability, and data accuracy.
- Troubleshoot and fix issues related to scraping, Django applications, APIs, and databases.
- Write clean, reusable, and maintainable Python code.
- Collaborate with developers, project managers, and other stakeholders to understand project requirements.
- Perform testing, debugging, and performance optimization.
Required Skills
- 3+ years of professional experience in Python/Django development.
- Strong knowledge of Python and Django.
- Hands-on experience with Web Scraping / Data Extraction.
- Experience with Scrapy, Selenium, BeautifulSoup, Requests, or equivalent tools.
- Good understanding of HTML, CSS, JavaScript, DOM, and HTTP/HTTPS.
- Experience working with REST APIs and JSON.
- Good knowledge of SQL and databases such as MySQL/PostgreSQL.
- Understanding of Git/version control.
- Strong debugging and problem-solving skills.
Good to Have
- Experience with Celery, Redis, Docker, and Linux.
- Knowledge of cloud platforms such as AWS.
- Experience handling large-scale scraping projects.
- Understanding of proxy rotation, browser automation, and anti-bot techniques.
- Experience with data pipelines and ETL processes.
What We Offer
- Opportunity to work on challenging and real-world projects.
- Collaborative and growth-oriented work environment.
- Learning and professional development opportunities.
- Competitive salary based on skills and experience.
How to Apply
Interested candidates can share their updated resume along with their current salary, expected salary, notice period, and total experience.
Data Engineer Short Hiring Post
🚨 Hiring: Data Engineer
🔹 Experience: 5–9 Years
🔹 Location: Bangalore / Hyderabad
🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling
🔹 Process: L1 Virtual → L2 F2F Karat Test
🔹 F2F: Bangalore / Hyderabad Location
🔹 Positions: Immediate requirement
⚠️ Note: Candidates must be available for F2F Karat immediately after L1.
#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners
Hiring for Data Scientist / Senior Data Scientist
Exp : 4 - 12 yrs
Edu : BE/B.tech/MCA
Work Location : Pune
Notice Period : Immediate - 15 days
Skills :
4+ years of experience in data engineering, data science, or related domains.
Hands-on experience with SQL, Python, and distributed data systems.
Knowledge of machine learning techniques and statistical analysis.
Experience with cloud data platforms (Azure Data Factory, AWS Glue, GCP BigQuery).
Familiarity with DevOps practices and CI/CD for data pipelines.
Platforms & Operations Experience (Preferred)
- Experience working with Azure, AWS, or Google Cloud data tools.
Operational experience with data orchestration tools (Airflow, ADF, Glue).
Understanding of Kubernetes, Docker, or containerized environments.
Hands-on experience with data warehousing platforms (Snowflake, Redshift, BigQuery).
Experience in monitoring, logging, and alerting operations for data workflows.
Job Description:
As a Data Science Intern, you will collaborate with our data science and analytics teams to work on meaningful projects involving data analysis, predictive modeling, and statistical modeling. You will have the opportunity to apply your academic knowledge in a practical, fast-paced environment, contribute to key data-driven projects, and gain valuable experience with industry-leading tools and technologies.
Responsibilities:
- Assist in collecting, cleaning, and preprocessing data from various sources.
- Perform exploratory data analysis to identify trends, patterns, and anomalies.
- Develop and implement machine learning models and algorithms.
- Create data visualizations and reports to communicate findings to stakeholders.
- Collaborate with team members on data-driven projects and research.
- Participate in meetings and contribute to discussions on project progress and strategy.
- Work with large datasets to clean, preprocess, and analyze data.
- Build and deploy statistical and machine learning models to generate actionable insights.
- Conduct exploratory data analysis (EDA) to uncover trends, patterns, and correlations.
- Assist in the creation of data visualizations and dashboards for reporting insights.
- Support the development and improvement of data pipelines and algorithms.
- Collaborate with cross-functional teams to understand data needs and translate them into actionable analytics solutions.
- Contribute to the documentation and presentation of results, findings, and recommendations.
- Participate in team meetings, brainstorming sessions, and project discussions.
Duration: 03 Months (with the possibility of extending up to 6 months)
MODE: Work From Home (Online)
Requirements:
- Any Graduate / PassOuts / Freasher can apply.
- Currently pursuing a Bachelor's or Master’s degree in Data Science, Computer Science, Mathematics, Statistics, or a related field.
- Proficiency in programming languages such as Python, R, or SQL.
- Strong foundation in statistics, probability, and data analysis techniques.
Benefits
Internship Certificate
Letter of recommendation
Stipend Performance Based
Part time work from home (2-3 Hrs per day)
5 days a week, Fully Flexible Shift
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Summary
We are seeking a motivated Data Engineer with strong skills in SQL, Python, and Linux to design, build, and maintain scalable data pipelines and support data-driven decision-making. The ideal candidate should have experience working with large datasets, ETL processes, and relational databases while ensuring data quality and performance.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write optimized SQL queries, stored procedures, and database objects.
- Develop Python scripts for data extraction, transformation, and automation.
- Work in Linux environments to manage scripts, cron jobs, and system processes.
- Monitor and troubleshoot data pipeline failures.
- Ensure data integrity, consistency, and quality across systems.
- Collaborate with data analysts, software engineers, and business stakeholders.
- Optimize database performance and query execution.
- Participate in code reviews and follow best engineering practices.
Required Skills
- Strong proficiency in SQL (joins, subqueries, window functions, CTEs, indexing, query optimization).
- Good programming experience in Python.
- Hands-on experience with Linux commands and shell scripting.
- Understanding of ETL/ELT concepts and data warehousing.
- Knowledge of relational databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
- Familiarity with Git for version control.
- Strong problem-solving and analytical skills.
Job Description
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Develop data processing solutions using Python.
- Write complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data ingestion and integration workflows.
- Implement data quality, validation, monitoring, and error-handling processes.
- Develop and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
- Collaborate with data analysts, data scientists, software engineers, and business teams.
- Optimize data pipelines for performance, reliability, and scalability.
- Troubleshoot production data issues and ensure timely resolution.
- Follow best practices for version control, code quality, testing, and deployment.
Mandatory Skills
- Python
- ETL
- SQL
- CI/CD
- DevOps
- Git / Version Control
- Strong problem-solving and debugging skills







