
Closed
Posted
Paid on delivery
I'm looking for an experienced Hadoop Big Data Developer skilled in Scala and PySpark. You will work with: - Raw zone containing nested JSON and XML data. - Managed zone where data needs to be stored in JSON or Parquet format, creating Hive tables per client requirements. Ideal Skills and Experience: - Proficiency in Hadoop ecosystem - Strong expertise in Scala and PySpark - Experience with Hive and data transformation - Ability to work with nested JSON and XML
Project ID: 39689043
16 proposals
Remote project
Active 9 mos ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
16 freelancers are bidding on average ₹569,844 INR for this job

Hello, No up front payment give me after completion of work and satisfaction. you can take my technical interview and then decide I have 9 years of experience in bigdata, hive , pyspark, hdfs, map reduce, Java development. My current role is Senior Java Developer, where my responsibilities include day-to-day Java development. My skills include Core Java 8, Spring Boot 3.0, Microservices, Rest API, JPA, Kafka (Spring Boot), Spring Security, graphql, Junit 5, Mockito, AWS, MySql, Oracle, Git, Angular, Spring Scheduler, Agile, Spring 5.3, and Spring Cloud, kubernativ I have experience working with agile methodology, testing through Junit and Mockito, and QA deployment. I am also familiar with stage deployment, pre-production deployment, and production deployment. In my previous role, I worked with agile methodology and delivered user stories based on story points within a two-week sprint. I participated in daily status scrum calls to discuss work progress and had end-of-sprint retrospectives to evaluate what went well and what didn't. I would greatly appreciate the opportunity to work with you.,
₹112,500 INR in 1 day
4.5
4.5

Hello, I liked your project Big Data Developer with Hadoop, Scala & PySpark, which aligns perfectly with my professional expertise. I bring extensive experience in Java, XML, Big Data Sales, I can ensure timely delivery within 3 days. I'm committed to exceeding your expectations and would welcome the opportunity to discuss your requirements in detail.
₹75,000 INR in 3 days
3.8
3.8

Check my resume and message me currently I have learned Hadoop Scala and other languages message me I will send my new resume
₹112,500 INR in 7 days
0.0
0.0

Hey, I have extensive experience in Hadoop ecosystem development with strong proficiency in both Scala and PySpark, making me confident in delivering your project efficiently and to a high standard. Here’s how I can help: Raw Zone Processing – Parsing and handling nested JSON and XML files with complex structures, ensuring no data loss during ingestion. Data Transformation & Enrichment – Using PySpark/Scala to clean, transform, and standardize data for downstream processing. Managed Zone Creation – Storing data in JSON or Parquet format, ensuring optimal compression and query performance. Hive Table Setup – Creating Hive tables as per client requirements with proper schema design and partitions for faster query execution. Optimized Workflow – Implementing efficient Spark jobs with parallel processing and resource optimization on YARN.
₹122,500 INR in 7 days
0.0
0.0

I have strong expertise in the Hadoop ecosystem with hands-on experience in Scala, PySpark, and Hive. I have worked extensively on transforming nested JSON and XML data in the raw zone into optimized formats like Parquet and JSON for the managed zone, while creating Hive tables per client requirements. My approach focuses on efficient schema design, performance optimization, and maintaining data quality. I can deliver clean, scalable, and well-documented pipelines tailored to your needs portfolio.
₹112,500 INR in 7 days
0.0
0.0

Hi There, I’m genuinely excited about this project! Back in my college days, Hadoop and Machine Learning were my strong suit, and since then, I’ve built on that foundation with Python and databases like Oracle. I have hands-on experience setting up Hadoop systems, distributing large datasets across multiple nodes to optimize fetch speed and ensure high data availability — I understand the scale and complexity involved in projects like this. I’ve also spent years creating APIs and working extensively with JSON, including complex, nested structures. I’m comfortable handling raw data in formats like JSON and XML, transforming it into client-ready outputs in Parquet or Hive tables, and ensuring the process is both efficient and scalable. With strong expertise in both Scala and PySpark, I can bridge raw and managed zones smoothly, aligning with your exact requirements.
₹130,000 INR in 6 days
0.0
0.0

As a seasoned full-stack and back-end engineer, I possess strong skills in handling big data and have vast experience working with Scala and PySpark in the Hadoop ecosystem. My expertise extends to effectively working with nested JSON and XML files, exactly what your project requires. In fact, I have previously managed raw zones containing similar data structures and transformed them into desired formats like JSON or Parquet, creating Hive tables as demanded by clients. Moreover, my ability to deliver cloud-native solutions across AWS, GCP, and Azure aligns perfectly with your requirement for storage of the managed zone's data. Having optimized APIs to handle over 100k daily users showcases my understanding of creating scalable systems without compromising performance. With me, you can be sure of clean architecture that not only works well today but is built to last. Last but not least, I bring transparent communication, offering frequent demos to maintain constant project alignment with clients. I understand your need for an efficient big-data developer aware of the best practices that ensure lower latency and a smoother workflow. So if you want a developer who is not just technically astute but also values collaboration and has an eye for detail to ensure bulletproof solutions - let's discuss!
₹112,500 INR in 5 days
0.0
0.0

As a Hadoop enthusiast and a proficient Scala and PySpark developer, I am the ideal candidate to tackle your big data challenge. My name is Nabil Anzum. While my professional journey started with data entry analyses, it quickly evolved into a passion for data management and analytics. Through my priceless experience in fields ranging from Java coding to Cisco network, I have developed an intrinsic understanding of data pipelines that will be essential for this project. My primary expertise lies in manipulating large-scale datasets in Hadoop using Scala and PySpark. I boast a robust understanding of the entire Hadoop ecosystem, which includes Hive - an integral component for managing data as per your client's requirements. Moreover, my skills extend to converting nested JSON and XML structures into more streamlined formats like JSON and Parquet, which aligns perfectly with your project needs in the managed zone. It's important that you place your data in the hands of someone you can trust, and I assure you that your project will receive absolute dedication if given to me. I'm confident my experience, skills, and passion for big data make me the best fit for your needs. Let's discuss your project further and get started!
₹112,500 INR in 7 days
0.0
0.0

I am well versed with Data ingestion, transformation and enrichment Of big data
₹112,500 INR in 15 days
0.0
0.0

I am interested in your project, and since I have experience in this field, I would like to suggest a few ideas based on how I envision the reqquirements. I am taking little more time to make sure there is no issue with the data in futre. Based on the details, we can start by testing the project manually and then proceed to automate it. using Slurm, we can schedule it to run automatically as per our requirements. I can also perform performance tuning to reduce storage usage and improve run time. Stages, 1. Will do the KPI mapping. 2. Data storage 3. Creating staging tables. 4. Final output.
₹112,500 INR in 10 days
0.0
0.0

Hi, I’m a Big Data Engineer with 7 years of experience delivering end-to-end solutions using AWS, Snowflake, Spark, Hadoop, Hive, Bigdata and DBT. I specialize in scalable ingestion pipelines, high-performance ETL/ELT, and data governance—ensuring your data is fast, reliable, and business-ready. I can quickly understand your requirements and deliver efficient, scalable solutions that meet deadlines and exceed expectations. Let’s connect and get started!
₹90,000 INR in 2 days
0.0
0.0

I’m an experienced Senior Data Engineer with 6.5+ years in designing and implementing enterprise-grade Big Data & Cloud solutions, and I believe my skill set aligns perfectly with your requirement for Azure cloud migration, data engineering, and large-scale processing. Here’s how I can contribute to your project: • Cloud Migration Expertise: Successfully led multiple migrations from SQL Server/Talend to Snowflake and AWS/Azure, handling CDC pipelines, SCD2 implementations, and performance optimization. • Big Data & ETL Skills: Proficient in PySpark, AWS Glue, Azure Data Factory, SSIS, and data quality frameworks for extracting, cleansing, and transforming massive datasets. • Programming & Data Modeling: Skilled in Python, SQL, JSON, XML, with experience in Java for integration and API-based data ingestion.
₹75,000 INR in 12 days
0.0
0.0

Big Data Developer with Hadoop, Scala & PySpark 3-Phase Milestone Plan Milestone 1: Project Kickoff, Agreement & Planning This phase includes initial project kickoff, legal agreement signing, and a detailed discovery and planning session. We will finalize project requirements, technical specifications, and a detailed project roadmap to ensure alignment with your goals. Milestone 2: Data Ingestion & Transformation We will set up the data pipeline to ingest nested JSON and XML data from the raw zone. The data will be transformed and structured for storage in the managed zone. We will use Scala and PySpark for efficient data processing and transformation. Milestone 3: Data Storage & Post-Launch Support In this final phase, we will store the processed data in the managed zone in either JSON or Parquet format. We will create and optimize Hive tables as per your requirements for easy access and analysis. This milestone concludes with a 7-day post-launch support period to ensure a smooth transition and operational stability. We are committed to building a complete, robust solution that meets all your needs from the start. We do not use pre-existing templates or third-party source code for the core product. Upon project completion, you will have exclusive ownership of the entire source code. Our team ensures confidentiality and professionalism in all client engagements. Warm regards, Associative Pune
₹7,500,000 INR in 180 days
0.0
0.0

Currently I’m working on the same skills and I do have some bandwidth to deliver this project with in the expected time.
₹112,500 INR in 10 days
0.0
0.0

Hello, I am an experienced Big Data Engineer with strong expertise in the Hadoop ecosystem, Scala, and PySpark, and I can help you design and implement scalable data pipelines. ✅ What I bring to your project: Processing nested JSON and XML data efficiently in the raw zone Transforming and storing data in Parquet/JSON formats as per requirements Creating and managing Hive tables for structured and optimized queries Performance tuning of PySpark/Scala jobs for faster data processing End-to-end pipeline development from raw to managed zones I have successfully worked on enterprise-level data engineering projects involving Hadoop, Spark, and Hive, and I can deliver optimized, well-documented, and reliable data workflows for your use case. Let’s connect to discuss your specific client requirements, and I can get started right away. Regards, Sumana
₹112,500 INR in 25 days
0.0
0.0

Hyderabad, India
Member since Aug 11, 2025
$250-750 USD
₹37500-75000 INR
$30-250 USD
₹12500-37500 INR
$10-300 USD
₹600-1500 INR
$2-8 USD / hour
₹12500-37500 INR
$15-25 USD / hour
$1500-3000 USD
$25-50 USD / hour
$250-750 USD
₹600-1500 INR
$750-1500 USD
₹5000-10000 INR
$10-30 USD
₹12500-37500 INR
min $50 USD / hour
€30-250 EUR
₹1500-12500 INR