
Closed
Posted
Paid on delivery
Key Responsibilities: experience 14+Years ● Design, develop, test, and maintain scalable ETL data pipelines using Python. ● Architect the enterprise solutions with various technologies like Kafka, multi-cloud services, auto-scaling using GKE, Load balancers, APIGEE proxy API management, DBT, using LLMs as needed in the solution, redaction of sensitive information, DLP (Data Loss Prevention) etc. ● Work extensively on Google Cloud Platform (GCP) services such as: ○ Data-flow for real-time and batch data processing ○ Cloud Functions for lightweight serverless compute ○ BigQuery for data warehousing and analytics ○ Cloud Composer for orchestration of data workflows (on Apache Airflow) ○ Google Cloud Storage (GCS) for managing data at scale ○ IAM for access control and security ○ Cloud Run for containerized applications Should have experience in the following areas : ○ API framework: Python FastAPI ○ Processing engine: Apache Spark ○ Messaging and streaming data processing: Kafka ○ Storage: MongoDB, Redis/Big table ○ Orchestration: Airflow ○ Experience in deployments in GKE, Cloud Run. ● Perform data ingestion from various sources and apply transformation and cleansing logic to ensure high-quality data delivery. ● Implement and enforc
Project ID: 40509362
14 proposals
Remote project
Active 6 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
14 freelancers are bidding on average ₹23,668 INR for this job

Your GCP architecture will face bottlenecks if you don't implement proper data partitioning in BigQuery and streaming deduplication in Kafka - I've seen pipelines collapse under 100K events/sec when these aren't designed upfront. Quick questions before I map the architecture: What's your current data volume (events per second at peak)? And are you processing PII that requires DLP integration with your existing compliance framework? Here's how I'd approach this: - PYTHON + FASTAPI: Build async API endpoints with Pydantic validation and structured logging to handle 50K requests/min while maintaining sub-100ms latency - GCP DATAFLOW + APACHE SPARK: Design streaming pipelines with windowing functions and exactly-once semantics to prevent duplicate processing during network failures - KAFKA + BIGQUERY: Implement partitioned topics with Avro schemas and streaming inserts to BigQuery using storage write API for cost optimization - GKE + CLOUD RUN: Set up horizontal pod autoscaling with custom metrics and blue-green deployments to handle traffic spikes without downtime - DBT + AIRFLOW: Build modular transformation layers with incremental models and SLA monitoring to catch data quality issues before they hit production I've architected 4 similar multi-cloud data platforms that process 2M+ events daily. I don't take on projects where data governance isn't clearly defined - let's discuss your DLP requirements and partition strategy in a 20-minute technical call before committing to implementation.
₹22,500 INR in 7 days
5.5
5.5

As a seasoned AI developer with 14+ years of experience, I have a deep understanding of not only Python, but also an array of other technologies mentioned in the project description such as Apache Spark, Kafka, MongoDB and more. My aptitude includes implementing efficient ETL pipelines using Python that are not just scalable but also high-performing, polishing data to assure top quality. My background also entails working extensively with Google Cloud Platform (GCP), handling data using services like Data-flow, BigQuery, Cloud Storage and IAM - exactly what your project needs. My team and I are no strangers to architecting enterprise solutions as well. We have worked with multi-cloud services, auto-scaling, API proxies (like APIGEE), DBT and more, ensuring that sensitive data is handled correctly following practices like redaction and DLP(Data Loss Prevention). Furthermore, we have expertise in containerizing applications using Cloud Run and deploying on GKE or other platforms in GCP. Our proficiency with Apache Airflow makes us more than capable of orchestrating complex data workflows which is another crucial aspect highlighted in your project description.
₹25,000 INR in 7 days
3.8
3.8

Hello, how are you doing? I have considerable experience building scalable ETL pipelines with Python and deploying on GCP, including Dataflow, BigQuery, Cloud Composer, GCS, Cloud Run, and GKE, plus Kafka streaming, Airflow orchestration, and API frameworks like FastAPI. I’ve worked with Spark-based processing, MongoDB, Redis, and managing secure data through IAM and DLP approaches. I can discuss ingestion, transformation, and data quality steps clearly and keep deployments efficient. Let me know further information, if interested.
₹37,500 INR in 5 days
3.4
3.4

Hi, there. I have 14+ years of experience in data engineering and cloud architecture, with strong hands-on expertise in Python, GCP, Kafka, Spark, Airflow, FastAPI, and BigQuery. I have designed and maintained 20+ enterprise-grade ETL pipelines, processed datasets exceeding 5TB daily, and built scalable solutions using GKE, Cloud Run, Dataflow, Cloud Composer, and GCS. My recent work also includes API management, data governance, DLP implementation, and LLM-powered data workflows. For your project, I will start by reviewing the existing architecture, data sources, and processing requirements. I will design robust ETL pipelines with proper ingestion, transformation, validation, and monitoring layers. The solution will leverage Dataflow, Kafka, Spark, BigQuery, and Airflow for reliable batch and real-time processing while ensuring security through IAM controls, sensitive data redaction, and governance best practices. I follow a structured engineering approach focused on scalability, observability, performance, and maintainability. Every component will be designed for high availability, clean deployment, and smooth integration across cloud services. My experience with multi-cloud environments, GKE deployments, API frameworks, and enterprise data platforms enables me to deliver a dependable solution aligned with business objectives. Thanks.
₹12,500 INR in 3 days
3.1
3.1

With a solid 9+ years of experience in IT services, I'm no stranger to the complexities of designing, developing, and testing data pipelines using Python- an essential skill for this GCP Data Engineer role. I've built applications and managed cloud services, including those within the GCP ecosystem, such as Dataflow, GCS, BigQuery, and IAM. My expertise extends to using these technologies for real-time and batch data processing, enabling lightweight serverless compute using Cloud Functions and managing data at scale with Google Cloud Storage. The technologies you've outlined aren't just familiar to me but are also ones I've successfully deployed in previous projects. Be it Kafka for streaming data processing or Airflow as an orchestration tool, I've worked with them all. Additionally, my diverse skill set extends to API frameworking (Python FastAPI), processing engine (Apache Spark), different storages like MongoDB and Redis/Big table and more - making me an ideal candidate for this project. I have been honing my programming skills for over a decade in languages such as Java, PHP along with python I understand that your business aims to ensure high-quality data delivery; this is something I prioritize too— implenting trandformation and cleansing logics is second nature to me.I pride myself on efficient problem-solving skills and a proactive mindset, ensuring that I add value at every stage of development. Consider my deep expertise matched with
₹25,000 INR in 7 days
2.0
2.0

Hi, GCP data engineering is my core - I'm currently building a standardized container framework on GCP (GKE) with Terraform IaC and CI/CD, so your stack maps directly to what I do. For this role: - ETL in Python: clean, testable, idempotent jobs with retries, schema validation and observability. - GKE: containerized, auto-scaling pipeline workloads with Helm; load balancing and resource tuning. - Streaming + transform: Kafka ingestion, DBT for the transform layer, orchestration via Airflow/Cloud Composer. - Multi-cloud where needed, APIGEE for API management, and LLMs slotted in where they add real value. - IaC + CI/CD so the whole platform is reproducible and auditable. I have ~20 years across infrastructure and data/software (Python, GCP, Kubernetes, Terraform) - I can architect, build and mentor. Two questions: (1) Batch ETL only, or real-time streaming from day one? (2) Greenfield, or extending an existing GCP/GKE setup? Any compliance regime to design around? Profile: https://www.freelancer.com/u/mengw3. I can start this week.
₹12,500 INR in 30 days
0.7
0.7

Hello, I am a principal data architect with 10+ years of experience designing scalable ETL pipelines and enterprise cloud solutions. I specialize in Python and the Google Cloud Platform (GCP), making me an excellent fit to engineer your data workflows. I have extensive experience orchestrating complex data pipelines using Cloud Composer (Airflow), BigQuery, and Dataflow. I regularly architect auto-scaling applications on GKE and Cloud Run using Python FastAPI and Apache Spark. My background includes streaming data processing with Kafka, and integrating LLMs while strictly enforcing Data Loss Prevention (DLP) for sensitive information—skills I actively leverage when building secure, AI-driven platforms like KIMB – The AI School. Why hire me? I bridge the gap between heavy data processing and secure API management. I don’t just write ETL scripts; I design fault-tolerant, high-performance architectures ensuring clean, reliable data delivery across MongoDB, Redis, and BigQuery. Deliverables: - Scalable Python ETL pipelines & Apache Spark processing - GCP orchestration (BigQuery, Dataflow, Cloud Composer, GKE) - FastAPI management, Kafka streaming & NoSQL storage integration - LLM integration with strict IAM access control and DLP enforcement Let’s connect to discuss your infrastructure.
₹35,000 INR in 7 days
0.0
0.0

Rahul here, I understand you're looking for a senior data engineering and cloud architecture professional with deep experience in designing scalable ETL pipelines, enterprise data platforms, and cloud-native solutions on GCP. With extensive experience in Python, FastAPI, data engineering, cloud architectures, API integrations, distributed systems, and enterprise application development, I have worked on building scalable data processing solutions, workflow orchestration, cloud deployments, and high-performance backend systems. My approach is to architect secure, scalable, and maintainable solutions leveraging GCP services, Kafka-based streaming, Spark processing, Airflow orchestration, FastAPI services, and modern data governance practices including DLP, access control, monitoring, and performance optimization. I'd be happy to discuss your current architecture, data volume, integration requirements, and project goals in more detail. Ready to start immediately. Thank you for your consideration.
₹21,500 INR in 9 days
0.0
0.0

Hi, I can fix your Lead GCP Data engineer I've solved this exact problem many times. Here is what I will do: Design scalable Apache Spark ETL pipelines on Google Cloud Platform for batch and real-time data. Build Airflow orchestration with Python, FastAPI, and secure GCP services for reliable delivery. Implement Kafka, BigQuery, GCS, Cloud Run, and GKE workflows with cleansing and data quality checks. 10 days free support after delivery Milestone-based payment Reply "YES" and I will share a similar sample within 1 hour. Best regards, Ribal Ali - write short and to the point do not make it longer
₹37,500 INR in 5 days
0.0
0.0

This aligns perfectly with my skill set. I understand the need for scalable ETL data pipelines, automated solutions, and seamless data processing in GCP. While I am new to freelancer, I have tons of experience and have done other projects off-site. I specialize in designing and maintaining robust ETL pipelines, implementing automation with Python, and integrating various GCP services effectively. I would love to chat more about your project! Regards, Warrick Van Eeden
₹16,900 INR in 7 days
0.0
0.0

Bengaluru, India
Member since May 21, 2026
₹750-1250 INR / hour
$30-250 USD
$250-750 USD
₹12500-37500 INR
min $50 USD / hour
$250-750 USD
₹600-1500 INR
$8-15 AUD / hour
₹1500-12500 INR
₹75000-150000 INR
₹1500-12500 INR
₹37500-75000 INR
₹1500-12500 INR
₹1000000-2500000 INR
$8-15 USD / hour
₹12500-37500 INR
₹12500-37500 INR
$250-750 USD
₹1500-12500 INR