
Closed
Posted
Paid on delivery
I want to build a dataset for training a medical-assistant bot, so I need Image data that appears inside publicly available hospital reports across the web. The task is to locate those reports, extract every embedded image, and hand the files over to me in readable format I’m open to any solid, well-documented approach—focused web crawling, API use, or site-specific scraping—as long as it scales, respects each site’s terms of service, Please tell me which tools you prefer Deliverables 1. Working scraper or repeatable method with a brief setup guide. 2. First sample batch of at least 1,000 JPEG images for validation. 3. Short report describing sources, filtering logic Acceptance criteria • Only Image data is collected; no text extraction needed. • No protected or copyrighted material is included. • Method can be re-run by me without modification. If anything is unclear, let me know so we can sort it out quickly and move forward.
Project ID: 40528554
120 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
120 freelancers are bidding on average $132 USD for this job

Hi, I've built medical data scrapers before — image extraction and dataset preparation for ML training is exactly what we do. I noticed you're collecting hospital report images specifically. We can handle the scraping, image processing, and dataset organization so it's ready for your bot training right away. I have delivered 1500+ web and mobile projects over 14+ years — happy to share relevant examples. Let's discuss your timeline and data volume. Send over the hospital sources you're targeting and we'll scope this properly. Thanks, Hasan
$200 USD in 7 days
8.7
8.7

⭐⭐⭐⭐⭐ Create a Dataset of Images from Public Hospital Reports ❇️ Hi My Friend, I hope you're doing well. I've reviewed your project requirements and noticed you're looking for a dataset of images from hospital reports. Look no further; Zohaib is here to assist you! My team has already completed 50+ similar projects for data extraction. I will locate the reports, extract embedded images, and deliver them in a readable format. I can use efficient web crawling or API methods, ensuring compliance with site terms. ➡️ Why Me? I can easily create your image dataset as I have 5 years of experience in web scraping, data extraction, and automation. My expertise includes using various tools and techniques for efficient data collection. Besides, I have a strong grip on Python, Beautiful Soup, and Selenium, ensuring a thorough and effective approach to your project. ➡️ Let's have a quick chat to discuss your project in detail and let me show you samples of my previous work. Looking forward to discussing with you in chat. ➡️ Skills & Experience: ✅ Web Scraping ✅ Data Extraction ✅ Python Programming ✅ Beautiful Soup ✅ Selenium ✅ API Integration ✅ Data Validation ✅ Image Processing ✅ Automation ✅ Report Generation ✅ Data Filtering ✅ Quality Assurance Waiting for your response! Best Regards, Zohaib
$150 USD in 2 days
8.1
8.1

SURE------I will do it as per the given specification so lets get started and complete it------- I am highly appreciative to work on this project. I am an Innovative PYTHON/Full stack developer having rich experience with so many successful Tasks. I will give you exact accurate budget after the proper detailed discussion . Let’s connect on chat for further discussion and start quickly. Thanks!!
$250 USD in 7 days
8.1
8.1

Hi I hope you are fine and doing great! sure I can make scraper that will collect images for you. let me know more details to proceed
$150 USD in 1 day
7.7
7.7

With my extensive experience spanning over 17 years and a project value exceeding €500K, there's no doubt I could deliver the precise dataset you require. Regarding your project, I propose we leverage Python's powerful scraping libraries like BeautifulSoup and Selenium to extract images from the hospital reports available on the web. These tools, coupled with my deep knowledge of data extraction techniques, will ensure an automated and streamlined process that respects each site's terms of service. One of the differentiating values I bring to the table is direct communication with me, the company owner, throughout the project's lifecycle. This guarantees not only quality work but also utmost confidentiality and reliability. As a seasoned freelancer from India, my working hours are flexible to cater to client needs from different time zones. To summarize, I offer 100% dedication to satisfying your needs by providing a working scraper method (equipped with a concise setup guide), a substantial sample batch of at least 1,000 JPEG images for validation, and a comprehensive report describing sources and filtering logic. Let's work together to build an efficient and compliant medical-assistant bot dataset!
$333 USD in 99 days
7.8
7.8

I can build a repeatable image-collection workflow to gather embedded images from publicly accessible medical reports while keeping source tracking and reuse constraints in mind. Using Python-based crawling and image extraction, I’ll deliver the scraper, setup guide, an initial sample batch, and a source/filtering report for future reruns.
$120 USD in 3 days
7.2
7.2

With a Bachelors in computer science I understand the importance of delivering comprehensive, reliable and easily scalable dataset. At BN-Droids Digital Services, we have successfully built massive databases with over 20 million records from various industries including Healthcare. Our specialized large-scale web scraping and data mining team, including myself, have always maintained high quality standards in every project conducted.
$30 USD in 7 days
7.0
7.0

Have over 18 years of experience in data mining/ Web scrapping/ Scraping Bots/ Chrome/Opera Extensions I have done it all. Tell us your source and we will put it in excel for you, Or we can even give you filtered results as per your requirement, In the format you want. You can also ask for data into a particular format - Excel, Json, Mysql, Databases, XMLs, you name them. Further Can help you with integrating it with ur databases, Can create json outputs. We are not only good with scraping but also with the tools that u may need after that. We can help you build you softwares round the data we have 99% Data Accuracy. We have Duplicate finder. etc., We can help with Statistics on the data We can help with creating Api's front the data We can create Softwares to manage that data We can build Sites round the data
$70 USD in 1 day
6.9
6.9

Hello Medical dataset for AI training. My experience: curated medical images for ML. Can build dataset, clean, label. Fast, accurate. Ready to start. Giáp Văn Hưng
$250 USD in 7 days
6.8
6.8

Hi There, Ready right now I'm ready to Scrape Hospital Medical Report Images. I will show you sample for your satisfaction and project accuracy then we will go to start, so please contact me and share more details thanks. Check My Profile: https://www.freelancer.pk/u/WelcomeClient I would like to work on this project and can complete with 100% accuracy within the time frame. https://www.freelancer.pk/projects/excel/business-profit-loss-reporting-excel/reviews https://www.freelancer.pk/projects/data-entry/copy-listings-from-website-another/reviews Thanks, Umer
$30 USD in 1 day
6.6
6.6

Hi there, I have reviewed the project requirements for scraping hospital medical report images to build a dataset for a medical-assistant bot. I understand the need to locate and extract embedded images from publicly available hospital reports. Let's chat and discuss it further. To handle your project, I will start with identifying relevant websites, utilizing web crawling techniques, and extracting images using Python libraries such as BeautifulSoup and Requests. My approach involves ensuring scalability and adherence to each site's terms of service. The deliverables will include a working scraper or method with a setup guide, a sample batch of 1,000 JPEG images, and a report detailing the image sources and filtering logic. Before signing-off my bid, I would like to ask a question, i.e., what specific file formats should the images be in for optimal use by the medical-assistant bot? Warm Regards, Aneesa.
$100 USD in 1 day
6.6
6.6

Hello There!!! ★★★★ (Build a scalable, repeatable scraper to collect publicly available medical report images safely and efficiently.) ★★★★ I read your project carefully and understand you need a reliable solution to locate publicly available hospital reports, extract only embedded images, and deliver a reusable scraping workflow with proper documentation. The process should be scalable, well-organised, and easy for you to run again. ⚜ Python Web Scraping ⚜ Automated Image Extraction ⚜ API & Web Crawling ⚜ Data Filtering & Validation ⚜ Repeatable Scraping Workflow ⚜ Setup Guide & Documentation ⚜ Sample Dataset Delivery I have experience building custom Python scraping solutions using Scrapy, BeautifulSoup, Requests, and Selenium where needed. I'll create a clean, reusable scraper with clear filtering logic, provide a setup guide, and deliver the requested sample batch while keeping the workflow organised and easy to maintain. I'd be happy to discuss the target sources and the best approach for your dataset. Looking forward to working with you. Warm Regards, Farhin B.
$110 USD in 10 days
6.6
6.6

Hello dear! I’m Md Toriqul Islam, an experienced Python developer and data extraction specialist with 10+ years of expertise, and I’m excited to partner with you. I can dive into your project immediately. I understand you need a scalable solution to collect image data from publicly available hospital reports while ensuring compliance, repeatability, and clean dataset delivery. I have rich experience building automated web crawlers, scraping pipelines, and data collection systems that extract, filter, validate, and organize large image datasets from public sources. I am skilled in Python, Scrapy, Selenium, Playwright, BeautifulSoup, APIs, data processing, and automation. I’m ready to start immediately and would be happy to discuss this project. Looking forward to hearing from you. Best regards, Md Toriqul Islam
$70 USD in 2 days
6.0
6.0

Hello, I can support this project with a clean and maintainable approach. My focus would be backend structure, integrations, and a reliable data flow. I also noticed the listed skills include PHP Python Web Scraping Data Mining Image Processing Web Crawling API Data Collection. I would first check the current setup, then complete the work in a way that is easy to review. To set this up properly: 1. Are there any infrastructure, security, scalability, or maintenance constraints I should follow? 2. Is there any current codebase, admin panel, or documentation I should review first? 3. What integrations should be connected first, and do you have API docs or test access ready? Regards, Houssame
$140 USD in 7 days
6.5
6.5

Hello, Thanks for the clear brief. I understand you need a repeatable scraper pipeline to collect only embedded images from publicly available hospital reports for a medical training dataset, with a first batch of 1,000 JPEGs and no text extraction. Before starting, can you confirm the data sources? Should I include only: 1. Official hospital websites, or 2. Also open-access medical repositories like PubMed Central (OA), WHO, NIH, and government health portals? Proposed Approach I’ll build a scalable and compliant image extraction system using: Python (core scripting) Scrapy / Playwright (web crawling) PyMuPDF (PDF image extraction) BeautifulSoup (HTML parsing) Deduplication via perceptual hashing Workflow Crawl approved medical/report sources Extract only embedded images (HTML + PDF) Convert to clean JPEG format Remove duplicates and non-medical images Deliver structured dataset + source logs Deliverables ✔ Working reusable scraper ✔ Setup guide ✔ First 1,000 cleaned JPEG images ✔ Source + filtering report Everything will be fully reproducible and compliant with public-access data policies. Please confirm the allowed sources so I can proceed immediately.
$140 USD in 7 days
5.8
5.8

Hello, I can help you build a scalable and fully documented image collection pipeline for medical-assistant dataset development while ensuring compliance with licensing requirements and source restrictions. My approach includes: • Identifying approved public sources and open-access repositories containing medical reports and images. • Building a reusable Python-based crawler/downloader to locate and process eligible documents. • Filtering duplicates, corrupted files, and non-image content. • Exporting images in JPEG format with organized folder structures and metadata manifests. • Providing complete setup documentation so the process can be re-run without modification. I have experience developing custom data-processing systems, web crawlers, API integrations, and large-scale automation workflows. The final solution will be designed for reliability, maintainability, and easy future expansion. I would be happy to discuss the expected image volume, approved data sources, and any licensing requirements before getting started. Looking forward to working with you. Best regards, Farhad
$80 USD in 7 days
5.9
5.9

Hello, I can help build a compliant, repeatable pipeline for collecting image data from publicly available hospital/medical reports, with clear filtering and documentation. For this type of dataset, I would focus only on sources that are legally usable: public-domain, open-license, or explicitly reusable reports. I would also avoid protected health information, copyrighted scans, restricted portals, or sites that disallow scraping in their terms/robots.txt. Workflow: Identify approved public/open sources. Crawl or download reports while respecting ToS and rate limits. Extract embedded images only, with no text extraction. Filter duplicates, broken images, logos, icons, and non-medical graphics. Convert valid images into readable JPEG format. Generate a source report with URLs, license notes, filtering logic, and counts. Package a first validation batch of 1,000 images. Deliverables: Working scraper or repeatable extraction method Setup guide First sample batch of 1,000 JPEG images Short report describing sources, licensing assumptions, and filtering rules I can make the pipeline re-runnable without modification and structured so additional approved sources can be added later.
$140 USD in 2 days
5.9
5.9

Hey there, I'm Vishal Maharaj, a seasoned developer with 25 years of experience in PHP, Python, API, Web Scraping, and Web Crawling, based in Perth, Australia. I am passionate about taking on your project. I understand that you need to scrape hospital medical report images for a medical-assistant bot dataset. My approach would involve utilizing web crawling techniques to locate and extract the embedded images from publicly available hospital reports while ensuring compliance with each site's terms of service. Let's discuss further details and kickstart this project. Looking forward to chatting with you. Cheers, Vishal Maharaj
$250 USD in 5 days
5.0
5.0

Hi, I can build a repeatable and well-documented image collection workflow for publicly available hospital reports, focused only on extracting embedded image data and saving it in a readable format such as JPEG. I have experience with Python web scraping, crawling, PDF/image extraction, BeautifulSoup, requests, PyMuPDF/pdfplumber, and automated data pipelines. I can create a respectful scraper that follows site rules, filters sources carefully, avoids protected/private material, and delivers an initial validation batch of 1,000 images with a short report explaining the sources, filtering logic, and setup steps. I will make the workflow clean, reusable, and easy for you to run again without modification.
$80 USD in 2 days
5.0
5.0

Hi there, Employer, Thank you for outlining your project requirements in detail. We at Demivision LLC are excited about the opportunity to collaborate with you on building a high-quality dataset of medical report images for your medical assistant bot. We understand the importance of not only obtaining a substantial volume of relevant images but also ensuring full compliance with copyright restrictions and site terms. Our team has extensive experience in web scraping, data mining, and image processing, particularly within the healthcare and research domains. We are adept at designing scalable, ethical web crawlers using Python (Scrapy, Requests, BeautifulSoup) and leveraging APIs where available. For image extraction, we utilize robust libraries such as Pillow and OpenCV to handle various image formats and ensure the delivered files are in the required JPEG format. To address your needs, we propose a focused approach: - Identifying reputable, publicly accessible hospital report repositories and medical journal sites that permit data collection. - Implementing site-specific scrapers or utilizing available APIs, with logic to filter out any non-image or protected content. - Compiling a repeatable, well-documented scraping method, including a setup guide, so you can independently re-run the process. As deliverables, you’ll receive a working scraper, an initial batch of at least 1,000 JPEG images for your validation, and a concise report outlining data sources and filtering criteria. We will ensure that only permissible, high-quality images are included, with no text extraction or copyright issues. Please let us know if there are specific sites or formats you would like prioritized, or if you have further questions. We look forward to working together to deliver a robust solution for your project.
$140 USD in 5 days
4.6
4.6

HK, Hong Kong
Payment method verified
Member since Jul 2, 2011
$250-750 USD
$30-250 USD
$2-8 USD / hour
$30-250 USD
$30-250 USD
$30-250 USD
$30-250 USD
$10-30 USD
₹600-1500 INR
₹750-1250 INR / hour
$30-250 USD
₹750-1250 INR / hour
$30-250 USD
₹250000-500000 INR
£10-50000 GBP
€8-30 EUR
£750-1500 GBP
$250-750 USD
$5000-10000 USD
£750-1500 GBP
$30-250 USD
£10-15 GBP / hour
$250-750 USD
$15-25 USD / hour
$30-250 USD