
Closed
Posted
I need every article on my Arabic-language website copied out of its static HTML pages and delivered as clean, UTF-8 plain-text files. The site has no dynamic elements or log-in walls, so a straightforward scraper (Python + BeautifulSoup, wget, or any similar tool you prefer) should be enough. Please remove all HTML, inline styles, menus, and ads; keep only the article title, body, sub-headings, and any author/date line that appears inside the article itself. Each piece should be saved as its own .txt file, named after the URL slug or the article title (whichever is easier to automate). Diacritics and right-to-left order must stay intact. Deliverables: • A zipped folder of the plain-text files, one per article • The script or command line you used so I can rerun the extraction later I’ll test a random sample against the live site to confirm nothing was missed and the encoding is correct. Let me know your timeframe and any questions you have about site access.
Project ID: 40523140
5 proposals
Remote project
Active 21 secs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs