
Closed
Posted
I want to turn the text already stored in my database into reliable predictions that can drive decisions. The job starts with pulling the raw records straight from the tables, cleaning and preparing the text, and then building a machine-learning pipeline that outputs clear, measurable forecasts. Python with libraries such as pandas, scikit-learn, spaCy or transformers is ideal, but I am open to other proven stacks if they suit the data volume. What matters most is accuracy and reproducibility. Every step—from extraction and tokenisation to model training and evaluation—needs to be scripted so I can rerun the process as new data lands in the database. A lightweight inference endpoint or CLI that lets me feed fresh text and receive the prediction scores in real time will wrap up the project nicely. Deliverables 1. Well-commented code or notebooks that connect to the database, preprocess the text and train the model 2. The trained model artefact plus a simple interface (REST, CLI or batch script) for inference 3. A concise README explaining setup, hyper-parameters and how to retrain with new data Acceptance criteria: the pipeline runs end-to-end on my machine, meets the agreed-upon accuracy benchmark on a hold-out set, and the inference script returns predictions within a reasonable latency.
Project ID: 40550661
38 proposals
Remote project
Active 57 yrs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs