
Open
Posted
•
Ends in 18 hours
Paid on delivery
I am looking for an experienced Python developer to help build a robust research and backtesting system for pre-race horse-racing trading on the exchange. I already have approximately 4 GB of historical Exchange Stream data, including .bz2 market files, together with API documentation and an initial trading hypothesis. The immediate objective is to build a technically accurate system that can: 1. Read and replay historical Exchange Stream data. 2. Reconstruct market and runner states accurately. 3. Generate fixed time-to-off market snapshots. 4. Simulate realistic back and lay orders. 5. Calculate correct gross and net profit after commission. 6. Test trading strategies using chronological out-of-sample data. 7. Produce complete and auditable trade reports. The first version will cover: - Horse racing only. - Great Britain and Ireland initially. - WIN markets only. - Pre-race trading only. - No positions intentionally held in-play. - One strategy position per market. - Current favourite and second favourite analysis. - Historical replay and backtesting before any live API work. Existing data The historical files are Exchange Stream .bz2 files containing data such as: - Market definitions. - Event and market IDs. - Runner IDs and names. - Market status and scheduled start time. - Best available-to-back prices. - Best available-to-lay prices. - Available amounts at the top price levels. - Last traded price. - Traded volume. - Market suspension and in-play status. The files contain incremental Exchange Stream updates, so the developer must understand that they are not ordinary CSV snapshots. A persistent market state must be reconstructed by correctly applying each update in timestamp order. Initial paid technical test The selected developer will first complete a small, fixed-price paid test using one sample historical WIN market file. The test must: 1. Decompress and parse the .bz2 file. 2. Identify: - Market ID. - Event ID. - Venue. - Scheduled market start. - Market type. - Runner names and IDs. 3. Reconstruct the market state by applying Exchange Stream delta messages. 4. Produce snapshots at: - Ten minutes before scheduled start. - Five minutes before scheduled start. - One minute before scheduled start. 5. For each snapshot, output: - Current favourite. - Favourite back price. - Favourite lay price. - Available amounts at the best prices. - Last traded price. - Runner traded volume, where available. - Total market traded volume. - Number of active runners. 6. Identify market suspension and in-play timestamps. 7. Explain how missing fields, zero-size ladder updates and unchanged values have been handled. 8. Include automated tests. 9. Commit the code and documentation to a private GitHub repository owned by me. Successful completion of this paid test may lead to the full project.
Project ID: 40678765
167 proposals
Open for bidding
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
167 freelancers are bidding on average £173 GBP for this job

⭐⭐⭐⭐⭐ Build a Robust Backtesting System for Horse Racing Trading ❇️ Hi My Friend, I hope you're doing well. I've reviewed your project needs and see you're looking for an experienced Python developer for your horse racing backtesting system. You don’t need to look any further; Zohaib is here to help you! My team has successfully completed over 50 similar projects in trading systems. I will build a system that accurately reads historical data, reconstructs market states, and produces reliable reports—all within your budget. ➡️ Why Me? I can easily create your robust backtesting system as I have 5 years of experience in Python development, specializing in data analysis, API integration, and automation. I also have a strong grip on database management and data visualization tools, ensuring a thorough approach to your project. ➡️ Let's have a quick chat to discuss your project in detail. I can provide samples of my previous work, showcasing my expertise in building trading systems. I look forward to discussing this with you in our chat. ➡️ Skills & Experience: ✅ Python Development ✅ Data Analysis ✅ API Integration ✅ Market State Reconstruction ✅ Historical Data Processing ✅ Automated Testing ✅ Data Visualization ✅ Database Management ✅ Backtesting Strategies ✅ Trading System Design ✅ Performance Optimization ✅ Report Generation Waiting for your response! Best Regards, Zohaib
£150 GBP in 2 days
8.1
8.1

Hi there, I understand you need a technically accurate Python research and backtesting system for pre-race exchange trading, starting with the paid technical test before moving into the full 4 GB historical dataset. The critical part is correctly reconstructing the persistent market state from incremental Exchange Stream deltas so that every snapshot and simulated trade is based on the actual historical market state. My approach is to first review the Exchange Stream documentation and sample .bz2 file structure, then build a reliable parser that decompresses and applies delta messages chronologically while preserving runner, price ladder, traded volume, suspension and market-status state. Next, I’ll generate the 10-, 5- and 1-minute pre-off snapshots with favourite/second-favourite identification, prices, available sizes, traded volumes and active-runner counts. I’ll explicitly document handling of missing fields, zero-size updates and unchanged values. I’ll then add automated tests covering state reconstruction and snapshot accuracy, identify suspension/in-play transitions, and commit the tested implementation and documentation to your private GitHub repository. Can you provide the sample WIN market file and Exchange Stream documentation so I can validate the delta-state reconstruction approach against the actual data format? I’m ready to start immediately. Warm Regards, Aneesa.
£250 GBP in 5 days
7.1
7.1

Hi there, I can build your paid Python technical test for the horse-racing Exchange Stream research system, accurately replaying the incremental .bz2 data to reconstruct persistent market and runner states. I’ll generate the required 10, 5, and 1-minute snapshots, analyze favourites, prices, liquidity and traded volumes, detect suspension/in-play timestamps, handle delta updates correctly, and provide automated tests with clean documentation in your private GitHub repository. Do you already have a representative sample WIN market file and the Exchange Stream documentation ready for review? Kindly send me a message to discuss more or directly award me. Thank you!
£230 GBP in 3 days
7.5
7.5

Greetings! I’m an expert in Python data pipelines and exchange-stream backtesting with 9+ years of experience, and I understand incremental deltas require persistent state reconstruction. Here's how I can help: * Parse/replay .bz2 Exchange Stream data * Reconstruct market/runner states chronologically * Generate 10/5/1-minute snapshots * Simulate back/lay orders with commission * Build auditable tests, reports & documentation Can you provide the sample file/API docs? Which exchange format/version should the parser target?
£135 GBP in 7 days
6.7
6.7

Hello, As an experienced team specializing in Python development and software architecture, we bring to the table the skill set perfect for your Horse Racing Backtesting project. We have vast experience dealing with large data volumes, building simulations, and generating precise and accurate reports - all of which are essential to your project's success. We understand the uniqueness of your requirements, including decompressing .bz2 files and working with incremental Exchange Stream updates, which require reconstructing a persistent state by accurately applying each update in timestamp order. Moreover, our proficiency is not limited to handling data; we also pride ourselves on building scalable web applications, scalable database architecture, and REST APIs - all key components of your project. We will ensure that your trading strategies are thoroughly tested using chronological out-of-sample data, generating auditable trade reports that hold up to the high standards that your business deserves. With our personalized approach to understanding clients' needs and our commitment to delivering high-quality solutions, we believe we can provide the robust research and backtesting system you’re looking for. Meticulousness is crucial as seen in our ISO 9001 and ISO 27001 certifications, ensuring that we maintain high quality standards while keeping client information secure. Thank you
£250 GBP in 5 days
7.0
7.0

Hello, Python Developer for Horse Racing Backtesting {{{ I HAVE CREATED SIMILAR PYTHON DATA PROCESSING, BACKTESTING, AND TRADING SYSTEMS BEFORE AND I CAN SHOW YOU }}} I have carefully reviewed your horse-racing backtesting requirements and understand that you need a technically accurate research system capable of replaying historical Exchange Stream data, reconstructing market and runner states, generating time-to-off snapshots, simulating back and lay orders, and producing auditable profit and trade reports. I have 11+ years of software and automation development experience and have created similar Python-based data processing, trading, API, backtesting, and automation systems before. I can show you relevant previous work so you can see my experience with this type of project. I can build the historical replay engine to decompress and parse the .bz2 Exchange Stream files, identify market and event information, and correctly reconstruct persistent market and runner states by applying incremental delta updates in timestamp order. Thanks, Christina
£200 GBP in 7 days
7.1
7.1

With expertise in Python development and data analysis, I understand the need to build a robust research and backtesting system for pre-race horse-racing trading. My experience includes working with large datasets and API integration for similar projects. How do you envision incorporating real-time data feeds into the system for live trading decisions? Regards, Yogesh Kumar
£60 GBP in 8 days
6.7
6.7

Hello There! I’m Md Toriqul Islam, and I’m excited to partner with you. I have strong experience in Python, data engineering, time-series processing, simulation, backtesting, and building reliable research systems from large historical datasets. I am skilled in Python, compressed data processing, event-driven systems, state reconstruction, pandas, automated testing, and reproducible backtesting workflows. I understand you need a technically accurate horse-racing exchange research system that can parse incremental .bz2 Exchange Stream data, reconstruct persistent market/runner states chronologically, generate fixed time-to-off snapshots, and later support realistic back/lay simulation and auditable P&L calculations. I have a few questions: 1) Which CRM are you currently using, and do you already have the required API/integration credentials? 2) Do you have a preferred WordPress builder such as Elementor, Divi, or Gutenberg? 3) How many landing pages are you planning to build after the pilot page? I’m ready to start with the paid technical test and would be happy to review the sample file and API documentation. Looking forward to hearing from you. Best regards, Md Toriqul Islam
£100 GBP in 4 days
6.6
6.6

The critical part here isn't decompressing the .bz2 files-it's reconstructing the exchange's incremental market state correctly. A backtest can appear accurate while being wrong if delta messages, timestamps, zero-size updates or unchanged ladder values are mishandled. I'd build the paid test around a persistent per-market state machine, applying updates strictly in chronological order and generating the 10/5/1-minute snapshots from the reconstructed state. Market and runner metadata, price ladders, traded volume, suspension and in-play transitions would be tracked separately, with explicit handling for missing and zero-size fields. I'd keep the replay engine independent from strategy logic so the later backtesting system can add different entry/exit rules, commission models and out-of-sample periods without rewriting the parser. Automated tests will make the replay deterministic and auditable. I'd also validate timestamp semantics early: snapshot selection should consistently reference scheduled start time while respecting actual stream timestamps and market lifecycle. Are the files standard Betfair Exchange Stream format, and can you provide the relevant stream specification/API documentation? For snapshot generation, should we use the last valid reconstructed state at or immediately before each target time? I can deliver the paid test with clean Python code, documentation, automated tests and the requested private GitHub commit. Juan Pablo
£200 GBP in 7 days
6.4
6.4

I can build the backtesting system around a persistent Exchange Stream state engine, correctly applying incremental updates, handling resets and zero-size ladder changes, and generating reproducible pre-race snapshots. I’ll begin with the paid test, including metadata extraction, favourite identification, suspension/in-play detection, realistic back/lay simulation, commission-aware P&L, automated tests, and clear audit documentation committed to your private repository.
£100 GBP in 7 days
6.3
6.3

Hi, I'm Denis, a developer who has built systems for processing and replaying high-frequency market data streams, including replaying historical exchange feeds for backtesting. This project involves reconstructing market state from incremental Exchange Stream updates, which is different from working with static snapshots. The key challenge is maintaining a persistent, accurate market state by applying each delta message in timestamp order. The snapshots at fixed times before the race start are a good way to validate state reconstruction. I’ve worked with similar streaming data where reconstructing the exact market state at any point in time was critical for strategy evaluation. For this task, I’d implement a clean state machine that processes each update sequentially, handles edge cases like zero-size ladder updates, and ensures no data is missed during decompression and parsing. I’d start by writing a small proof-of-concept using the sample file to confirm the parsing logic and state reconstruction. Once validated, I’d expand it into a full backtesting pipeline with automated tests, clear documentation, and a focus on correctness and maintainability. The main risk here is handling malformed or incomplete delta updates, especially during market suspension. I’d address this by logging unhandled cases and ensuring the state remains consistent even if some updates are skipped. I can start working right away. Let's connect and discuss the details. Thanks, Denis.
£66 GBP in 3 days
6.0
6.0

Hi, We’re building a backtester for pre-race WIN markets across GB and IE, turning Exchange Stream updates into reproducible research. Relevant experience signal: I’ve built Python tooling to decompress and replay tick data, reconstruct market states, and generate auditable outputs. Execution approach: I’d start by validating the current flow with a single sample file, decompressing, parsing needed IDs, and reconstructing state by applying delta messages in timestamp order; then implement ten-minute, five-minute, and one-minute snapshots and a test harness to verify data integrity against known outcomes. One technical risk / key challenge: Gaps in delta updates or missing ladder updates can drift the state if not handled properly. Two clarification questions: 1) How detailed should the trade reports and field mappings be, and what formats are preferred (CSV/JSON) for auditable output? 2) What are the expectations for handling missing fields or zero-size ladder updates during replay? If we're aligned, I can outline the implementation plan before we get started. Best regards, Brandon
£300 GBP in 3 days
6.0
6.0

Hi, I am a Python and algorithmic trading developer with 8 years of rich experience in software development, with a background in historical market-data processing, backtesting systems, trading logic, and API integrations. I am familiar with Python, C++, data processing, event-stream replay, backtesting, statistical analysis, market-state reconstruction, automated testing, GitHub, databases, and trading-system architecture. I can start with your paid technical test by parsing the .bz2 Exchange Stream file, rebuilding the persistent market and runner state from incremental updates, generating the 10/5/1-minute pre-off snapshots, identifying suspension and in-play transitions, and producing the required favourite, price, liquidity, LTP, volume, and active-runner data. I will also document how missing values, zero-size ladder updates, and unchanged fields are handled and include automated tests. I'm an individual freelancer and can work on any time zone you want. Please contact me with the best time for you to have a quick chat. Looking forward to discussing more details. Thanks.
£250 GBP in 7 days
6.0
6.0

Your requirement is clear, and this is a strong match for my Python/data-processing experience. I understand that the .bz2 Exchange Stream files are incremental delta updates, not ordinary snapshots, so the key part is correctly maintaining persistent market/runner state while replaying every message in timestamp order. For the paid technical test, I’ll build a clean Python parser that decompresses the sample file, identifies the market metadata and runners, applies the stream updates correctly, and generates the 10-minute, 5-minute and 1-minute pre-off snapshots. I’ll calculate favourite/second-favourite state, back/lay prices, available liquidity, LTP, traded volume, active runners, market suspension and in-play timestamps. I’ll explicitly document handling of zero-size updates, unchanged fields and missing data rather than silently making assumptions. I’ll also include automated tests and structured/auditable output so the same architecture can later be extended to the full 4 GB historical dataset and eventually the backtesting engine with realistic back/lay execution and commission calculations. I’m comfortable working chronologically to avoid look-ahead bias and keeping the entire development inside your private GitHub repository. I can start immediately with the paid sample test and, once the replay logic is validated against your expected results, continue with the complete historical backtesting system.
£50 GBP in 1 day
5.6
5.6

Hi, I'm Rafael. I’ll build a Python replay-and-backtest core that decompresses and parses your .bz2 Exchange Stream, reconstructs persistent market state by applying each delta in timestamp order, and captures fixed snapshots at 10, 5 and 1 minute(s) before scheduled start. I’ll include automated tests and commit the test code to your private GitHub so you can audit every step. A practical approach: keep an in-memory state machine keyed by market/runner IDs so snapshots are deterministic and cheap to generate; validate with pytest fixtures that check favourite, top back/lay and top-level volumes. For the paid test, do you prefer snapshot outputs as JSON or CSV? Please contact me through Freelancer chat to discuss this project in detail and estimate realistic timing and cost.
£250 GBP in 5 days
6.1
6.1

Betfair's Exchange Stream API doesn't send full order books, it sends deltas against a cached mcSnapshot, so the actual work in this test isn't decompressing bz2 files, it's correctly replaying imBook and mcBook increments against the right runner ladder without drifting after a few thousand messages. That's the part most attempts get subtly wrong. I'd build the pipeline as a stateful reader: stream-unzip the bz2 in chunks rather than loading it whole, parse each line as its own JSON message, and maintain one market cache keyed on marketId and selectionId that gets mutated in place as mc updates arrive. Zero-size levels mean remove that price point, not zero it out and leave a gap in the ladder. A field absent from a delta means unchanged, not zero, so the diff logic has to distinguish "not present" from "present as 0" at the parsing layer. I'd pull three snapshots at your chosen timestamps by replaying up to that point and serializing state, then back it with tests that construct a small synthetic delta sequence with known edge cases and assert the reconstructed book matches by hand-checked values, not just that the code runs without throwing. First milestone gets the reader working end to end against your actual file: bz2 streaming, delta application, one correct snapshot output, so you can check the numbers before I write the other two and the test suite around them. M1: streaming bz2 reader + delta state engine + first verified snapshot, GBP 120, 1 day. M2: remaining two snapshots, edge case handling (zero-size, unchanged, missing fields) and automated correctness tests, GBP 180, 2 days. Bid is off the brief as posted. If the actual flow files run bigger or the market has more runners than typical, I'll flag that before M2 starts rather than after.
£300 GBP in 3 days
5.8
5.8

Hi, As per my understanding: The real risk isn't backtesting logic, it's that these are incremental deltas, not snapshots, so a delta applied out of order or misread as a removal silently corrupts every downstream snapshot and P&L figure without an obvious error. That's why the test is scoped around reconstruction, not strategy. Implementation approach: I'd build market state as a per-runner state machine, applying each delta strictly in timestamp order and diffing against the last known good state so zero-size ladder updates and unchanged fields aren't misread as removals. The 10/5/1-minute snapshots read reconstructed state at that exact instant, using the last update before it rather than the nearest raw message, since messages rarely land exactly on those boundaries. Missing fields or ambiguous deltas get logged during processing so the write-up reflects what the sample file actually contains. Automated tests target the reconstruction logic specifically, since that's the piece most likely to drift wrong silently. A few quick questions: 1. For snapshots, hold the last known price on a timing gap, or interpolate? 2. Is commission fixed, or does it vary by market and need to be configurable? 3. Should tests cover reconstruction only, or also snapshot/report output format?
£98 GBP in 5 days
5.8
5.8

Hi There! I specialize in Python data processing and backtesting with 9+ years of experience and a team of 62 professionals. Here’s how we can help: 1. Parse and replay Exchange Stream .bz2 data accurately. 2. Reconstruct market states and generate time-to-off snapshots. 3. Build automated tests and auditable backtesting reports. Can you share the sample market file and API documentation for the paid technical test?
£135 GBP in 7 days
5.4
5.4

My name is Mahad Sheikh, and I believe I possess the expertise and skills that align perfectly with your project's requirements. With a solid background in API development, data processing, Python, and C++ programming, I assure you that your historical data and stream updates are in good hands. I understand the meticulousness of accurately reconstructing market states from incremental updates in timestamp order; this requires a deep understanding of data flow as well as an almost impeccable attention to detail. Having worked on numerous projects centered around data analysis and processing, I have become exceptionally skilled at handling complex datasets effectively and producing consistent and reliable results - skills that are absolutely indispensable for your project. Furthermore, I am ready to commit myself fully to ensure the proper completion of the initial technical test. It will be an opportunity for you to see firsthand how my skills can turn the vast amount of .bz2 data into snapshots at specific intervals before scheduled starts, manage runners' up-to-date statistics including their traded volumes, identify market suspensions correctly, handle missing fields adroitly, and much more.
£200 GBP in 5 days
5.3
5.3

Hi there, I went through your project description and understand you need a technically rigorous Python replay and backtesting system that will reconstruct Exchange Stream market states from incremental .bz2 data before any live trading integration. I will first build the paid technical test around one WIN market, decompressing and parsing the stream chronologically while maintaining persistent runner and market state. I will correctly apply delta updates, handle zero-size ladder changes and unchanged fields, identify suspension and in-play transitions, and generate validated snapshots at 10, 5, and 1 minute before the scheduled start. I will structure the parser and state engine with Python, pandas where appropriate, and automated pytest coverage. Outputs will include runner-level prices, available liquidity, LTP, traded volumes, active runners, and market totals in an auditable format. I will document every state-handling assumption and commit the complete test implementation to your private GitHub repository. Within 24-48 hours of being awarded, I will share a project blueprint covering the stream parser, state reconstruction, snapshot engine, testing, and reporting architecture. Could you provide the sample .bz2 file and Exchange Stream API documentation so I can validate the message schema before implementation? Let’s establish a correct replay engine first, then build the broader backtesting system on that foundation. Cheers, Imran.
£325 GBP in 3 days
5.5
5.5

Wakefield, United Kingdom
Payment method verified
Member since Jul 16, 2023
£20-250 GBP
£20-250 GBP
£750-1500 GBP
£20-250 GBP
$30-250 USD
$30-250 USD
₹750-1250 INR / hour
₹12500-37500 INR
$250-750 USD
$250-750 USD
$3000-5000 USD
$15-25 USD / hour
£20-250 GBP
$100-200 USD
₹12500-37500 INR
₹750-1250 INR / hour
$100-200 USD
$250-750 USD
$15-25 USD / hour
€8-30 EUR
₹750-1250 INR / hour
$30-250 USD
₹750-1250 INR / hour
$250-750 USD
$30-250 USD