
Closed
Posted
We are seeking an applied process-mining research engineer to implement a time-bounded, preregistered comparative analysis of a historical loan-application event log. This is not a conventional dashboarding or process-optimization consulting assignment. The primary objective is to build a reproducible analytical pipeline that applies three frozen diagnostic selection methods to the same discovery data, evaluates their resulting decisions on a concealed chronological holdout, and clearly distinguishes what the event log establishes from what remains unknown or merely hypothesized. The work will support a practitioner research presentation in early October. The selected contractor must be able to execute a written protocol faithfully, identify ambiguities before analysis, and resist changing analytical rules after seeing favorable or unfavorable results. Responsibilities You will: Inspect and validate the a Loan Application event log and its lifecycle semantics. Construct a canonical, case-level analytical representation from the raw event data. Document mappings among cases, applications, offers, activities, resources, lifecycle events and timestamps. Implement preregistered discovery and chronological holdout partitions. Calculate specified transition intervals, waiting measures, case durations, loops, rework, handoffs and variant features. Distinguish inter-event gaps from supported measures of active work or organizational waiting. Implement the frozen wait-first and automation-first comparison methods. Implement the preregistered calculations required by the Flow-First arm without silently adding analyst discretion. Support outcome-linked candidate analysis using frozen exposure definitions, adjusted models, uncertainty estimates and materiality rules. Run specified robustness and sensitivity tests, including business-calendar and lifecycle-definition checks. Preserve separation between discovery and holdout evidence. Produce audit-ready tables and exhibits for technical review and presentation. Maintain automated data-quality and calculation tests. Document analytical limitations and identify claims that the data cannot support. The methodology owner will perform the interpretive Flow-First diagnosis. The contractor will support its computational execution but will not silently redefine its target, mechanism or selection rule. Required qualifications Strong Python skills, particularly pandas or Polars and reproducible analytical workflows. Demonstrated experience with event logs, process mining, transaction histories or longitudinal operational data. Familiarity with XES and process-mining libraries such as PM4Py. Experience reconstructing process lifecycles and validating timestamp semantics. Ability to calculate case-, event- and interval-level measures without conflating them. Working knowledge of regression or time-to-event methods, bootstrapping and uncertainty estimation. Experience with chronological validation, holdouts or other leakage-sensitive evaluation designs. Proficiency with Git, automated tests and documented analytical environments. Ability to explain the difference among descriptive evidence, observational association, causal inference and simulated effects. Strong written documentation and attention to preregistered rules. Helpful but not required Experience with financial-services application or case-management workflows. Experience evaluating competing analytical or diagnostic methods. Familiarity with Celonis, Disco, UiPath Process Mining or comparable platforms. Experience preparing reproducibility packages or technical appendices for research. Familiarity with calendar-aware duration calculations and process-variant analysis. Commercial process-mining platforms may be used for exploration or visual validation, but the authoritative analysis must be reproducible through documented code without requiring a proprietary license. Deliverables Validated analytical dataset documented entity and event mappings; lifecycle reconstruction; discovery and holdout assignments; data-quality findings and exclusions; case-, interval- and candidate-level analytical tables. Reproducible analysis pipeline version-controlled source code; environment and dependency specification; deterministic configuration; automated data-quality and calculation tests; one-command or clearly documented rerun procedure. Frozen-method implementation wait-first calculation and selection output; automation-first calculation and selection output; Flow-First computational measures defined by the protocol; audit trail showing the input, calculation and selected candidate for each method. Holdout and robustness results frozen holdout evaluation; specified robustness and sensitivity tests; uncertainty estimates; explicit treatment of calendar effects, lifecycle ambiguity, missingness and segment instability. Technical findings package presentation-ready tables and charts; process and variant views where analytically useful; concise technical appendix; limitations and claim-boundary register; reproducibility handoff documentation. Optional exploratory interface A compact dashboard or notebook permitting inspection of individual cases, variants and key distributions may be included if it does not displace the authoritative analytical deliverables. Acceptance criteria The work will be accepted when: all frozen analytical rules are implemented without undocumented discretion; discovery and holdout data remain appropriately separated; every reported number can be traced to source events and executable code; the same inputs and configuration reproduce the same outputs; automated tests cover critical mappings, durations, partitions and selection calculations; exclusions, transformations and missing-data treatments are documented; observed associations are not presented as proven causal effects; simulations, if included, expose their assumptions and uncertainty and are not presented as observed gains; another qualified analyst can rerun and review the work without the contractor’s intervention; presentation exhibits agree with the authoritative computational outputs. Expected engagement Approximately 15–20 hours per week. Roughly 90–120 total hours, subject to final scope. Start as soon as possible. Core analytical work completed by late September to allow independent review before the early-October presentation. Short scheduled reviews at protocol clarification, pipeline validation, pre-holdout freeze and final-results stages. Application materials Please provide: one relevant example involving event-sequence or operational-process data; a short description of how you validated event and lifecycle semantics; your experience with Python, PM4Py or comparable tools; your approach to preventing holdout leakage and undocumented analytical discretion; your availability, hourly rate and anticipated hours through late September.
Project ID: 40663092
79 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
79 freelancers are bidding on average $21 USD/hour for this job

I am a skilled applied process-mining research engineer with extensive experience in event logs and process mining. With strong proficiency in Python and libraries like pandas, Polars, and PM4Py, I am well-equipped to meet the needs of your Applied Process-Mining project. My background includes implementing reproducible analytical workflows and working with financial services data, making me a strong fit for validating historical loan-application event logs. In previous projects, I constructed canonical analytical representations from raw event data and meticulously documented mappings among cases and events. I am familiar with preregistered analysis methodologies and have successfully managed holdout evaluations and sensitivity tests to prevent analytical discretion and ensure robust, reproducible results. I excel in ensuring all calculations and results can be traced back to source events and maintained through version-controlled environments. I am available to begin immediately and commit 15–20 hours weekly to meet your September timeline. I would like to discuss how my skills align with your specific needs and am happy to provide more detailed examples of my work.
$25 USD in 40 days
8.4
8.4

Hello, Answers to your questions: 1. Example: Built event-based lead-scoring/automation systems tracking lifecycle stages, timestamps, rules, and outcomes. 2. Validation: Verified entity mappings, event order, timestamps, lifecycle transitions, duplicates, and missing events while documenting assumptions. 3. Tools: Strong Python/pandas, PostgreSQL, Git, Docker, and automated testing; familiar with PM4Py/XES workflows. 4. Leakage: Use chronological discovery/holdout splits, prevent future data from selection logic, version rules, and audit every change. 5. Availability: Mon–Fri, 9:30 AM–7 PM IST, 40 hrs/week. $15/hr; 90–120 hours through late September. I can deliver a reproducible Python pipeline covering lifecycle validation, process analysis, strict holdout controls, regression, bootstrap uncertainty, sensitivity checks, automated testing, and fully traceable, audit-ready outputs. Looking forward to your reply! Best, Niral
$15 USD in 40 days
8.0
8.0

Hi there, I have carefully reviewed the project requirements for the Applied Process-Mining research engineer position. I understand the need for implementing a time-bounded, preregistered comparative analysis of a historical loan-application event log. Let's chat and discuss it further. To handle your project, I will start with inspecting and validating the Loan Application event log, constructing a canonical analytical representation, and documenting mappings among cases, applications, offers, activities, resources, and timestamps. I will then implement the necessary discovery and holdout partitions, calculate specified transition intervals, waiting measures, and distinguish inter-event gaps accurately. The deliverables for this project will include a validated analytical dataset, a reproducible analysis pipeline, frozen-method implementation results, holdout and robustness findings, and a technical findings package. Before signing-off my bid, I would like to ask a question, i.e., how crucial is the integration of financial-services application workflows in this analysis? Warm Regards, Aneesa.
$15 USD in 40 days
6.9
6.9

With a background in process mining and Python analytics, I'm well-equipped to handle your project needs. I'll conduct a precise comparative analysis of the historical loan-application event log using frozen diagnostic selection methods. Deliverables will include a validated dataset, reproducible analysis pipeline, implementation of frozen-method rules, holdout results, and technical findings. I propose 15-20 hours/week, totaling 90-120 hours, to be completed by late September for an independent review ahead of the October presentation. Let's work together to create a robust analytical pipeline for your research presentation.
$22.50 USD in 5 days
6.3
6.3

Hi, The part that matters most here is the discipline, not the mining. On a recent contract we built a deterministic ingestion pipeline with a secure CI setup where the same inputs always reproduced the same outputs, with automated tests guarding the critical steps. Secure CI Pipeline & Deterministic Ingestion: reproducible data workflow One thing I'd pin down before any analysis: the exact lifecycle-event semantics you want treated as authoritative, since inter-event gaps versus supported active-work or waiting measures depend entirely on how the log's start and complete events are read. I'd document that mapping first, freeze it, then split discovery and holdout so no rule shifts after seeing results. Which log are you using, the BPI 2017 loan application set or a private one? Adil
$23.44 USD in 40 days
5.9
5.9

Dear , We carefully studied the description of your project and we can confirm that we understand your needs and are also interested in your project. Our team has the necessary resources to start your project as soon as possible and complete it in a very short time. We are 25 years in this business and our technical specialists have strong experience in Python, Git, Data Visualization, Data Analysis, Pandas, Regression Analysis and other technologies relevant to your project. Please, review our profile https://www.freelancer.com/u/tangramua where you can find detailed information about our company, our portfolio, and the client's recent reviews. Please contact us via Freelancer Chat to discuss your project in details. Best regards, Sales department Tangram Canada Inc.
$25 USD in 5 days
7.3
7.3

You’re requesting a reproducible, preregistered applied process-mining research pipeline, specifically a time-bounded comparative analysis over a historical loan-application event log, with frozen diagnostic selection rules and a concealed chronological holdout. I can implement the end-to-end computational workflow in Python (pandas/Polars), including XES ingestion via PM4Py, canonical case-level representation, lifecycle reconstruction, timestamp semantic validation, and audit-ready mapping tables (cases ↔ applications/offers ↔ activities/resources ↔ lifecycle events). I will build deterministic partitioning (discovery vs holdout), calculate interval- and variant-level measures (wait/active-work separation, gaps vs active work, durations, loops/rework/handoffs, rework/iteration and transition intervals), and implement the frozen wait-first and automation-first selection methods exactly as specified by the protocol. I will ensure each reported quantity is traceable to source events and executable code, with automated data-quality/calculation tests, controlled missingness handling, and leakage-sensitive evaluation. Outputs will include holdout results, robustness/sensitivity checks (including calendar-aware duration calculations and lifecycle ambiguity checks), uncertainty estimation, and a limitations/claim-boundary register suitable for early-October presentation.
$20 USD in 47 days
5.4
5.4

The hardest part of this job is faithfully executing a written protocol on data that is not yet fully understood, so I will build a Python pipeline to load the historical loan-application event log using Pandas, and then implement the three diagnostic selection methods as functions, each taking the event log as input. I will build the data ingestion and cleaning first, so that the data is ready for analysis, then I will implement the three selection methods, and finally the evaluation on the holdout data, because the protocol dictates this order. I will pick the side of treating the event log as a single source of truth for the analysis, as the brief implies a need to distinguish established facts from hypotheses. What is the exact format and granularity of the "concealed chronological holdout" data and how should it be accessed for evaluation? For what it is worth, every job I have taken on Freelancer has gone out on time and on budget, 100% on both. I need the event log data file to start.
$25 USD in 7 days
5.3
5.3

Hi, I got that you are looking for an applied process-mining research engineer to implement a time-bounded, preregistered comparative analysis of a historical loan-application event log. This is what I can help you with, let's chat. My approach is to meticulously inspect and validate the Loan Application event log, construct a case-level analytical representation, and implement the frozen wait-first and automation-first comparison methods using Python, specifically leveraging pandas for data manipulation. By ensuring the separation of discovery and holdout evidence, I will maintain the integrity of the analysis. The deliverables will include a validated analytical dataset, a reproducible analysis pipeline with version-controlled source code, and the implementation of frozen-methods with audit trails for transparency. As final deliverables, you will receive a validated analytical dataset, documented entity and event mappings, a reproducible analysis pipeline, frozen-method implementation outputs, holdout and robustness results, and a technical findings package. One thing I'd like to confirm before we start: Are there any specific calendar effects or lifecycle ambiguities that need to be addressed? Looking forward to discussing this project further. Regards, Imran
$15 USD in 40 days
5.4
5.4

I understand you need a rigorous, data-driven comparison of diagnostic methods applied to a time-bounded event log, mirroring the analytical depth required for a preregistered study. My experience with similar event log analysis projects, focusing on reproducible pipelines and objective evaluation of findings against holdout data, aligns perfectly with your requirements. My approach will involve utilizing Python with libraries like `pm4py` for process discovery and analysis. I’ll implement a robust ETL process to clean and structure the loan-application event log, ensuring temporal integrity. The three diagnostic selection methods will be implemented as distinct, parameterized functions. I will then execute these methods on the training data, followed by a blinded evaluation against the chronological holdout set, quantifying performance metrics to objectively distinguish established insights from hypotheses. Could you clarify the specific nature of the "concealed chronological holdout" and the expected format for presenting the "clearly distinguished" findings? I'm keen to discuss how my technical plan can be tailored to meet these specific output requirements and ensure a successful, reproducible analysis.
$25 USD in 7 days
4.5
4.5

Hi there, Thank you for sharing such a thorough and thoughtfully designed project brief. I’m excited by the opportunity to contribute to your applied process-mining analysis and help ensure your practitioner research presentation is supported by a robust, transparent, and fully reproducible analytical pipeline. I bring extensive experience working with event log and process-mining datasets—including projects analyzing operational processes in both financial services and healthcare. In a recent assignment, I reconstructed loan origination event lifecycles, mapped granular activities to business entities, and implemented case-level analytical features for downstream modeling. Validating event and lifecycle semantics involved cross-verifying timestamp consistency, reconciling activity-resource mappings, and confirming event ordering directly with SME feedback and data-quality scripts, ensuring analytical measures reflected true business process dynamics. To prevent leakage and undocumented discretion, I adhere to preregistered protocols, implement automated tests for all critical partition and calculation steps, and document every transformation and exclusion. I also ensure that observed associations are clearly separated from causal claims, with all uncertainty estimates and simulation assumptions transparently reported.
$25 USD in 10 days
4.6
4.6

Hello, I am Dr. Rajesh Rolen, PhD in Computer Science & Engineering, with experience of over 20+ years in Data Visualization, Python As a preferred freelancer in the top 1%, I have done 400+ projects here on freelancer.com, I have 4.9 ratings out of 5 on average, which showcases my quality of work and timely delivery. Key Highlights: - Free Hosting Support on the Cloud or any desired platform. - Free 3 months of post-delivery support to ensure that our client doesn’t face any challenges after the launch of the project. - Free Dedicated tester on projects to ensure quality delivery, so clients don’t need to act as a tester. - 10+ Years experience UI/UX team to ensure intuitive UI. Portfolio: https://www.freelancer.com/u/Microlent Please open the chat and send me a message, so we can have a more detailed discussion about the project to give you the project timeline and cost. Thank you for considering my services. I look forward to engaging in a productive conversation and understanding how I can be of assistance in bringing your project to life. Regards Rajesh Rolen
$15 USD in 40 days
5.5
5.5

Good day! I can implement this as a reproducible, protocol-driven process-mining analysis rather than a conventional optimisation exercise. My approach would prioritise faithful execution of the frozen methodology, careful validation of event and lifecycle semantics, and a clear separation between what the historical event data demonstrates and what remains observational or uncertain. Using Python with pandas/Polars and PM4Py where appropriate, I can construct the canonical case-level representation, validate entity mappings and timestamps, implement discovery and chronological holdout partitions, and calculate the specified interval, duration, rework, handoff and variant measures. I will maintain strict separation between discovery and holdout data to prevent leakage and avoid undocumented analytical discretion after results are observed. I can deliver a version-controlled, reproducible pipeline with deterministic configuration, automated data-quality and calculation tests, frozen method outputs, robustness analyses, uncertainty estimates and audit-ready tables and charts. My focus will be ensuring every reported result is traceable to source events and executable code, with clear documentation of exclusions, limitations and claim boundaries for independent review and the final presentation.
$20 USD in 40 days
4.6
4.6

Hello, As a result of a detailed review of your protocol and acceptance criteria, I understand this is a reproducibility-focused research implementation, not a dashboard or process-optimisation exercise. I’m available to start immediately and can commit 15–20 hours/week through late September at $20/hour. For comparable event-sequence work, I reconstruct canonical case/event tables first and validate case IDs, activity/lifecycle transitions, timestamp ordering, duplicate events, missingness and impossible durations before calculating process measures. For your analysis, I would encode the preregistered rules as versioned configuration and tested functions, then freeze discovery/chronological-holdout assignments before evaluation. Wait-first, automation-first and Flow-First computations would share the same validated source representation while maintaining separate selection logic and audit trails. I’d use deterministic pipelines and tests for mappings, intervals, loops, handoffs, variants, partitions and selection calculations. Every exhibit would be generated from authoritative outputs, with associations, causal claims and simulated effects explicitly separated. One question: is the event log already available as XES, or will the canonical lifecycle need to be reconstructed from CSV/database extracts? Best regards, Carlos.
$15 USD in 40 days
4.4
4.4

Hi there, I've built reproducible pipelines for longitudinal operational data, and this preregistered process-mining design with chronological holdouts is exactly where strict separation between discovery and evaluation matters. I'll implement the three frozen diagnostic methods faithfully with a clean audit trail. What I'll do: ✅ Build the canonical case-level representation from the Loan Application event log, including lifecycle reconstruction and entity mapping ✅ Implement wait-first, automation-first, and Flow-First computational methods exactly as preregistered with full audit trails ✅ Run holdout evaluation, robustness tests, and uncertainty estimates with calendar-aware duration calculations ✅ Back up all intermediate datasets and version every pipeline stage before modifying anything ✅ Python (pandas), PM4Py, and event log analysis ✅ Reproducible workflows with automated data-quality and calculation tests ✅ Chronological validation and holdout evaluation design ✅ Regression, bootstrapping, and uncertainty estimation ✅ Microsoft® Certified: MCSA | MCSE | MCT ✅ 300+ projects delivered, 280+ five-star reviews I'm available around the clock and respond fast, so you won't be waiting during the push to your October presentation. Is the event log in XES format or raw tabular data, and is the lifecycle event mapping already defined or derived from the data? I can deliver in 5 days at $20 USD/hr and I'm ready to start immediately.
$20 USD in 5 days
4.5
4.5

Hi, Your emphasis on frozen analytical rules, chronological holdout separation and traceability is exactly how I would structure this engagement. I work in Python with pandas/Polars, statistical modelling, reproducible data pipelines, Git and automated testing. For this project I would use PM4Py/XES tooling where useful, while keeping the authoritative transformations and analytical outputs reproducible through documented Python code. I would first validate case IDs, activities, lifecycle transitions, timestamps, resources and application/offer relationships before calculating durations or waiting measures. This avoids treating inter-event gaps as active work or organisational waiting without evidence. Discovery/holdout assignment would be deterministic and frozen before holdout evaluation. Method parameters, exposure definitions, exclusions and selection rules would live in version-controlled configuration, with tests protecting partitions and calculations from leakage or undocumented changes. Outputs would include traceable analytical tables, robustness results, uncertainty estimates, presentation exhibits and a reproducibility handoff.
$25 USD in 40 days
4.5
4.5

Hello, The key part of this project is **turning a historical loan-application event log into a reproducible, preregistered analysis without leaking holdout evidence**. I can help you handle this accurately and efficiently without overcomplicating the process. I have hands-on experience with **Python, Pandas, and Data Analysis**, including cleaning longitudinal datasets, building audit-ready tables, and documenting calculation logic. For your project, I would focus on **validating event and lifecycle semantics**, **constructing the canonical case-level representation**, and **implementing the frozen discovery/holdout pipeline with traceable outputs**, while making sure the final result is **reproducible, testable, and clear about what the data can and cannot support**. I can start **immediately** and expect to complete this within **18 weeks at 15-20 hours per week**. One detail I'd like to confirm before starting: **will the event log be provided in XES, CSV, or another format, and is the preregistered protocol already frozen?** Best regards, Miguel
$20 USD in 18 days
4.2
4.2

Hi, I am a professional web developer and I can do this project "Applied Process-Mining", I have 5 years of experience in web development. I have done many projects like this. I can do this job for you. I can start right now. Please contact me. Thanks
$15 USD in 2 days
3.9
3.9

Nice to talk you , After reading in detail the requirements of your project and concluding that they match my areas of knowledge and skills, I would like to introduce myself. My name is Anthony Muñoz and I am the lead engineer for DS Pro IT agency. I have worked for over 10 years in Backend and software development and have successfully done multiple jobs. It will be a pleasure to work together to make your project a reality. Please feel free to contact me. I´m looking forward to working with you. I really appreciate your time and remain attentive to any request or question. Greetings
$21 USD in 40 days
3.8
3.8

Preregistered and time-bounded means clean event logs and a reproducible pipeline from day one. I would use Python and Pandas to structure the log, run regression on cycle times, and version everything in Git so it holds up under review. Can start today, first pass ready in 3 days. Budget and timeline are initial estimates from the post, refined after a quick scope chat. Want me to send a quick scope doc so we can get moving?
$25 USD in 14 days
3.6
3.6

Falls Church, United States
Member since Aug 22, 2026
₹600-1500 INR
₹75000-150000 INR
₹12500-37500 INR
$30-250 USD
$15-25 USD / hour
min ₹2500 INR / hour
₹600-1500 INR
$750-1500 USD
₹600-1500 INR
₹600-1500 INR
$750-1500 USD
$25-50 USD / hour
₹750-1250 INR / hour
$15-25 USD / hour
£250-750 GBP
₹1000-10000 INR
$10-20 USD / hour
₹750-1250 INR / hour
$14-30 NZD
₹600-1500 INR