Predictive healthcare ML
Predictive healthcare ML: from cohort definition to operational action
A portfolio of readmission, length-of-stay, utilization, healthcare quality, and FWA models built around calibrated risk, explainable review, and workflow fit.
- Problem
- Healthcare risk scores could be statistically sound yet unusable when cohorts, horizons, thresholds, or reviewer capacity did not match the operating decision.
- Decision
- Treat cohort definition, calibration, threshold policy, and workflow capacity as one release contract for human prioritization.
- Business / proof change
- Supported operational prioritization across payer and hospital contexts, including an 18% reduction in transportation waste in the documented FWA workflow.
- Ownership
- Data science and ML engineering across cohort design, features, modeling, evaluation, deployment, monitoring, and stakeholder adoption
- Scale / team
- Cross-functional delivery with clinical, care-management, quality, operations, product, and data-platform partners
- Evidence
- 3 source notes · 3 explicit limits
Supported prioritization and operational improvement across payer and hospital contexts, including an 18% reduction in transportation waste in the documented FWA workflow.
A model earns release one decision at a time.
Switch the evaluation lens. The leading row changes when probability quality, slice stability, and operating cost matter alongside discrimination.
- AUROC
- .74
- Brier
- .168
- Calibration
- Stable
Interpretable release reference
- AUROC
- .79
- Brier
- .151
- Calibration
- Corrected
Challenger after calibration
- AUROC
- .80
- Brier
- .165
- Calibration
- Variable
Rejected: complexity did not earn release cost
One discipline across several prediction problems
Readmission, length of stay, emergency-department utilization, quality gaps, and transportation anomalies have different labels and operating rhythms. They share a harder requirement: the score has to arrive with enough context for a real team to act responsibly.
This case study is a public-safe synthesis of multiple engagements. It shows the modeling and delivery discipline without collapsing distinct clients, cohorts, or outcomes into one fictional product.
My role
I worked across cohort definition, data-quality rules, feature engineering, baseline and challenger models, calibration, explainability, evaluation slices, deployment, monitoring, and the analytics surfaces used by operations and care teams.
The product question was always paired with the statistical one: who will use the score, what action can follow, what capacity exists, and what failure is more costly?
Cohorts and labels before algorithms
A technically clean model can still be wrong for the workflow if the index event, observation window, prediction horizon, exclusions, or outcome definition drift from the operational decision. I treated those definitions as versioned contracts.
Claims, ADT events, demographics, diagnoses, medications, prior utilization, quality signals, and social-risk features were joined with explicit provenance and point-in-time boundaries.
- Prevent post-outcome leakage
- Keep training and scoring definitions aligned
- Measure missingness and freshness by source
- Document proxy and access concerns
Baselines, challengers, and calibration
Interpretable baselines established whether additional complexity earned its operational cost. Logistic models, rules, and simple risk scores were compared with boosted trees and task-specific challengers.
AUROC alone was never the release decision. Precision-recall behavior, Brier score, calibration curves, threshold tradeoffs, and slice-level stability mattered because teams consume ranked probabilities under limited capacity.
Explainability as a review aid
Feature attribution helped debug data and communicate why a record moved in the ranking. It did not establish causality or justify an intervention by itself.
Useful reviewer surfaces paired risk bands with current evidence, temporal context, missingness, and the action boundary. Sensitive or weak proxies required explicit review rather than a prettier explanation chart.
From score to operating system
Batch and API delivery paths were designed around the receiving workflow: refresh cadence, available capacity, escalation rules, feedback capture, and the difference between informational and action-triggering outputs.
Monitoring separated service health, data drift, score distribution, calibration, slice behavior, downstream action, and outcome. A healthy endpoint was not treated as a healthy model.
What failed or underperformed
Leakage from post-event fields, changes in coding practice, sparse subgroups, shifting utilization patterns, and threshold choices disconnected from operational capacity all created failure modes that a single aggregate metric would hide.
Where a model did not beat an interpretable baseline or could not be connected to a responsible action, the right outcome was to simplify, keep it in shadow mode, or stop.
Documented outcomes
In the transportation FWA workflow, explainable anomaly review supported an 18% reduction in waste. Other models supported care-management prioritization, utilization insight, quality work, and the reusable analytics platform described in the adjacent case study.
These outcomes belong to their documented workstreams. This page does not imply clinical validation, universal performance, or that a model alone produced a program result.
What I would improve next
I would formalize dataset and model cards in the release path, expand temporal and subgroup validation, pair calibration monitoring with action-rate monitoring, and require a shadow-mode checkpoint before every material cohort or feature change.
Source notes
What you can inspect—and what remains private.
Career record
The predictive-healthcare scope, model families, engagement contexts, and qualified FWA outcome are supported by the governed career record.
Public reference API
A separate synthetic FHIR/ML API makes an adjacent public contract inspectable; it is not a production clinical model.
Open source ↗Research foundations
The Research page connects calibration, explainability, clinical reporting, and model-quality references to this work.
Open source ↗Limits
- No production datasets, model weights, thresholds, or client definitions are published.
- The illustrated comparison uses synthetic data and does not reproduce a private model result.
- Model scores supported human prioritization and workflow decisions; they did not replace clinical judgment.
All cohorts, thresholds, feature examples, curves, and workflow artifacts are synthetic reconstructions. No member data, client records, proprietary definitions, or production model files are published.
Illustrative evaluation contract
evaluation_release:
cohort: synthetic_adult_inpatient
horizon_days: 30
baseline: logistic_regression
challenger: gradient_boosted_trees
metrics: [auroc, auprc, brier, calibration_slope]
slices: [age_band, service_line, prior_utilization]
decision: reviewer_prioritization_only
release_state: shadow_ready