Predictive ML
How should a risk score be trained, calibrated, explained, and reported?
Curated implementation map
A map of the papers, standards, and benchmarks that have shaped my work—from calibrated healthcare models and RAG to tool-using agents and representation engineering.
Editorial boundary: external work is never presented as mine. Status labels distinguish stable foundations, practical guidance, and emerging research.
Research behind the work
Three project-first paths connect source material to a design consequence. Each source remains available in the complete atlas below.
How should a risk score be trained, calibrated, explained, and reported?
How should evidence enter a model loop, and how should tool behavior be evaluated?
What can representation-level interventions reveal or change?
Research atlas
The default view favors durable, cross-domain foundations. Recent context and harness work is available in Frontier watch, not allowed to dominate the opening view.
12 records in Foundations
Field-shaping work across statistical learning, calibrated prediction, retrieval, tool use, representation engineering, and healthcare interoperability.
A durable, scalable baseline for structured risk modeling.
Additive feature attribution can aid debugging and communication, but is not causal proof.
Probability calibration belongs beside discrimination metrics when scores drive thresholds.
Post-hoc calibration methods can improve probability quality for decisions.
Combines parametric generation with retrieved evidence.
Relevant information placement changes long-context performance.
Interleaves reasoning traces with environment actions.
Studies learned decisions about when and how to call external tools.
Treats high-level concepts as directions that can be inspected and influenced.
Describes an app platform built on open healthcare interoperability standards.
Defines the R4 healthcare data exchange resources and API semantics.
Synthesizes published implementation evidence around the FHIR standard.
No records in this view match both filters. Reset the filters or choose another view.
Guidelines, evaluations, and benchmarks that influence how systems are documented, tested, constrained, and reviewed.
A reporting framework for transparent clinical prediction-model studies using regression or machine learning.
A risk-of-bias and applicability lens for prediction-model evidence.
Introduces reference-aware dimensions for evaluating retrieval-augmented systems.
Combines retrieval decisions with generation and critique signals.
Uses linguistic feedback and memory to alter later attempts.
Frames LM pipelines as programs that can be optimized against evaluation.
Evaluates agents across interactive environments and tasks.
Tests policy-constrained tool use in realistic user interactions.
Grounds coding-agent evaluation in real repository issues and tests.
Shows that the agent-computer interface materially shapes coding performance.
Defines workflow-triggered clinical decision-support service interactions.
No records in this view match both filters. Reset the filters or choose another view.
Emerging work tracked for useful signals—not treated as settled practice. Preprints remain visibly separate from stable foundations.
Maps the systems that assemble, manage, and evolve model context.
Treats context as an evolving system rather than a fixed prompt.
Proposes explicit language for agent context structure.
Examines the engineering layer around coding agents.
Studies natural-language structures that govern agent execution.
A separate research paper on optimizing model harnesses; unaffiliated with Shailesh’s project.
Connects harness adaptation with observability signals.
Surveys code as the execution structure around agents.
Targets measurement of harness effects apart from model effects.
Investigates how harness design changes repository search behavior.
Analyzes why representation-level interventions can generalize.
Studies where and when model representations remain steerable.
No records in this view match both filters. Reset the filters or choose another view.
Two authored records with publication metadata confirmed against publisher-hosted sources. Full text remains on the publisher archive.