Curated implementation map

Research & reading.

A map of the papers, standards, and benchmarks that have shaped my work—from calibrated healthcare models and RAG to tool-using agents and representation engineering.

Editorial boundary: external work is never presented as mine. Status labels distinguish stable foundations, practical guidance, and emerging research.

Research behind the work

Start with an engineering question.

Three project-first paths connect source material to a design consequence. Each source remains available in the complete atlas below.

Predictive ML

How should a risk score be trained, calibrated, explained, and reported?

Implementation consequence

Treat discrimination, probability quality, attribution, and reporting quality as separate engineering concerns.

Retrieval + agents

How should evidence enter a model loop, and how should tool behavior be evaluated?

Implementation consequence

Keep retrieval quality, context placement, action contracts, and task evaluation observable rather than collapsing them into one score.

Research atlas

The reading collection.

The default view favors durable, cross-domain foundations. Recent context and harness work is available in Frontier watch, not allowed to dominate the opening view.

01 / Foundations

What durable ideas shaped the work?

Field-shaping work across statistical learning, calibrated prediction, retrieval, tool use, representation engineering, and healthcare interoperability.

Foundation2016

XGBoost: A Scalable Tree Boosting System

A durable, scalable baseline for structured risk modeling.

Source
KDD
Status
Conference record
Cluster
Predictive models
Foundation2017

A Unified Approach to Interpreting Model Predictions

Additive feature attribution can aid debugging and communication, but is not causal proof.

Source
NeurIPS
Status
Conference record
Cluster
Predictive models
Foundation2017

On Calibration of Modern Neural Networks

Probability calibration belongs beside discrimination metrics when scores drive thresholds.

Source
ICML
Status
Conference record
Cluster
Uncertainty and calibration
Foundation2005

Predicting Good Probabilities With Supervised Learning

Post-hoc calibration methods can improve probability quality for decisions.

Source
ICML
Status
Conference record
Cluster
Uncertainty and calibration
Foundation2020

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Combines parametric generation with retrieved evidence.

Source
NeurIPS
Status
Conference record
Cluster
Retrieval and context
Foundation2024

Lost in the Middle: How Language Models Use Long Contexts

Relevant information placement changes long-context performance.

Source
TACL
Status
Peer reviewed
Cluster
Retrieval and context
Foundation2022

ReAct: Synergizing Reasoning and Acting in Language Models

Interleaves reasoning traces with environment actions.

Source
ICLR
Status
Conference record
Cluster
Agent loops and tools
Foundation2023

Toolformer: Language Models Can Teach Themselves to Use Tools

Studies learned decisions about when and how to call external tools.

Source
NeurIPS
Status
Conference record
Cluster
Agent loops and tools
Foundation2023

Representation Engineering: A Top-Down Approach to AI Transparency

Treats high-level concepts as directions that can be inspected and influenced.

Source
arXiv
Status
Preprint
Cluster
Representation engineering
Foundation2016

SMART on FHIR: A Standards-Based, Interoperable Apps Platform for Electronic Health Records

Describes an app platform built on open healthcare interoperability standards.

Source
JAMIA
Status
Peer reviewed
Cluster
Healthcare interoperability
Foundation2019

HL7 FHIR R4 Specification

Defines the R4 healthcare data exchange resources and API semantics.

Source
HL7
Status
Technical specification
Cluster
Healthcare interoperability
Foundation2021

The Fast Healthcare Interoperability Resources (FHIR) Standard: Systematic Review

Synthesizes published implementation evidence around the FHIR standard.

Source
JMIR
Status
Peer reviewed
Cluster
Healthcare interoperability
02 / Applied engineering

What changes an engineering decision?

Guidelines, evaluations, and benchmarks that influence how systems are documented, tested, constrained, and reviewed.

Applied engineering2024

TRIPOD+AI Statement

A reporting framework for transparent clinical prediction-model studies using regression or machine learning.

Source
BMJ
Status
Reporting guideline
Cluster
Clinical-model quality
Applied engineering2025

PROBAST+AI

A risk-of-bias and applicability lens for prediction-model evidence.

Source
PROBAST
Status
Reporting guideline
Cluster
Clinical-model quality
Applied engineering2023

RAGAS: Automated Evaluation of Retrieval Augmented Generation

Introduces reference-aware dimensions for evaluating retrieval-augmented systems.

Source
arXiv
Status
Preprint
Cluster
Retrieval and context
Applied engineering2023

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Combines retrieval decisions with generation and critique signals.

Source
arXiv
Status
Preprint
Cluster
Retrieval and context
Applied engineering2023

Reflexion: Language Agents with Verbal Reinforcement Learning

Uses linguistic feedback and memory to alter later attempts.

Source
NeurIPS
Status
Conference record
Cluster
Agent loops and tools
Applied engineering2024

DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Frames LM pipelines as programs that can be optimized against evaluation.

Source
ICLR
Status
Conference record
Cluster
Agent loops and tools
Applied engineering2023

AgentBench: Evaluating LLMs as Agents

Evaluates agents across interactive environments and tasks.

Source
arXiv
Status
Preprint
Cluster
Agent loops and tools
Applied engineering2024

τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Tests policy-constrained tool use in realistic user interactions.

Source
arXiv
Status
Preprint
Cluster
Agent loops and tools
Applied engineering2023

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Grounds coding-agent evaluation in real repository issues and tests.

Source
ICLR
Status
Conference record
Cluster
Coding agents and harnesses
Applied engineering2024

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Shows that the agent-computer interface materially shapes coding performance.

Source
arXiv
Status
Preprint
Cluster
Coding agents and harnesses
Applied engineering2024

CDS Hooks Specification

Defines workflow-triggered clinical decision-support service interactions.

Source
HL7
Status
Technical specification
Cluster
Healthcare interoperability
03 / Frontier watch

What is promising, recent, or still uncertain?

Emerging work tracked for useful signals—not treated as settled practice. Preprints remain visibly separate from stable foundations.

Frontier watch · recent preprint2025

A Survey of Context Engineering for Large Language Models

Maps the systems that assemble, manage, and evolve model context.

Source
arXiv
Status
Preprint
Cluster
Retrieval and context
Frontier watch · emerging record2025

Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Treats context as an evolving system rather than a fixed prompt.

Source
arXiv / conference record
Status
Conference record
Cluster
Retrieval and context
Frontier watch · emerging record2026

A Language for Describing Agentic LLM Contexts

Proposes explicit language for agent context structure.

Source
CAIS 2026 record
Status
Conference record
Cluster
Retrieval and context
Frontier watch · recent preprint2026

Harness Engineering for Agentic AI Coding Tools: An Exploratory Study

Examines the engineering layer around coding agents.

Source
Conference-related preprint
Status
Preprint
Cluster
Coding agents and harnesses
Frontier watch · recent preprint2026

Natural-Language Agent Harnesses

Studies natural-language structures that govern agent execution.

Source
arXiv
Status
Preprint
Cluster
Coding agents and harnesses
Frontier watch · recent preprint2026

Meta-Harness: End-to-End Optimization of Model Harnesses

A separate research paper on optimizing model harnesses; unaffiliated with Shailesh’s project.

Source
arXiv
Status
Preprint
Cluster
Coding agents and harnesses
Frontier watch · recent preprint2026

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Connects harness adaptation with observability signals.

Source
arXiv
Status
Preprint
Cluster
Coding agents and harnesses
Frontier watch · recent preprint2026

Code as Agent Harness

Surveys code as the execution structure around agents.

Source
arXiv
Status
Preprint
Cluster
Coding agents and harnesses
Frontier watch · recent preprint2026

Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Targets measurement of harness effects apart from model effects.

Source
arXiv
Status
Preprint
Cluster
Coding agents and harnesses
Frontier watch · recent preprint2026

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

Investigates how harness design changes repository search behavior.

Source
arXiv
Status
Preprint
Cluster
Coding agents and harnesses
Frontier watch · recent preprint2025

Why Representation Engineering Works: A Theoretical and Empirical Analysis

Analyzes why representation-level interventions can generalize.

Source
arXiv
Status
Preprint
Cluster
Representation engineering
Frontier watch · recent preprint2024

A Timeline and Analysis for Representation Plasticity in Large Language Models

Studies where and when model representations remain steerable.

Source
arXiv
Status
Preprint
Cluster
Representation engineering
04 / My publications

What can be attributed with confidence?

Two authored records with publication metadata confirmed against publisher-hosted sources. Full text remains on the publisher archive.

Authored publication2018

Verified author position · 3 of 4

A salutary biotechnical approach for explosive identification and border patrol using electrophysiological signals

C Santhanakrishnan · T Peermeer Labbai · Shailesh S. Dudala · Y Sai Santhosh Nag

A lightweight sensor rover controlled through EEG and EOG signals, with PIR and ultrasonic sensing plus ZigBee telemetry for remote observation.

Venue
International Journal of Engineering and Technology 7(2.31), 106–109 (2018)
Publisher
Science Publishing Corporation
Author position
Shailesh S. Dudala · 3 of 4
Persistent ID
DOI 10.14419/ijet.v7i2.31.13408
Authored publication2018

Verified author position · 3 of 3

Monitoring of Suspicious Discussions on Online Forums Using Data Mining

Tanya Srivastava · R. Mangalagowri · Shailesh S. Dudala

A Python and NLTK forum-monitoring framework using preprocessing, similarity, keyword categories, sentiment, and classifier methods, with explicit scalability and security limits.

Venue
International Journal of Pure and Applied Mathematics 118(22), 257–262 (2018)
Publisher
Academic Publications, Ltd., Sofia
Author position
Shailesh S. Dudala · 3 of 3
Persistent ID
No journal DOI verified