Hi! I’m Sully. I’m a 27-year-old medical student at Duke University School of Medicine. I’m also a staff research scientist at OpenAI, where I work on developing fundamental methodologies for advancing reasoning and text generation. I’m broadly interested in improving the standard of living and quality of care for all people through two main paths:
Developing highly specialized, domain-specific AI systems that directly impact clinical care and/or biomedical research. This is largely the focus of my academic research.
Hi! I’m Sully. I’m a 27-year-old medical student at Duke University School of Medicine. I’m also a staff research scientist at OpenAI, where I work on developing fundamental methodologies for advancing reasoning and text generation. I’m broadly interested in improving the standard of living and quality of care for all people through two main paths:
Developing highly specialized, domain-specific AI systems that directly impact clinical care and/or biomedical research. This is largely the focus of my academic research.
Working toward artificial general intelligence (AGI), which has already begun to tangibly reshape our world. I’m particularly hopeful that AGI will dramatically accelerate the pace of scientific discovery, tear down barriers and disparities in healthcare, and provide high quality, personalized education to all people.
My life path has been super unconventional, and as a result I get a lot of questions. Here’s a quick timeline/description of my life to help answer a few:
January 1999 - June 2017:
I was born in Los Angeles in January 1999. I loved video games as a kid, which led me to learn C++ around the time I turned 10-years-old in hopes of becoming a video game developer. It turns out you need a lot of math to render graphics, so I learned most of what would be taught in undergraduate multivariable calculus as well as some linear algebra by the end of middle school. In high school, I became completely captivated by artificial intelligence – I was fascinated by the concept of “intelligence”, how it could be replicated by machines, and generally how complex behaviors could arise from simple-ish rules. My ultimate dream was to build a world where everyone could live a life of abundance, free from the burdens of disease, poverty, and labor. I did some open source work on self-driving cars and educational stuff near the end of high school, leading to some cool internship opportunities from Nvidia and Ultraleap.
June 2017 - May 2021:
I ended up pursuing a degree in computer science at California Polytechnic University, but I dropped it after taking my first CS class (learning computer science for a grade, to me, kills the fun). I switched my major to mathematics, then wrote my what would have been my senior thesis in mathematics on analytical number theory during my freshman year, which was later published in a journal by the European Mathematical Society. I had a somewhat sporadic change of heart and decided to pursue medicine, so I transferred to the University of Southern California, where I majored in biochemistry.
January 2022 - October 2023:
I stayed at the NIH for about 4 months, when I was abruptly offered a full-time research position at OpenAI. I was thrilled to be one of the first few hundred employees at OpenAI, working directly under Ilya and Mark on what would later become the reasoning team. I primarily worked on test-time compute during this time, and I can say without a doubt these are some of my fondest memories. Ilya and Mark are amazing people, incredibly brilliant scientists and leaders, and were so immensely supportive of my growth as a researcher. I worked there for 7 months during which time I was also abruptly offered a spot at Duke University School of Medicine. I was so eager to start that I left my position at OpenAI to start medical school. I finished my preclinical year at Duke, as well as my anesthesia/surgery rotation, passed USMLE Step 1, before taking a research leave to return to OpenAI.
October 2023 - August 2025:
On research leave working at OpenAI :)
August 2025 - May 2027:
Currently finishing up my MD degree!
Large-Scale Multi-omic Biosequence Transformers for Modeling Protein-Nucleic Acid Interactions
We developed OmniBioTE to study whether one model can learn the shared language of proteins, DNA, and RNA. It connects biological sequences across these domains and predicts how proteins interact with nucleic acids, including binding strength and residues involved in contact.
Protein–DNA complex with the protein colored by predicted contact involvement.
Approach
We pretrained encoder transformers on 250 billion tokens from GenBank and UniRef100, comparing them with protein-only and nucleic-acid-only models given equal compute across multiple model sizes. We trained the models to reconstruct masked sequence content, then fine-tuned them for specific tasks. Our evaluation covered binding-energy prediction, gene–protein representation alignment, binding specificity, contact prediction, and established protein and nucleic acid benchmarks.
Findings
We found that joint training aligned gene and protein representations without paired pretraining. OmniBioTE outperformed the tested single-domain controls on binding-energy prediction, and attention maps after affinity training contained information about molecular contacts. Across sequence benchmarks, mixed-domain training often improved performance per unit of compute, although absolute results varied across the individual tasks and tokenization methods.
Sparse learned kernels for interpretable and efficient medical time series processing
We developed SMoLK to make medical signal analysis easier to inspect and cheaper to run. Our compact model recognizes waveform patterns and shows how those patterns contribute to predictions, addressing both corrupted optical pulse recordings and atrial fibrillation in single-lead ECGs.
ECG traces colored to show local contributions to the model’s prediction.
Approach
A single sparse layer combines learned convolution kernels: short waveform templates that scan a recording. Their contributions can be traced directly to the output during inference, enabling precise signal-level attribution. Weight absorption and pruning of correlated kernels reduce redundant parameters. We compared this architecture with larger neural networks on pulse-signal artifact segmentation and ECG classification.
Findings
We found that SMoLK matched models with substantially more parameters on the evaluated tasks. For ECG classification, it matched a deep residual network with less than one percent of its parameter count and performed better in low-data experiments. The smallest pulse-signal model used twelve inspectable kernels, while additive contributions exposed which waveform regions supported or opposed a classification.
LLM-assisted systematic review of large language models in clinical medicine
We reviewed how large language models have been evaluated in clinical medicine. We separated tests on patient data from simulations and exam questions to examine differences in clinical realism, study design, and human comparison groups.
Reported model outperformance rates vary across human comparator training levels.
Approach
We searched PubMed, Embase, and Scopus for studies from January 2022 through September 2025. We used GPT-5 to screen titles and abstracts, assign evidence tiers, and extract study characteristics. Independent human review validated screening on 500 candidate studies and assessed tiering on a separate sample. We compared clinical tasks, specialties, datasets, and reported human–model comparisons across the evidence base.
Findings
We included 4,609 studies, but only 1,048 used real patient data and 19 were prospective randomized trials. Simulations and exam-style tasks dominated. Across 1,046 head-to-head comparisons, models outperformed humans in 33 percent, with results depending on task realism and comparator training. At least a quarter of studies used sample sizes smaller than thirty in their evaluations.
An interpretable peptide-HLA model emergently learns binding energetics and structure
We developed LAMINA to predict how peptides bind to human leukocyte antigen proteins, which display molecular fragments to the immune system. We designed its predictions to be attributable to interactions between sequence patterns, connecting interpretability with binding energetics and molecular structure.
Molecular rendering of a peptide–HLA complex with a bound peptide.
Approach
Our model represents every contiguous segment of the candidate peptide and the HLA pseudosequence as a learned, flexible motif. It scores pairs of peptide and HLA motifs, then combines those interaction scores into a prediction. This construction exposes the components of the prediction directly. We trained on sequence data, then analyzed learned representations using structural complexes.
Findings
With 4.7 million parameters and approximately ten hours of training on one desktop workstation, LAMINA matched or exceeded the tested state-of-the-art predictors on held-out binding-affinity regression and remained competitive on rank correlation. Its internal states also correlated with interaction energies and structural measurements in characterized peptide–HLA complexes, despite receiving no structural information during the sequence-based training process.
Posted to bioRxiv on September 14, 2026, as a preprint.
The irrationality measure of π as seen through the eyes of cos (n)
We connect the behavior of powers of cosine at integer inputs to how closely rational numbers can approximate π. A seemingly elementary sequence becomes a way to investigate the conjecture that π has irrationality measure exactly two.
Numerical cosine-power sequences for different values of the exponent γ.
Approach
We study |cos(n)| raised to n^γ as γ varies. We relate integers close to multiples of π to good rational approximations, then use bounds on cosine and asymptotic estimates to establish a transition in the sequence's limiting behavior. Numerical plots and analysis of arithmetic subsequences complement the proof and motivate further questions about the boundary case.
Findings
The transition occurs at γ = 2μ(π) − 2: below this threshold, the limit superior is one; above it, the sequence approaches zero. Consequently, showing that the limit superior at γ = 2 differs from one would establish μ(π) = 2. Computations reveal persistent peaks and increasingly sparse subsequences, providing numerical evidence consistent with that conjecture.
Intraoperative brain tumor classification via laser-induced fluorescence spectroscopy and machine learning
With TumorID, we combine laser-induced natural tissue fluorescence with machine learning to classify brain specimens during surgery. We examined whether this rapid, nondestructive measurement can distinguish glioma, meningioma, pituitary adenoma, and nonneoplastic tissue without adding fluorescent dyes.
Illustration of an optical probe directed toward a brain tumor.
Approach
We scanned removed tissue from 46 patients in the operating room using a 405-nm laser. Each fluorescence measurement required 0.5 seconds. A support vector machine learned to distinguish the four tissue categories from these spectra. We also compared emission regions associated with free and bound NADH, flavin adenine dinucleotide, and neutral porphyrins to examine the biological signals behind classification.
Findings
Our model achieved a multiclass receiver operating characteristic area under the curve of 0.809 ± 0.002. Neutral-porphyrin emission regions differed significantly and contributed most strongly to model output. These results demonstrate rapid classification of removed specimens and support investigating whether the device could eventually help surgeons identify tissue types and tumor boundaries during an operation.
We describe GPT-4, a multimodal model that accepts images and text and generates text. We document performance across academic and professional benchmarks, investigate predictable scaling, and report work to improve factuality and alignment with intended behavior.
Approach
We pretrained GPT-4 to predict the next token using public and licensed data, then refined it through reinforcement learning from human feedback. We evaluated professional examinations, language tasks, and code generation. We also fitted scaling relationships using smaller training runs to predict the final model's loss and selected coding capabilities before the full training run was complete.
Findings
GPT-4 reached approximately the top tenth of test takers on a simulated bar examination and improved on a broad collection of language benchmarks. Some aspects of its performance could be predicted from models using substantially less training compute. We also describe post-training improvements in factuality and desired behavior, alongside evaluations of remaining limitations and safety risks.
We use satellite imagery and neural networks to locate terrestrial waste accumulations at regional scale. We combine automated detection with human confirmation, then track site footprints and nearby environmental features to help characterize potential sources of plastic pollution.
Satellite imagery showing a cleared site and surrounding landscape.
Approach
Our system analyzes Sentinel-2 imagery with separate neural networks for individual pixel spectra and larger image patches. Paired observations capture changes over time, while semi-supervised training expands the patch classifier's learning data. We checked candidate detections against higher-resolution imagery and other available evidence. Confirmed sites can then be monitored over time and cross-referenced with geographic and social datasets.
Findings
We identified 374 confirmed waste sites in Indonesia, more than twice the number in the public databases used for comparison. Deployment across Southeast Asia produced 996 confirmed sites without additional regional training. Nineteen percent lay within 200 meters of a waterway, highlighting locations where accumulated waste could enter rivers and ultimately contribute to pollution in the ocean.