Work

Research

The AgentRelBench paper was accepted as a poster at the NeurIPS 2026 workshop “Who Verifies the Agents? Toward Reliable Agent Development” (Sydney, December 2026), and sqlpup is open source, with an official entry on the BIRD leaderboard. In both, the headline finding went against what I was hoping for, and the write-ups say so.

  • sqlpup33.15% BIRD test execution accuracyCode
  • AgentRelBench2,128 pre-registered agent runsarXiv Code
2026 08.11[LLM, Pretraining, Research]

sqlpup

394M Text-to-SQL Model Built From Scratch

A 394M-parameter text-to-SQL decoder where every stage is built for this one task: a 9.67B-token corpus, a byte-level BPE tokenizer, a Llama-style GQA architecture, a preemption-safe 4-GPU training loop, a sandboxed execution-accuracy harness, and fine-tuning with SFT and GRPO over an execution reward. No pretrained weights anywhere in the stack.

  • Official BIRD leaderboard entry, scored by the BIRD team on the hidden test set: 33.15% execution accuracy with 7-sample voting, 27.33% single pass (dev: 23.51% and 19.23%)
  • Fine-tuned identically, Qwen2.5-0.5B (matched non-embedding size, ~1,861× more pretraining data) beats sqlpup by 5.87 points on BIRD dev at greedy decoding; more test-time compute widened the gap to 8.65 points instead of closing it
  • Corpus, 32,768-entry tokenizer, and training loop all built from scratch, with exact-resume checkpointing that survives spot preemption
2026 08.15[Agents, Evaluation, Research]

AgentRelBench

Reliability Instrument for Action-Taking LLM Agents

An instrument that prices agent damage from database state diffs, deterministically and with no LLM judge anywhere in the measurement path, then measures whether that damage repeats across runs. Built over EnterpriseOps-Gym (ServiceNow AI Research and Mila).

  • Paper accepted as a poster at the NeurIPS 2026 workshop “Who Verifies the Agents? Toward Reliable Agent Development” (Sydney, December 2026)
  • Pre-registered campaign of 2,128 runs across 20 tasks and nine models in six model families; no task caused damage on every run
  • On the development pool, a single clean run missed a damage-producing (model, task) pair 80% of the time; one model family made a gated, irreversible change while claiming it had refused
2026 04.02[LLM, Systems, AWS]

ModelRouter

Multi-Model Inference Cost Optimizer

A routing layer that scores each query’s difficulty with a fine-tuned DistilBERT and sends it to the cheapest of Llama 8B, Llama 70B (Groq), or Claude Sonnet 4 (Bedrock) that can handle it.

  • Fine-tuned DistilBERT difficulty classifier: 96% accuracy at 30ms inference
  • Multi-factor routing engine scoring models on quality, cost, latency & success rate
  • 58.6% measured cost savings vs an all-Tier-3 baseline across 24 live routed queries, zero errors
2026 03.11[Agents, LLM, Full-Stack]

ReproAgent

ML Paper Reproducibility Analysis

A multi-agent AI system that parses arXiv papers, extracts training methodology with per-field confidence scoring, and generates self-contained PyTorch scripts using Llama 3.3 70B via Groq.

  • Keyword-density chunking to prioritize experiment-heavy sections within token budgets
  • 0.91 average extraction confidence across parsed papers
  • Reproducibility scores of 94.3 (ResNet) and 89.1 (Attention Is All You Need)
2026 02.19[ML, Statistics, Full-Stack]

CalibratedReturns

Stock Predictions with Coverage Guarantees

A conformal-prediction model that outputs 5-day return intervals with a coverage guarantee, trained on 20+ engineered features and checked against 3,755 backtested predictions.

  • Quantile regression over 20+ engineered features
  • 92.3% empirical coverage across 3,755 backtested predictions
  • Full-stack app (FastAPI + React) deployed on Render and Vercel
2025 04.22[Computer Vision, Deep Learning]

Group Activity Recognition

Volleyball Video Analysis & Rally Prediction

A hierarchical multi-task deep-learning system for volleyball video that jointly predicts individual actions, group activities, and rally outcomes by combining AlexNet features, BiLSTM temporal modeling, self-attention, and a Spatio-Temporal GCN.

  • Jointly predicts 9 player actions, 8 group activities, and rally outcomes across 1,337 clips
  • 73.90% rally accuracy, +24.4 pts over the heuristic baseline (49.51% to 73.90%)
  • Weighted multi-task loss (Focal + Margin Ranking + MSE) with class re-weighting