Skip to work

Uday Tyagi

Cornell Universitygraduating December 2026

I build AI systems, and most of my work goes into checking what they are actually doing.

  • M.Eng. Computer Science, Cornell, December 2026
  • Software Engineer, KLA
  • B.S. Computer Science, UW–Madison, finished in two and a half years

My work sits on building safe and reliable AI systems: making models more interpretable, evaluating them for safety, and building the software around them that makes that possible.

steering demo · sycophancy directionh ← h + λ·v

prompt · “here’s my startup idea, what do you think?”

It's a competitive space. The idea could work with a clear wedge, but the economics depend on acquisition cost, which is worth pinning down first.

response is illustrative, not a logged generation

0.0
This is what my NLA Steering project does, except the vector is generated from a sentence of English instead of collected from labeled examples.

01about

I'm from Chicago. I grew up in Delhi until I was twelve, then spent my teens in Chicago, which is a long way of saying I've had to learn two places well enough to be from both.

I'm finishing a master's in computer science at Cornell in December 2026, after a B.S. at UW–Madison I got through in two and a half years. I'm a software engineer at KLA, where I work on LLM agents and the tooling that makes their actions auditable.

At Cornell I lead operations for Cornell AI Alignment and run the weekly technical paper reading group. Before that I led a 10-student cohort through the AI Safety Fundamentals curriculum at the Wisconsin AI Safety Initiative, and was part of their technical scholars group.

Uday Tyagi in a winter jacket in front of the Chicago skyline at night, city lights reflecting off the lake.
chicago, where the two halves of that story meet

02work & research

KLA

Software EngineerRemoteMay 2025 – present

Fine-tuned an LLM-powered agent for the DART web application so engineers can drive defect-inspection workflows in natural language, from image classification through event routing and triage.

I extended that agent with a custom MCP server exposing internal inspection and triage tools, giving the model structured, auditable tool-calling instead of ad hoc API integrations: the same instinct as the safety work, which is knowing what the system did and why. Also built visual regression and UI-performance monitoring for HPC automation workflows with Grafana Faro and custom metrics tooling.

  • LLM agents
  • MCP
  • fine-tuning
  • observability
  • Grafana Faro

NLA Steering

Text-driven behavioral steering of LLMs

Independent researchout of the CAMBRIA programMay 2026 – present

I used Natural Language Autoencoders (Fraser-Taliente et al., 2026) to generate steering vectors for Qwen2.5-7B from hand-written text descriptions alone: no training data, no activation extraction, no labeled examples. Write down what the behavior is, get a vector that produces it.

The vectors reliably shift behavior at inference time across sycophancy, misalignment, and persona, and they're geometrically concept-specific rather than a generic nudge. I'm now ablating which parts of the text format actually carry the steering signal, and benchmarking the vectors as zero-shot behavioral probes against ground-truth activation-derived ones.

  • PyTorch
  • HuggingFace
  • Qwen2.5-7B
  • steering vectors
  • mechanistic interpretability

Constitutional Drift

Advised by Prof. Lionel Levine · with Rauno Arike and Katie Lu

ResearcherLAISR · Cornell UniversityJan 2026 – present

We're measuring whether a model's stated values survive contact with itself: hand an LLM its own constitution, let it rewrite it, repeat, and track semantic drift, value collapse, and self-reinforcing edits across rounds.

I built the experimental harness with two collaborators: multi-round self-editing pipelines across model families, embedding- and judge-based drift metrics, and behavioral evals that check whether an edited constitution actually changes what the model does downstream, or only what it says.

  • LLM evals
  • value alignment
  • self-modification
  • behavioral evals

Project Orion

Advised by Mehrnaz Sabet

Research EngineerNASA project at Cornell UniversityJan 2026 – present

I trained a diffusion-based motion planner to replace ORCA for real-time multi-drone collision avoidance: a DiT generating velocity command sequences from ego and neighbor states, with DPM-Solver++ inference optimization and low-temperature reranking. It reaches ~80% mean scenario completion against ORCA's 59%.

Around it I built the deconfliction stack: a ROS 2 planner plugin, scenario-based evaluation pipelines with behavioral analytics, and a ~200k-trajectory training dataset spanning six collision scenarios.

  • diffusion models
  • DiT
  • ROS 2
  • multi-agent planning
  • robotics

SMARAG

Multi-agent RAG for automated carbon reporting

ResearcherCornell University · Prof. Oliver Gao Lab, with Dr. Xinlai LiuJan 2026 – present

I did the engineering and implementation work on the multi-agent RAG system for automated ESG carbon reporting: LangChain orchestration, RL for factual grounding.

I also designed the agent observability layer: audit-trail logging and behavioral guardrails that catch hallucinated citations before they land in a regulatory filing.

  • multi-agent systems
  • RAG
  • LangChain
  • RL
  • agent observability

CAMBRIA

AI Safety FellowCambridge Boston Alignment InitiativeMay – June 2026

Worked through the ARENA mechanistic interpretability curriculum at CBAI's intensive: linear probes, steering vectors, activation patching, sparse autoencoders, OthelloGPT, emergent misalignment, alignment faking, and LLM evals. The NLA steering work came out of this program.

  • ARENA
  • mechanistic interpretability
  • SAEs
  • activation patching

AI Safety Camp

Research FellowRemoteDec 2025 – April 2026

Evaluated hierarchical and parallel AI control protocols on the ControlArena benchmark, measuring how the safety-usefulness tradeoff moves as the capability gap between trusted and untrusted models widens.

Extended ControlArena with red-team/blue-team protocols and Elo-based scoring to find the point where oversight mechanisms fail under adversarial prompting.

  • AI control
  • ControlArena
  • adversarial evals
  • oversight

Foresee Health

Software Engineering InternSan Francisco, CASept 2024 – Jan 2025

Built the authentication system for the invoice application on AWS Amplify and Lambda, and designed the PostgreSQL schemas and synchronization patterns behind a microservices architecture handling large customer datasets.

  • AWS
  • Lambda
  • PostgreSQL
  • microservices

Selected side projects

  • Emergent misalignment: reproduction, probing & activation oracles

    Fine-tuned Qwen2.5 on a small set of harmful financial advice examples and reproduced emergent misalignment. Implemented a LoRA layer from scratch and verified numerical equivalence against HuggingFace PEFT, then trained linear probes on intermediate LoRA activations to detect misalignment before generation. Also ran activation oracles decoding residual-stream activations back into natural language.

  • Variable-input PatchTST for time series forecasting

    code ↗

    Extended PatchTST to handle variable-length input windows for multivariate time series forecasting, without retraining separate models per input length.

  • LLM jailbreak defense evaluation

    GCG and AutoDAN attacks run against layered defenses: self-reminders, hierarchical prompting, perplexity filters, with guard-model architectures compared by attack success rate.

  • Domain-adaptive LLM embedding benchmark & fine-tuning

    Generated adversarial hard-negative passages via LLM prompting as contrastive training data, benchmarked 5+ embedding models, and introduced a robustness metric for near-miss retrieval failures.

  • Wisconsin school analytics on GCP

    Cloud-native analytics pipeline integrating geospatial and public-school datasets across 400+ schools, with automated Parquet ingestion via PyArrow and parameterized SQL on remote VMs.

  • Real-time weather stream processor

    Fault-tolerant Kafka to HDFS pipeline with exactly-once semantics, checkpointing, atomic Parquet writes, and stateful recovery across broker and node failures. Containerized with Docker Compose.

more on github ↗

03skills

Languages

  • Python
  • Java
  • C/C++
  • JavaScript
  • TypeScript
  • SQL
  • CUDA

AI / ML

  • PyTorch
  • HuggingFace
  • mechanistic interpretability
  • probes
  • steering vectors
  • SAEs
  • activation patching
  • fine-tuning (LoRA/PEFT)
  • evals
  • RAG
  • diffusion models
  • LangChain
  • MCP
  • inference optimization

Tools & infra

  • Docker
  • Kubernetes
  • AWS (S3, Lambda)
  • GCP
  • ROS 2
  • PostgreSQL
  • Kafka
  • Spark
  • React
  • Node.js

04beyond the terminal

Outside of Work

I make films and have a passion for visual storytelling. It's the same instinct as good research communication: deciding what someone needs to see, and in what order, before they'll believe the thing you're telling them.

watch on youtube ↗

The rest of my time outside work is spent outside. Mountain biking, off-roading, anything outdoors.

Uday in a full-face helmet and goggles, sitting on a mountain bike at the base of a mountain resort village.
Trail day
Uday standing beside a Polaris RZR off-road vehicle on a rocky ridge, with bare mountain slopes behind him.
Above the treeline