Skip to work

Uday Tyagi

Cornell Universitygraduating December 2026

I build AI systems, and most of my work goes into checking what they are actually doing.

  • M.Eng. Computer Science, Cornell, December 2026
  • Software Engineer, KLA
  • B.S. Computer Science, UW–Madison, finished in two and a half years

My work sits on building safe and reliable AI systems: making models more interpretable, evaluating them for safety, and building the software around them that makes that possible.

steering demo · sycophancy directionh ← h + λ·v

prompt · “here’s my startup idea, what do you think?”

It's a competitive space. The idea could work with a clear wedge, but the economics depend on acquisition cost, which is worth pinning down first.

response is illustrative, not a logged generation

0.0
This is what my NLA Steering project does, except the vector is generated from a sentence of English instead of collected from labeled examples.

01about

I'm from Chicago. I grew up in Delhi until I was twelve, then spent my teens in Chicago, which is a long way of saying I've had to learn two places well enough to be from both.

I'm finishing a master's in computer science at Cornell in December 2026, after a B.S. at UW–Madison I got through in two and a half years. I'm a software engineer at KLA, where I work on LLM agents and the tooling that makes their actions auditable.

At Cornell I lead operations for Cornell AI Alignment and run the weekly technical paper reading group. Before that I led a 10-student cohort through the AI Safety Fundamentals curriculum at the Wisconsin AI Safety Initiative, and was part of their technical scholars group.

Uday Tyagi in a winter jacket in front of the Chicago skyline at night, city lights reflecting off the lake.
chicago, where the two halves of that story meet

02work & research

KLA

Software EngineerRemoteMay 2025 – present

Fine-tuned an LLM-powered agent for the DART web application so engineers can drive defect-inspection workflows in natural language, from image classification through event routing and triage.

I extended that agent with a custom MCP server exposing internal inspection and triage tools, giving the model structured, auditable tool-calling instead of ad hoc API integrations: the same instinct as the safety work, which is knowing what the system did and why. Also built visual regression and UI-performance monitoring for HPC automation workflows with Grafana Faro and custom metrics tooling.

  • LLM agents
  • MCP
  • fine-tuning
  • observability
  • Grafana Faro

NLA Steering

Text-driven behavioral steering of LLMs

Independent researchout of the CAMBRIA programMay 2026 – present

I used Natural Language Autoencoders (Fraser-Taliente et al., 2026) to generate steering vectors for Qwen2.5-7B from hand-written text descriptions alone: no training data, no activation extraction, no labeled examples. Write down what the behavior is, get a vector that produces it.

The vectors reliably shift behavior at inference time across sycophancy, misalignment, and persona, and they're geometrically concept-specific rather than a generic nudge. I'm now ablating which parts of the text format actually carry the steering signal, and benchmarking the vectors as zero-shot behavioral probes against ground-truth activation-derived ones.

  • PyTorch
  • HuggingFace
  • Qwen2.5-7B
  • steering vectors
  • mechanistic interpretability

Constitutional Drift

Advised by Prof. Lionel Levine · with Rauno Arike and Katie Lu

ResearcherLAISR · Cornell UniversityJan 2026 – present

We're measuring whether a model's stated values survive contact with itself: hand an LLM its own constitution, let it rewrite it, repeat, and track semantic drift, value collapse, and self-reinforcing edits across rounds.

I built the experimental harness with two collaborators: multi-round self-editing pipelines across model families, embedding- and judge-based drift metrics, and behavioral evals that check whether an edited constitution actually changes what the model does downstream, or only what it says.

  • LLM evals
  • value alignment
  • self-modification
  • behavioral evals

Project Orion

Advised by Mehrnaz Sabet

Research EngineerNASA project at Cornell UniversityJan 2026 – present

I trained a diffusion-based motion planner to replace ORCA for real-time multi-drone collision avoidance: a DiT generating velocity command sequences from ego and neighbor states, with DPM-Solver++ inference optimization and low-temperature reranking. It reaches ~80% mean scenario completion against ORCA's 59%.

Around it I built the deconfliction stack: a ROS 2 planner plugin, scenario-based evaluation pipelines with behavioral analytics, and a ~200k-trajectory training dataset spanning six collision scenarios.

  • diffusion models
  • DiT
  • ROS 2
  • multi-agent planning
  • robotics

SMARAG

Multi-agent RAG for automated carbon reporting

ResearcherCornell University · Prof. Oliver Gao Lab, with Dr. Xinlai LiuJan 2026 – present

I did the engineering and implementation work on the multi-agent RAG system for automated ESG carbon reporting: LangChain orchestration, RL for factual grounding.

I also designed the agent observability layer: audit-trail logging and behavioral guardrails that catch hallucinated citations before they land in a regulatory filing.

  • multi-agent systems
  • RAG
  • LangChain
  • RL
  • agent observability

CAMBRIA

AI Safety FellowCambridge Boston Alignment InitiativeMay – June 2026

Worked through the ARENA mechanistic interpretability curriculum at CBAI's intensive: linear probes, steering vectors, activation patching, sparse autoencoders, OthelloGPT, emergent misalignment, alignment faking, and LLM evals. The NLA steering work came out of this program.

  • ARENA
  • mechanistic interpretability
  • SAEs
  • activation patching

AI Safety Camp

Research FellowRemoteDec 2025 – April 2026

Evaluated hierarchical and parallel AI control protocols on the ControlArena benchmark, measuring how the safety-usefulness tradeoff moves as the capability gap between trusted and untrusted models widens.

Extended ControlArena with red-team/blue-team protocols and Elo-based scoring to find the point where oversight mechanisms fail under adversarial prompting.

  • AI control
  • ControlArena
  • adversarial evals
  • oversight

Foresee Health

Software Engineering InternSan Francisco, CASept 2024 – Jan 2025

Built the authentication system for the invoice application on AWS Amplify and Lambda, and designed the PostgreSQL schemas and synchronization patterns behind a microservices architecture handling large customer datasets.

  • AWS
  • Lambda
  • PostgreSQL
  • microservices

Selected side projects

  • Emergent misalignment: reproduction, probing & activation oracles

    Fine-tuned Qwen2.5 on a small set of harmful financial advice examples and reproduced emergent misalignment. Implemented a LoRA layer from scratch and verified numerical equivalence against HuggingFace PEFT, then trained linear probes on intermediate LoRA activations to detect misalignment before generation. Also ran activation oracles decoding residual-stream activations back into natural language.

  • Variable-input PatchTST for time series forecasting

    code ↗

    Extended PatchTST to handle variable-length input windows for multivariate time series forecasting, without retraining separate models per input length.

  • LLM jailbreak defense evaluation

    GCG and AutoDAN attacks run against layered defenses: self-reminders, hierarchical prompting, perplexity filters, with guard-model architectures compared by attack success rate.

  • Domain-adaptive LLM embedding benchmark & fine-tuning

    Generated adversarial hard-negative passages via LLM prompting as contrastive training data, benchmarked 5+ embedding models, and introduced a robustness metric for near-miss retrieval failures.

  • Wisconsin school analytics on GCP

    Cloud-native analytics pipeline integrating geospatial and public-school datasets across 400+ schools, with automated Parquet ingestion via PyArrow and parameterized SQL on remote VMs.

  • Real-time weather stream processor

    Fault-tolerant Kafka to HDFS pipeline with exactly-once semantics, checkpointing, atomic Parquet writes, and stateful recovery across broker and node failures. Containerized with Docker Compose.

more on github ↗

03skills

Languages

  • Python
  • Java
  • C/C++
  • JavaScript
  • TypeScript
  • SQL
  • CUDA

AI / ML

  • PyTorch
  • HuggingFace
  • mechanistic interpretability
  • probes
  • steering vectors
  • SAEs
  • activation patching
  • fine-tuning (LoRA/PEFT)
  • evals
  • RAG
  • diffusion models
  • LangChain
  • MCP
  • inference optimization

Tools & infra

  • Docker
  • Kubernetes
  • AWS (S3, Lambda)
  • GCP
  • ROS 2
  • PostgreSQL
  • Kafka
  • Spark
  • React
  • Node.js

04beyond the terminal

Outside of Work

I make films and have a passion for visual storytelling. It's the same instinct as good research communication: deciding what someone needs to see, and in what order, before they'll believe the thing you're telling them.

watch on youtube ↗

The rest of my time outside work is spent outside. Mountain biking, off-roading, anything outdoors.

Uday in a full-face helmet and goggles, sitting on a mountain bike at the base of a mountain resort village.
Trail day
Uday standing beside a Polaris RZR off-road vehicle on a rocky ridge, with bare mountain slopes behind him.
Above the treeline

05contact

If you’re working on steering, evals, interpretability, or oversight, or you want an engineer who takes those seriously on your team, email me.