AI engineer and full stack developer

LLM systems
that hold up
in production.

I build agents, guardrails, evals, and computer vision pipelines, then sit with the people using them until the thing actually works. Most of my time goes into the part after the demo.

Portrait of Suyash Bhattarai
  • World 2nd in CS Official Cambridge Recognition
  • Top in country in CS Outstanding Cambridge Learner Awards
  • 1540 SAT 99th percentile globally
  • A*s in CS and Maths Cambridge International A Levels
  • Best Academic Performance Award Budhanilkantha School
  • CAIE AS Level Topper Budhanilkantha School
  • As in Economics, Business and EGP Cambridge International A Levels

Selected work

Agents, infrastructure, and applied ML.

01 Live demo

Open Switch

LLM gateway with circuit breakers, half-open probes, hedged requests, cost-aware routing, per-key budgets, and request traces.

  • Rust
  • Tokio
  • Axum
  • WASM
Live demo
02 Live demo

Screener

Typo tolerant hybrid search with faceted filters that runs fully in the browser.

  • Rust
  • BM25
  • FST
  • WASM
Live demo
03 Live demo

Evalgate

LLM evaluation runner with its own YAML check format, flakiness detection, and a pull request gate that writes a Markdown report.

  • Rust
  • CLI
  • YAML
  • WASM
Live demo
04 Live demo

Corridor FX

FX exposure and hedging engine with a quantile forecast that is honest about what it cannot predict.

  • Python
  • XGBoost
  • FastAPI
  • React
Live demo
05 Case study

Flamide

Multilingual AI agent that answers student enquiries, fills a CRM, and hands off decisions to a counselor.

  • Next.js
  • TypeScript
  • Supabase
  • DeepSeek
06 Case study

Sekkin

Schema-first code generator where deterministic structure keeps the LLM focused on application logic.

  • Python
  • FastAPI
  • SQLAlchemy
  • React Flow

How I work

A green test suite is a claim, not a proof.

01

Run it live

My worst bugs passed every mocked test. I run the real provider, unintercepted, before I call anything done.

02

Guardrails outside the model

If a rule matters, code checks it after generation. Prompts ask nicely. Validators refuse.

03

Say what wasn't measured

Every report I write lists what it can't claim. A 0.55 AUC gets reported as a coin flip, not rounded up.

04

Sit with the users

I've shipped to factory floors, school offices, and police briefings. The spec is whatever they actually need on day two.

Contact

Let's build something
that holds up.