All services

What I do

AI Integration

Getting AI out of the demo and into a product you can actually trust.

I build production AI: retrieval-augmented systems, agents and the pipelines behind them, with the evaluation and guardrails that make them safe to ship. The interesting part isn't the model, it's everything around it that decides whether you can trust the output.

What I bring

  • RAG systems and agentic workflows built to be grounded, not just fluent.
  • Evaluation as a first-class part of delivery: golden datasets, LLM-as-judge, regression gates in CI.
  • Guardrails and observability: schema validation, entitlement checks, tracing and cost tracking.
  • Local and cloud inference, chosen on the real economics rather than the headline number.
  • Working AI-first: I build through agents and hold the quality bar through review and evals.

How I work

My rule is simple: don't ship an AI feature you can't measure. I put a scored gate between a model's output and the decision it drives, so a change can't silently make things worse. It's the same discipline whether the system writes prose or raises a pull request. I've open-sourced the reference implementation of it as agent-eval-starter.

A good fit when

  • You have an AI prototype that works in the demo but you can't trust it in production.
  • You're adopting AI in engineering and want it done with measurement, not vibes.
  • You want someone who ships AI and can prove it's actually working.