the map — one forward pass, top to bottom
input
hidden layer 1 · education
hidden layer 2 · experience
hidden layer 3 · essays
hidden layer 4 · research
output layer
H = −Σ p log2 p
The Hitchhiker's Guide to Evals. How to tell whether your AI product actually works, past vibes-based QA.
Cognitive bias benchmark for LLMs grounded in Kahneman-Tversky research. Ran 6 frontier models through 15 core biases. Opus 4.6 least susceptible at 11%.
Production AI safety pipeline for Waves AI. Constitutional AI patterns, crisis response testing.
Deceptive alignment detection in small models. Published on PyPI. Kaggle competition entry.
GPT from scratch in PyTorch. Multi-head attention, transformer blocks, text generation.
Benchmark testing linguistic relativity in LLMs. Do models "think" differently in different languages?
Custom plugins for Claude Code CLI. Workflow automation, code review utilities.