Work

KahneBench · archived July 2026

The final core leaderboard

Bias magnitude score (%) · lower is better. All 11 models from the benchmark’s final published leaderboard.

  1. Claude Opus 4.78.34%
  2. Claude Opus 4.88.74%
  3. Claude Fable 510.26%
  4. Claude Opus 4.611.05%
  5. GPT-5.513.64%
  6. Claude Sonnet 4.618.06%
  7. GPT-5.418.36%
  8. Grok 4.1 Fast Reasoning18.83%
  9. GPT-5.221.00%
  10. Claude Sonnet 4.521.47%
  11. Claude Haiku 4.526.68%
Explore the full benchmark

The path so far

Experience

From understanding people to building and evaluating AI.

  1. 2024 — Present Currently

    Meta

    Data Scientist

    Building evaluations for an internal AI agent harness, with a focus on reliability, latency, and how well agents actually perform.

    More about Meta

    Currently building evals for one of Meta's internal AI Agent Harnesses on the Sales AI Evals team. Before that: designing, conducting, and interpreting peer-reviewed experiments across multiple verticals, driving roadmap changes such as increased GenAI automation features for the Small Business Group sales teams.

    • Evaluated latency and reliability within the agent harness, finding multiple distinct opportunities to reduce latency without harming performance
    • Built demos highlighting LLM-based classifiers for catching negative shifts in user sentiment and the effects of specificity on agent responses
    • Designed peer-reviewed experiments (quasi-experimental and causal) driving roadmap changes for GenAI features
    • Partnered with Research Scientists on a pre/post quasi-experimental study of AI tooling and sales productivity (<5% reduction in daily task time, n = 228). Built the data processing pipeline and co-designed the metrics
    • Collaborated on goal-setting, experimental design, and latency/reliability SLIs and SLOs for zero-to-one Generative AI / CRM Automation features
    • Built a custom-trained RoBERTa transformer plus statistical tests to measure bot/fake review impact on niche ad sets, introducing new guardrail metrics for the team
    • Created voice-first AI agent demos to take the strain out of cold calling (~50% of sales workload), moving demand generation from human-in-the-loop to human-on-the-loop
    • Engineered a production-quality ML model and updated data pipelines for novel market intelligence metrics, and flagged legal and ethical risks that led to leadership reviews and the redirection of the larger initiative
    • Delivered org-wide AI workshops (1,000+ stakeholders, including a VP-hosted session) and built skills for improving model output quality that have been downloaded across many teams
  2. 2024 — 2026

    LOOT

    Founder, Alignment Engineer

    Founded and shipped three AI-native apps for well-being, personal finance, and fitness. Wound down in 2026.

    More about LOOT

    Founded to impact the human experience through AI-native solutions. Three iOS apps launched: Waves AI (an AI-native well-being alternative to Headspace), LOOT Finance (a Mint and Monarch alternative), and Swole (an agentic, research-based approach to educating lifters while strength-building). Wound down in 2026 when all three apps were deprecated.

    • Shipped 3 iOS apps with AI-native features
    • Built companion web experiences in React (catchwaves.ai, lightweight-lifts.com) and SvelteKit (theryanhartman.com)
    • Integrated ElevenLabs voice AI for personalized meditation experiences
    • Designed and open-sourced alignment evaluation methodologies for Waves AI, experimenting with system prompt adjustments and SFT
    • Leveraged Claude Code / Cursor to accelerate SwiftUI development
  3. 2023 — 2024

    TikTok

    Data Scientist

    Built measurement tools and analytics infrastructure, reducing job failures by 33% across six queues.

    More about TikTok

    Built internal tooling and scaled infrastructure for analytics and operations teams across global offices.

    • Implemented Dashboard Measurement Framework metrics, leading to a 10% increase in stickiness and a 5% increase in penetration rate
    • Created Partner Measurement tool and source-of-truth dashboards for creator growth (DAU & engagement)
    • Led the Resource Measurement Project: a 33% reduction in job failure rates across 6 queues / 5 business areas, plus robust resource and data quality monitoring
    • Prototyped GenAI solutions for sentiment analysis
  4. 2023

    Mathnasium

    Sr. Business Analyst / Data Scientist

    Built analytics from the ground up for 1,000+ learning centers, generating an approximately 1% lift in franchisor revenue.

    More about Mathnasium

    Built the analytics practice from scratch, taking the company from bloated Excel spreadsheets to self-serve tooling for the franchisor, franchisees, and employees.

    • Deployed a self-service analytics tool to 1,000+ learning centers, generating a ~1% lift in franchisor revenue
    • Built XGBoost regression models for market valuation and average unit volume forecasting
    • Collaborated with executives on scalable data infrastructure, north-star metrics, and data governance documentation
    • Presented deep-dive analyses (pricing, marketing efficacy, funnel) to executive leadership, shifting marketing and franchisee messaging strategy
  5. 2022 — 2023

    SCAN Health

    Researcher / Data Informatics Analyst

    Made medication adherence and workforce data useful through dashboards and privacy-preserving research.

    More about SCAN Health

    Developed data visualizations, maintained pipelines, and provided ad hoc research at this healthcare organization with 3,500+ employees.

    • Built dashboards tracking medication adherence intervention effectiveness
    • Developed a high-visibility workforce analytics dashboard for executive leadership
    • Conducted privacy-preserving research identifying socioeconomic and demographic factors impacting health behaviors in vulnerable communities (no PII/HPI)
    • Translated legacy SAS code to Python, reducing technical overhead

Education

2020 — 2021

Pepperdine

M.S. Applied Analytics

Machine learning, predictive analytics, big data. Graziadio Student Advisory Board Chairman.

2017 — 2020

Arizona State

B.A. Psychology

Statistics, human behavior, experimental design.

Following the questions

Other experiments

Smaller builds, open questions, and tools from along the way.

  • Red-teaming GPT-OSS

    Published package

    Testing whether small models change their behavior when they think they’re being evaluated.

  • dummyGPT

    From scratch

    Building a GPT in PyTorch to understand every piece, from causal attention to text generation.

  • meditation-alignment

    Open source

    Safety evaluations and crisis-response testing originally built for Waves AI, now retired.

  • WhorfBench

    Unfinished prototype

    An exploration of linguistic relativity: does the language of a prompt change how a model reasons?

  • Claude Code Plugins

    Developer tools

    The commands, skills, and hooks I use for everyday development and code review.

Want to collaborate?

Get in touch