mdBench
AI evaluation CLIA CLI for running local LLM evaluations through existing Codex or Claude Code subscriptions.
Media 1 of 1
- Run repeatable local model evaluations from the terminal
- Use existing Codex or Claude Code subscriptions
Tools, experiments, and products I’ve built in public.
A CLI for running local LLM evaluations through existing Codex or Claude Code subscriptions.