Rubrics for evaluating AI agents
A repeatable, rubric-based way to evaluate how AI agents behave when they act on real systems through MCP tools — consistent scores instead of one-off manual checks.
Things I have built, tested or written about. More coming soon.
A repeatable, rubric-based way to evaluate how AI agents behave when they act on real systems through MCP tools — consistent scores instead of one-off manual checks.
Using Distrobox containers on an immutable Linux desktop to keep every toolchain isolated, disposable and reproducible.
This site: a Jekyll landing page plus a Chirpy blog, built in one pipeline and served from Cloudflare's edge, mirrored on GitHub Pages.