utilities
agent-evaluation
Evaluate AI agent tools against three structural dimensions (Persistent Memory, Inspectable Surfaces, Compounding Context) and generate calibrated delegation specs.
Target Audience: Founders, operators, and team leads evaluating AI agent tools for knowledge work delegation.
What it needs
- tool-name — Name of the agent tool being evaluated
- task-description — Specific task or workflow intended for delegation
- quality-definition — What 'good' looks like for this task
- tools — Comma-separated list of tool names to compare
- task-description — The task all tools will be evaluated against
What you get
- comparison (markdown)
- scorecard (markdown)
Ask it like this
- evaluate this agent
- should I use this tool
- agent comparison
- delegation spec
- is this tool good enough
