Methodology โ first public run pending
Evaluate a coding harness and model combination
For teams choosing an agent for repeated engineering work.
Run this recipe
Work through it here
Progress stays in this browser. The downloaded Markdown kit works in any notes app or repository.
0/5 complete
Procedure
- Select representative tasks with hidden acceptance tests and resettable repositories.
- Pin model, harness, instructions, permissions, tools, add-ons, and budget.
- Run each configuration on identical fresh fixtures; repeat nondeterministic tasks.
- Blind-review correctness, regressions, intervention, elapsed time, usage cost, and diff size.
- Publish artifacts and failures; choose a setup per job instead of declaring one universal winner.
Acceptance artifact
A versioned run manifest, raw artifacts, and dimension-by-dimension results
Do not use it blindly
Do not compare configurations on different prompts, budgets, or repository states.