Editorial recipe โ not benchmarked โ reviewed September 3, 2026
Compare open-source coding agents on your own repository task
For engineering teams choosing among provider-neutral or locally operated coding harnesses.
Run this recipe
Work through it here
Progress stays in this browser. The downloaded Markdown kit works in any notes app or repository.
0/5 complete
Is this recipe useful?
Procedure
- Choose one representative task, freeze the repository fixture and prompt, and define hidden acceptance checks plus a human artifact rubric.
- Pin each harness, model endpoint, instructions, permissions, tools, context budget, and optional add-on; change only the configuration being compared.
- Run every configuration from a fresh copy at least three times, capturing exit state, interventions, elapsed time, tokens, cost, commands, and changed files.
- Run the same verifier on every artifact and blind the harness labels before reviewing correctness, maintainability, security, and unnecessary scope.
- Publish raw artifacts and failures, then select a harness for this job and environment rather than declaring a universal winner.
Acceptance artifact
A reproducible fixture, pinned run manifests, raw outputs, verifier results, blind scores, and job-specific selection
Do not use it blindly
A harness license does not make every model, extension, or provider open source; compare exact configurations and keep credentials out of published artifacts.