โ† AI library

Editorial recipe โ€” not benchmarked โ€” reviewed September 3, 2026

Compare open-source coding agents on your own repository task

For engineering teams choosing among provider-neutral or locally operated coding harnesses.

Stack
Fixed repository fixture โ†’ Qwen Code โ†’ Goose โ†’ OpenCode โ†’ Executable verifier โ†’ Blind review

Run this recipe

Work through it here

Progress stays in this browser. The downloaded Markdown kit works in any notes app or repository.

0/5 complete

Is this recipe useful?

Procedure

  1. Choose one representative task, freeze the repository fixture and prompt, and define hidden acceptance checks plus a human artifact rubric.
  2. Pin each harness, model endpoint, instructions, permissions, tools, context budget, and optional add-on; change only the configuration being compared.
  3. Run every configuration from a fresh copy at least three times, capturing exit state, interventions, elapsed time, tokens, cost, commands, and changed files.
  4. Run the same verifier on every artifact and blind the harness labels before reviewing correctness, maintainability, security, and unnecessary scope.
  5. Publish raw artifacts and failures, then select a harness for this job and environment rather than declaring a universal winner.

Acceptance artifact

A reproducible fixture, pinned run manifests, raw outputs, verifier results, blind scores, and job-specific selection

Do not use it blindly

A harness license does not make every model, extension, or provider open source; compare exact configurations and keep credentials out of published artifacts.

Evidence and setup