AgentBench Cost Lab
An open benchmarking harness that scores Cursor Auto vs Claude Code vs Copilot on a team's own codebase so engineering leads can pick one subscription rather than stack them.
Tech leads at small dev shops deciding whether to consolidate their AI coding subscriptions
- Drop in a repo, run identical coding tasks across Cursor Auto, Claude Sonnet, GPT-4.1 and rank by pass-rate and latency
- Generates a 'marginal seat' report: how much extra throughput does adding a second tool give?
- Exports a one-page memo leadership can use to justify dropping a subscription
The Cursor Bridge trend is a symptom of developers already running A/B tests on their AI tool stack in production; this gives them a structured way to do it.
Abundant free comparison content exists (artificialanalysis.ai, paperclipped.de, future-of-software.com) showing real interest in picking among Cursor/Claude Code/Copilot, but most demand is satisfied by articles, not paid harnesses.Coding Agents Comparison: Cursor, Claude Code, GitHub Copilot, and more ↗AI Coding Assistants Compared 2026: Cursor vs Claude Code vs Copilot vs ↗
Public open benchmarks already cover this space (Vexp-swe-bench, codesota.com SWE-bench dashboard, jimmc414's swebench harness), so the 'score the candidates' idea is largely taken; only the private-codebase twist is differentiated.GitHub - Vexp-ai/vexp-swe-bench: Open benchmark for AI coding agents on ↗SWE-bench 2026: Compare Devin, Codex, Claude Code, Cursor, OpenHands ↗
Free open-source harnesses and free comparison articles dominate; no evidence of willingness to pay for a benchmarking product when eng leads can read a review or run SWE-bench themselves.GitHub - jimmc414/claudecode_gemini_and_codex_swebench ↗Best AI Coding Assistants 2026: Cursor vs Copilot vs Claude Code ↗
AI coding model capabilities shift every few months (e.g., new Claude/GPT/Cursor versions), so any single benchmark loses relevance quickly; the Cursor Bridge trend itself is a workaround for a pricing model that may not last.Cursor Bridge: Use Cursor in Claude Code or Codex! ↗
Building a harness is technically straightforward given existing open-source templates (Vexp-swe-bench) and public APIs, but differentiation against those free projects is the harder part.GitHub - Vexp-ai/vexp-swe-bench: Open benchmark for AI coding agents on ↗