bm.deuts.org
August 7, 2026
00:15
Kimi K3 Agent Harness Benchmark — 8 Harnesses, 25 Tasks
A controlled benchmark comparing eight agent harnesses—Oh My Pi, Kimi Code, Hermes Agent, Claude Code, Pi Agent, OpenCode, Grok Build, and Codex—using the same Kimi K3 model via OpenRouter, identical tools, and 25 complex business-app tasks. Results show a 20-point pass-rate spread, major differences in speed and cost, and that more tool calls did not guarantee better outcomes.