[research]By ByteBulletin Editor
Study finds harness choice barely moves agentic coding scores
A contamination-controlled benchmark of 256 tasks shows that swapping the agent framework around the same model yields statistically indistinguishable results, while cost per solved task varies significantly.
