Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out

Researchers at Armature have conducted an extensive analysis of 17,000 coding agent runs to understand how AI models like Claude, Codex, and Cursor interact with development tools. The study aims to demystify the 'black box' behavior of autonomous coding agents by tracking which command-line tools, libraries, and utilities these models prioritize when tasked with software development workflows. By measuring the frequency and success rates of various tool invocations, the report provides developers and engineers with empirical data on the current capabilities and preferences of leading AI coding assistants. The findings highlight significant patterns in how these models navigate file systems, execute tests, and manage dependencies. This data-driven approach offers valuable insights for those looking to optimize their development environments for AI-assisted workflows, ultimately helping to bridge the gap between human intent and machine execution in modern software engineering.
This is a summary. Read the full article at the original source:
Hacker News (YC)Related stories
Meta has unveiled its latest innovation, the Muse AI agent, designed to act as a highly personalized assistant for users. According to recent reports,…
What's going on with OpenAI and the Navier-Stokes controversy?
OpenAI has recently claimed a significant breakthrough in mathematics, specifically regarding the Navier-Stokes equations, which describe the motion o…
Large language models develop novel social biases through adaptive exploration
A recent research paper published on OpenReview explores how large language models (LLMs) can acquire and manifest new social biases during the proces…


