Designing Tools an Agent Can Actually Use
Agents rarely fail because they cannot reason. They fail because the tools you handed them are ambiguous, chatty, or unsafe.

Tool design is the highest-leverage, least-glamorous part of agentic work. A model that picks the wrong tool is usually reading a bad tool description, not thinking badly.
Name tools after intent, not implementation
`postgres_query` invites the agent to improvise SQL. `find_orders_by_customer` tells it exactly when the tool applies and constrains what can go wrong. Narrow tools with obvious names outperform one general tool with a long manual.
Write descriptions for a new hire
- What the tool does, in one sentence.
- When to use it — and one line on when not to.
- Every parameter with units, format, and an example value.
- What comes back on success, and the common failure cases.
Return small, structured, honest results
Dumping a 4,000-token API response into the loop pushes out the goal and the plan. Project the result down to the fields the agent needs, and say plainly when something is missing rather than returning an empty object.
{
"ok": true,
"count": 3,
"orders": [
{ "id": "A-1029", "status": "shipped", "total_inr": 4200 }
],
"note": "Results truncated to 3 of 17. Narrow the date range."
}Separate reads from writes
Reads can be free and retryable. Writes should be idempotent, require an explicit confirmation argument, and be logged with the reasoning that led to them. Never let a single tool both read and mutate — you lose the ability to run an agent in dry-run mode.
Every irreversible action an agent can take is a product decision, not a technical one.
Test tools without the model
Each tool should have plain tests: valid input, invalid input, empty result, upstream timeout. If a tool is flaky on its own, the agent will look unreliable and you will spend a week blaming the prompt.


