TAU-HOME.COM
LOADING

Cutting Agent Harness Costs by Up to 3x with Lean Prompts and Tool Schemas

Optimizing the system prompt and tool schema sizes in an agent harness can reduce execution costs by up to 3x without degrading benchmark accuracy when using th

tau · October 9, 2026

#AI-Agents #Agent-Harness #Prompt-Optimization #Cost-Reduction #SWE-bench

Cutting Agent Harness Costs by Up to 3x with Lean Prompts and Tool Schemas

AI research and education community DAIR.AI (@dair_ai) highlighted findings from a recent study investigating how agent harness design affects benchmark accuracy and operating cost when holding the underlying foundation model constant.

Performance and cost comparison chart across agent harnesses on SWE-bench Verified

Image source: DAIR.AI (@dair_ai)

The central takeaway is straightforward: keeping the harness system prompt and tool schemas concise can cut token expenditure by up to 3x without lowering model accuracy.

Performance Variance Across Agent Harnesses

The researchers evaluated three major coding agent harnesses—Claude Code, mini-SWE-agent, and OpenCode—using the same underlying LLM across the 447 tasks in the SWE-bench Verified benchmark.

  • 447-Task Evaluation: Claude Code and mini-SWE-agent finished within 5 points of each other across all 447 benchmark tasks.
  • Harness Swap vs. Rerun Noise: Switching between different harnesses produced performance variations comparable to simply rerunning the same harness multiple times.
  • Hard Subset Behavior: On a challenging 45-task subset, both conditions—swapping the harness and rerunning the same harness—flipped 13% of tasks.

These empirical results show that under identical model conditions, harness architecture differences contribute far less to overall outcome variance than the cost overhead introduced by bloated prompts and schema definitions.

The Root Cause of Cost Disparity: Per-Step Payload Accumulation

The up to 3x cost disparity between harnesses stems directly from how system prompts and tool schemas are transmitted across execution turns.

Most agent harnesses retransmit the entire system prompt and every declared tool schema payload at each intermediate step in the loop. As the agent takes additional reasoning steps and tool calls to resolve a task, the token cost of that static baseline payload multiplies linearly with step count.

When harnesses carry verbose instructions or expansive JSON schemas for unused tools, token consumption rises sharply without delivering measurable accuracy gains.

Practical Recommendations for Lean Agent Harnesses

  1. Trim Tool Schemas: Keep tool descriptions and parameter JSON Schemas concise, restricting definitions strictly to necessary arguments and essential fields.
  2. Streamline System Instructions: Remove non-essential role definitions, edge-case prose, and verbose prompt scaffolding from the harness baseline.
  3. Control Multi-Turn Step Overhead: Actively audit repeated prompt payloads per iteration to prevent unnecessary token accumulation across multi-step tasks.

Original source