$0.47 per run and about four minutes per task marked Asana’s optimized browser-agent workflow on GPT-6.1 Sol, down from at least $36.21 and 22.5 minutes on its original Model B setup.
A 144-run study found the gains came from extending caching to browsing history, raising the history budget to 480,000 characters, and trimming screenshots in batches instead of nearly every step.
On GPT-6.1 Sol alone, the new policy cut costs 4x from $1.97 to $0.47 because 89% of input tokens were served from cache at 5% of the uncached price.
Larger history also improved reliability: GPT-6.1 Sol produced correct answers in all 18 runs with the bigger budget, versus three of 18 with the smaller one.
GPT-6 Astra in Codex ran the experiments in about a week instead of an estimated one to two months, and Asana has already shipped the changes in StackAI while expanding agent-led product testing.