OpenAI published a GPT-6 model guide on 2 October that addresses an everyday problem for agent builders: choosing a model and judging whether the complete task works. This is operational guidance, rather than a separate agent product launch.
What the guide recommends
OpenAI recommends testing representative tasks and measuring success, latency and cost per successful task before deployment. It also discusses caching shared context and compacting longer conversations to manage context and spending.
For long-running work, the guide covers asynchronous tools, changes to instructions during a run and delegation of independent tasks. It asks developers to define the expected result, what the agent can do independently and when it should request input.
Model choice and reasoning effort should fit the workload, according to the guide. Those recommendations come from the model provider. They do not establish which model will perform best on your particular tasks.
Our assessment
For an agent, the unit you need to measure is the completed job. A cheap individual response may lead to an expensive workflow if the agent repeatedly retries or produces work someone must redo.
Consider an agent that reads a support request and prepares a proposed account update. A useful evaluation would check whether it chose the correct account, used current evidence and stopped at the agreed approval point. Record total runtime and spending across every step, including unsuccessful attempts.
Use the same tasks and acceptance criteria when comparing models. Keep the tool permissions and source material consistent, then inspect failures individually. These are suggested evaluation steps from Agentic Horizon; we have not benchmarked the models covered in the guide.
Source: OpenAI’s original model guide.
Read the original material
OpenAI (opens in a new tab)The article date belongs to Agentic Horizon. The source publication date is listed separately. Vendor claims remain attributed; we have not independently tested the reported capability.

