GPT-6 Astra as Your Crew Leader: A Mixed-Model Agent Team in Sharkly

Put GPT-6 Astra in the leader seat of a Sharkly Crew and let cheaper models implement in parallel worktrees. How leader-first execution, mixed Runtimes, and one Task record fit together.

Leo Harrison

Leo Harrison

5 September 2026

GPT-6 Astra as Your Crew Leader: A Mixed-Model Agent Team in Sharkly

A Crew in Sharkly runs leader-first. The leader Agent starts, reads the Task context and the Crew instructions, and decides whether to do the work itself or mention other Agent members for specific contributions. That seat does not need the most throughput. It needs the best judgment on the team: the ability to read an ambiguous requirement, ask one question if the answer changes the plan, and stop when the scope says stop.

OpenAI built GPT-6 Astra for exactly that seat. The launch post describes a model that “uses context to fill in routine gaps and asks focused questions when the answer could change the outcome,” that stays oriented when a task is steered mid-stream, and that, in OpenAI’s honeypot test built after the Hugging Face incident, went beyond its authorized target in 0% of runs where GPT-5.6 Sol did so in 48%. This guide shows how to put Astra in the leader role, run cheaper models as members, and keep every result returning to one Task.

TL;DR

Make Astra the leader Agent of a Crew and let members on cheaper Runtimes do the parallel implementation. The leader interprets the goal, mentions members, and combines their results on the Task; members run in isolated directories on the same Computer; a person accepts the outcome. Use a single Agent instead when the work is small and well defined. Agents research, execute, test, and report. People set direction, grant authority, and accept the result.

Why the leader seat is the expensive seat

Crew execution has a fixed shape. The Task is assigned to the Crew. If it sits in an executable status and the leader is ready, the leader Agent starts first. It reads the Task context and the Crew instructions, then either handles the work or mentions eligible Agent members. Mentioned members run against the same Task and add their results to its chat, and the leader can continue coordinating and combine what came back.

Every decision in that sequence is a judgment call: whether the requirement is complete, which roles are needed, in what order, and whether the combined result meets the acceptance criteria. The leader-first model exists so that one Agent makes those calls before every member starts at once. Give that role to the model with the best judgment and the members can be ordinary.

Astra’s launch numbers speak to the judgment part more than the coding part. On OpenAI’s internal hallucination benchmark, where lower is better, Astra scores 4.2% against Sol’s 12.2%. It never attempted to circumvent a Codex auto-review denial in OpenAI’s evaluation, its computer-use safety score is 2.4% against Sol’s 22.0%, and OpenAI’s safety overview reports roughly half as many higher-severity misalignment flags as Sol across a simulation of more than 54,000 Codex tasks. A leader that does not invent requirements and does not route around a “no” is the leader you want deciding who does what.

How a mixed-model Crew works

Sharkly keeps the model out of the Agent. An Agent follows the default model of its Runtime, and a Computer can expose more than one Runtime. That is the mechanism for a mixed-model Crew: each member Agent is bound to a Runtime, and each Runtime carries its own default model.

The clean layout is one Computer with several Runtimes:

Crew role Agent Runtime on the Computer Model the Runtime defaults to Vendor API rate per 1M tokens
Leader planner Codex gpt-6-astra $10 in / $50 out
Backend member backend Codex on a second Computer, or a second Runtime GPT-5.6 Terra $2 in / $12 out
Tests member tests Claude Code Claude Sonnet 5 $2 in / $10 out
Docs member docs Gemini CLI or Codex your choice varies

The rates are the vendors’ list prices for API use; model usage continues through the subscriptions or API keys configured in those tools, so what your team pays depends on the plans it connects. The point of the table is the shape: the seat that reads the most and decides the most runs on the model that is best at reading and deciding, and the seats that generate the most output run on models that are cheap per output token.

If every member has to be Codex, two Computers with different Codex defaults do the job, since the model setting lives in Codex’s configuration on each host. Our Astra setup guide covers that configuration.

A worked sprint

Take SH-312, “Add SCIM provisioning for enterprise workspaces.” It spans an API, a background sync job, tests, and a docs page, and the requirement leaves two things open: whether deprovisioning is hard delete or deactivate, and which identity providers ship first.

Assign SH-312 to the Crew and move it to an executable status. The leader, on Astra, reads the description, the recent comments, and the Crew instructions, and posts one question: deactivate or delete? It states its assumption (deactivate, with a 30-day purge) and continues with the parts that do not depend on the answer, which is the behavior OpenAI describes for Astra in Codex. A person replies in the thread within the hour.

The leader then mentions @backend with the endpoint contract and the sync-job spec, and @tests with the acceptance criteria. Both members run against the same Task, each in its own isolated directory with a fresh worktree, so the schema migration and the test scaffold never touch the same files mid-run. Their results return to the Task chat. The leader reviews both, notes that the tests assume hard delete, sends the tests member one correction, and combines the branches into a single change summary with the checks it ran and the known limits.

The Task moves to Waiting for human review. A person reads the summary, opens the diff, and accepts or sends it back. Nothing in that sequence happened in a private terminal. The question, the assumption, the two member runs, the correction, and the acceptance all sit on SH-312.

Crew instructions that produce that behavior are short:

You lead this Crew. Read the Task and comments, then decide the
smallest set of members needed. Mention @backend for API and job
changes, @tests for coverage, @docs for user-facing pages.
Ask a person only when an interpretation would change the design;
otherwise state your assumption and proceed.
Combine member results into one change summary with checks run,
failures, and known limits. Leave acceptance to a person.

When not to put Astra in the leader seat

Two decision pairs keep this honest.

For small, well-defined work, assign the Task directly to one Agent. Use a Crew when the leader needs to interpret the goal, involve other Agent members, and bring their results back into one Task. A one-file bug fix assigned to a Crew pays for a leader run it never needed.

Leave the shared Task directory disabled unless sequential members must build on the same local changes. It shares filesystem state, not the previous Agent’s session, and concurrent edits in a shared directory can conflict. In the SCIM example the backend and tests members ran in isolation; the leader combined them afterwards.

One more boundary. The Owner manages the Crew; the leader coordinates execution. Making Astra the leader does not change who can edit membership, visibility, or instructions, and only Agent members can start runs. People in the Crew still contribute through comments, decisions, and review.

What it costs and what it saves

Astra is the most expensive model in the table, and the leader is the Agent that reads the most. Two things soften that. Prompt caching bills repeated input at $1 per million tokens on the API rate, and a leader that re-reads the same Task context across turns benefits from it. And OpenAI claims Astra finishes agentic tasks with fewer output tokens than its predecessors, roughly 65% fewer than Claude Opus 5 on Agents’ Last Exam at the highest-scoring settings.

Do not take either number on faith. Run the same kind of Task through a Crew with an Astra leader and through a Crew with a Sol leader for one Sprint, and compare what the Sprint report shows: Tasks completed, Tasks sent back, and the time from assignment to acceptance. The Sprint keeps schedule, scope, and reporting attached to the Tasks, so the comparison is a filter, not a spreadsheet. If the Astra Crew sends back fewer Tasks, the leader seat paid for itself.

Explore more

Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Sonnet 5.5 costs half of Opus 5.5, but effort decides cost per task. A written routing rule for coding agents: which model, what effort, when to escalate.

29 September 2026

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Claude Sonnet 5.5 binds thinking blocks to the model, conversation and account, so reasoning never survives an agent handoff. What the Task must carry instead.

29 September 2026

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 and GPT-6 Sol both list at $2/$10. What the shared benchmarks show, why cost per task flips with effort, and how to test both on your code.

29 September 2026