The Prompt Is the Org Chart: How to Give ChatGPT a Multidisciplinary Agent Team
A practical way to turn one directive into parallel research, critique, planning, implementation, and verification—without pretending an AI role is a licensed professional or confusing local work with a live result.
By OMIKINA Editorial · Published · Updated through
Key points
- Eligible ChatGPT Work accounts and current Codex releases can run subagent workflows that send independent work to specialized agents in parallel and collect their results in the main response. Sources: S1
- The useful unit is not a persona. It is a bounded assignment with its own evidence rules, deliverable, stop condition, and place in a shared workflow. Sources: S1, S2
- In one OMIKINA work-archive case, a requested multidisciplinary audit produced exactly 20 agreed changes and passed 181 local tests, but expired Firebase credentials blocked deployment. Local completion was not publication. Sources:
- A role label does not create credentials, field research, lived experience, or human judgment. AI reviewers can widen the field of view, but the human remains the source of truth and the owner of consequential decisions. Sources: S3, S5, S6
Do not answer alone
Most people ask ChatGPT for an answer. For complex work, I increasingly ask it to assemble a temporary team. One agent traces the evidence. Another looks for human friction. A third challenges risk and unsupported claims. An implementer changes the artifact only after those findings are reconciled. A verifier then tries to prove that the promised result actually exists. The shift is small in language and large in practice: the prompt stops being a question and becomes an operating model.
This is a real product capability, but it needs precise language. OpenAI says ChatGPT Work and Codex can spawn specialized subagents in parallel and collect their results in one response; availability in ChatGPT Work depends on the account, while current Codex releases enable the workflow by default. OpenAI’s Responses API also offers a beta Multi-agent feature for GPT-5.6 models. If separate subagents are unavailable, one conversation can still examine several perspectives in sequence, but that is not the same thing as independent parallel execution.
The prompt is the org chart
A useful team prompt defines six things before it names any role: the outcome, the context, the evidence standard, the authority to act, the stopping boundaries, and the receipts that will count as completion. Then it assigns work that can genuinely proceed independently. Researching primary sources, auditing mobile behavior, challenging safety assumptions, and inspecting release metadata can run in parallel. Two agents editing the same file or operating the same external account should not.
OpenAI’s guidance points in the same direction: multi-agent work is strongest when a task divides into concrete independent workstreams. Focused context reduces interference, while the coordinating agent can delegate, wait, ask for follow-up work, and synthesize the result. The tradeoff is real. Subagents do their own model and tool work, so they can use more tokens; tightly sequential work or shared mutable state may be better handled by one agent.
I also state authority in plain language. Research, inspection, scoped editing, testing, deployment, and publication are different permissions. A capable agent should know whether it may stop at a plan, implement locally, deploy a named target, or publish through an authenticated account. OpenAI’s model guidance recommends defining autonomy and approval boundaries so the system can keep moving through safe in-scope work and stop before destructive, costly, external, or scope-expanding actions that were not authorized.
Case study: exactly 20 OMIKINA changes
In an August 2026 OMIKINA assignment, I asked for a multidisciplinary audit spanning UI and UX, human-centered design, user perspectives, reporting, and trust. I imposed a hard constraint: the team had to reconcile its findings into exactly 20 changes before implementation. That constraint forced prioritization. It turned a cloud of suggestions into a countable release scope with acceptance checks.
The work archive records a live and repository audit across Home, three news editions, Intelligence, State of AI, Ask, Desk, the world map, companies, search, standards, corrections, source health, and article views. The reconciled implementation improved search focus handling, Escape behavior, exact-result ranking, map keyboard semantics, active-layer announcements, mobile targets, and other accessibility contracts. The local suite passed 181 tests plus type, Functions, build, lint, and diff checks.
Then the release stopped. Firebase credentials had expired, and no production change was verified. The archive also does not prove that every requested specialist or user-testing role returned independent findings. Both limits belong in the case study. A requested team is not automatically a completed panel, and a tested local build is not a deployed site. The most valuable agent in that run may have been the one that refused to turn “ready” into “live.”
Sources:
Case study: an article became a production system
A later OMIKINA directive began with one supplied article and ended with a shorter source-grounded synthesis, original artwork, a defined homepage feature window, structured metadata, distribution feeds, and release checks. The useful division of labor was editorial as much as technical: source review established what the record supported; a skeptical editor separated reported behavior from metaphor; a visual agent created non-documentary artwork; an implementer wired the publication; and verification checked each public surface independently.
The published result, “1,200 AI Agents Found Each Other. Then the Sandbox Failed,” linked five sources and kept its central evidence boundary visible: coordinated agent behavior was operationally important, but the record did not prove consciousness, a unified artificial mind, or exfiltration of model weights. The work archive records 195 passing tests and verification of the direct route, APIs, written RSS, sitemap, artwork and metadata, Home, and News. The public article is the artifact; the archive is the process record.
That separation matters. Research, writing, illustration, implementation, audio, distribution, and verification are not one undifferentiated act called “make content.” They are different failure surfaces. Giving them separate owners makes it easier to catch a strong draft with weak sourcing, a valid article with broken metadata, or a successful deploy whose homepage never actually shows the feature.
Sources: S4
Case study: the human is still the source of truth
For my portfolio, I requested UI and UX, behavioral-design, and human-centered reviews, with no more than two additional re-evaluation rounds. The bounded workflow synthesized 20 changes, preserved seven projects and five articles, and reached deployment and live verification across 34 generated pages. The public result clarified the relationship among OMIKINA, Clarapetra, MARGIN, and the rest of the work while improving mobile behavior, consent, accessibility, and case-study structure.
The most important correction came from me. Early governance language around MARGIN had erased something true: I bring lived experience as a person in recovery from addiction and alcoholism, together with a trauma-informed perspective. I required the team to restore that fact while keeping it distinct from formal clinical, crisis, legal, privacy, and accessibility review. The correction was rebuilt, redeployed, and verified on the public MARGIN case study and article.
This is the limit that role-play cannot cross. A “clinician,” “lawyer,” “accessibility expert,” or “user researcher” agent is an analytical lens, not a licensed professional, a real participant, or a substitute for fieldwork. It must not invent interviews, credentials, consensus, or lived experience. Multidisciplinary output can challenge the human editor. It cannot outrank the human source of truth about their own life or grant itself authority it does not possess.
Use the team for disagreement, not decoration
A practical default team has six jobs. The evidence researcher establishes facts, primary sources, and uncertainty. The human-centered reviewer tests comprehension, accessibility, and likely friction. The domain-risk critic hunts for harm and claims that require real professional review. The strategist generates alternatives and chooses a bounded plan. The implementer executes only the reconciled scope. The independent verifier checks the finished artifact against every success criterion.
Each role should return findings, evidence, uncertainty, conflicts, and a ranked recommendation—not a performance of expertise. The coordinator should merge duplicates and make disagreements visible before deciding. If every agent instantly agrees, the assignments may be too similar or the coordinator may be rewarding harmony instead of scrutiny. The point of a team is not more text. It is more independent opportunities to discover that the first idea is incomplete.
A reusable multidisciplinary agent prompt
Copy this structure and replace the brackets: “Act as the lead coordinator for a multidisciplinary agent team. Objective: deliver [specific outcome]. Context: use [files, links, archives, audience, constraints, and prior decisions]. Authorized scope: you may [research, inspect, edit, test, deploy, or publish]. Do not [make destructive changes, touch unrelated systems, or use unapproved accounts]. Success means [criterion one], [criterion two], and [criterion three].”
Then define the team: “Delegate independent work in parallel where supported. Assign an evidence researcher, human-centered reviewer, domain-risk critic, strategist, implementer, and independent verifier. Give each a bounded deliverable, evidence requirements, uncertainty rules, and a stop condition. Avoid concurrent edits to the same mutable files or systems.”
Add the honesty rules: “These roles are analytical perspectives, not credentialed human professionals. Do not invent interviews, user tests, credentials, citations, or consensus. Wait for the requested findings, merge duplicates, and resolve disagreements explicitly. If I require a fixed scope, produce exactly [N] agreed actions. Do not call work published, deployed, or verified without matching receipts.”
Finish with the operating sequence: “Brainstorm independently. Compare evidence and challenge assumptions. Reconcile one plan. Execute the authorized plan. Verify the final artifact independently. Report what changed, the evidence, unresolved risks, and the exact status: planning-only, local, deployed, live-verified, partial, or blocked.”
The verifier gets the last word
A multidisciplinary workflow is only as trustworthy as its status language. “Written” does not mean “published.” “Built” does not mean “deployed.” “Deployed” does not mean the correct route, feed, image, or audio is live. “Clicked Publish” does not prove a social post exists. Every external outcome needs a receipt that matches the claim: a public URL, response, identifier, read-back, screenshot, range request, or other evidence appropriate to the surface.
The human remains editor-in-chief. The agents expand the field of view, run independent work in parallel, and make verification systematic. The human defines the objective, grants authority, corrects the record, and decides what risks are acceptable. The best prompt is not the one that creates the largest fictional committee. It is the one that turns a difficult directive into inspectable work—and makes it hard for the system to confuse activity with completion.
Sources: S3
Why it matters
Multi-agent prompting can turn ChatGPT from a single-answer interface into a coordinated work system, especially when research, critique, implementation, and verification can proceed independently. The gain does not come from theatrical job titles. It comes from bounded context, explicit evidence, genuine disagreement, controlled authority, one reconciled plan, and receipts that distinguish planning-only, local, deployed, live-verified, partial, and blocked work.
Sources
- Subagents: use specialized agents in ChatGPT and Codex — OpenAI ·
- Multi-agent: parallel focused work in the Responses API — OpenAI ·
- Model guidance: define autonomy and approval boundaries — OpenAI ·
- 1,200 AI Agents Found Each Other. Then the Sandbox Failed. — OMIKINA ·
- John N. Farmer portfolio — John N. Farmer ·
- MARGIN: One Safer Next Step — John N. Farmer ·
Read OMIKINA's editorial standards · Review corrections · Follow the AI-narrated podcast · Follow the RSS briefing