Agents SDK で Codex を使用する
Codex を MCP server として呼び出し、マルチエージェント開発ワークフローを構築します
Codex を MCP server として実行する
Codex を MCP server として実行し、ほかの MCP クライアント(たとえば、OpenAI Agents SDK MCP インテグレーションで構築したエージェント)から接続できます。
Codex を MCP server として起動するには、次のコマンドを使用します:
codex mcp-serverModel Context Protocol Inspector を使用して Codex MCP server を起動できます:
npx @modelcontextprotocol/inspector codex mcp-servertools/list request を送信すると、次の 2 つのツールが表示されます:
codex:次のプロンプトと設定の上書きを使用して Codex セッションを実行します:
| プロパティ | 型 | 説明 |
|---|---|---|
prompt(必須) |
string |
Codex の会話を開始するための最初のユーザープロンプトです。 |
approval-policy |
string |
モデルが生成したシェルコマンドの承認ポリシーです:untrusted、on-request、never。 |
base-instructions |
string |
デフォルトの指示の代わりに使用する一連の指示です。 |
compact-prompt |
string |
会話を圧縮するときに使用するプロンプトです。 |
config |
object |
$CODEX_HOME/config.toml の内容を上書きする個別の設定です。 |
cwd |
string |
セッションの作業ディレクトリです。相対パスの場合は、サーバープロセスの現在のディレクトリを基準に解決されます。 |
developer-instructions |
string |
developer ロールのメッセージとして注入される開発者指示です。 |
model |
string |
モデル名の任意の上書きです(例:gpt-5.6-terra)。 |
sandbox |
string |
サンドボックスモード:read-only、workspace-write、danger-full-access。 |
codex-reply:thread ID とプロンプトを指定して Codex セッションを続行します。codex-reply ツールは次のプロパティを受け取ります:
| プロパティ | 型 | 説明 |
|---|---|---|
prompt(必須) |
string | Codex の会話を続けるための次のユーザープロンプトです。 |
threadId(必須) |
string | 続行する thread の ID です。 |
conversationId(非推奨) |
string | threadId の非推奨エイリアスです(互換性のために維持されています)。 |
tools/call response の structuredContent.threadId にある threadId を使用してください。承認プロンプト(exec/patch)の params payload にも threadId が含まれます。
response payload の例:
{
"structuredContent": {
"threadId": "019bbb20-bff6-7130-83aa-bf45ab33250e",
"content": "`ls -lah` (or `ls -alh`) — long listing, includes dotfiles, human-readable sizes."
},
"content": [
{
"type": "text",
"text": "`ls -lah` (or `ls -alh`) — long listing, includes dotfiles, human-readable sizes."
}
]
}最新の MCP クライアントは通常、存在する場合はツール呼び出しの結果として "structuredContent" のみを報告します。ただし、Codex MCP server は古い MCP クライアントのために "content" も返します。
マルチエージェントワークフローを作成する
Codex CLI は、単発のタスクを実行する以上のことができます。CLI を Model Context Protocol(MCP)server として公開し、OpenAI Agents SDK でオーケストレーションすることで、単一のエージェントから完全なソフトウェアデリバリーパイプラインまで拡張できる、決定論的でレビュー可能なワークフローを作成できます。
このガイドでは、OpenAI Cookbook で紹介されているものと同じワークフローを説明します。次のことを行います:
- Codex CLI を長時間稼働する MCP server として起動します。
- プレイ可能なブラウザーゲームを生成する、目的を絞った単一エージェントのワークフローを構築します。
- ハンドオフ、ガードレール、後から確認できる完全なトレースを備えたマルチエージェントチームをオーケストレーションします。
開始する前に、次のものを用意してください:
codexコマンドを使用できるように、Codex CLI をローカルにインストールします。pipを備えた Python 3.10 以降。- 上記の MCP Inspector の例を実行する場合は Node.js 18 以降。
- ローカルに保存した OpenAI API key。OpenAI dashboard でキーを作成または管理できます。
ガイド用の作業ディレクトリを作成し、API key を .env ファイルに追加します:
mkdir codex-workflows
cd codex-workflows
printf "OPENAI_API_KEY=sk-..." > .env依存関係をインストールする
Agents SDK は、Codex、ハンドオフ、トレースにまたがるオーケストレーションを処理します。最新の SDK パッケージをインストールします:
python -m venv .venv
source .venv/bin/activate
pip install --upgrade openai openai-agents python-dotenv
Codex CLI を MCP server として初期化する
まず、Codex CLI を Agents SDK から呼び出せる MCP server にします。サーバーは 2 つのツール(会話を開始する codex() と、会話を続ける codex-reply())を公開し、複数のエージェント turn にわたって Codex を稼働状態に保ちます。
codex_mcp.py というファイルを作成し、次の内容を追加します:
import asyncio
from agents import Agent, Runner
from agents.mcp import MCPServerStdio
async def main() -> None:
async with MCPServerStdio(
name="Codex CLI",
params={
"command": "codex",
"args": ["mcp-server"],
},
client_session_timeout_seconds=360000,
) as codex_mcp_server:
print("Codex MCP server started.")
# More logic coming in the next sections.
return
if __name__ == "__main__":
asyncio.run(main())スクリプトを一度実行し、Codex が正常に起動することを確認します:
python codex_mcp.pyスクリプトは Codex MCP server started. と出力した後に終了します。次のセクションでは、より充実したワークフロー内で同じ MCP server を再利用します。
単一エージェントのワークフローを構築する
まず、Codex MCP を使用して小さなブラウザーゲームを完成させる、スコープを絞った例から始めましょう。このワークフローでは 2 つのエージェントを使用します:
- ゲームデザイナー:ゲームの概要を作成します。
- ゲーム開発者:Codex MCP を呼び出してゲームを実装します。
次のコードで codex_mcp.py を更新します。上記の MCP server の設定を維持しつつ、両方のエージェントを追加します。
import asyncio
import os
from dotenv import load_dotenv
from agents import Agent, Runner, set_default_openai_api
from agents.mcp import MCPServerStdio
load_dotenv(override=True)
set_default_openai_api(os.getenv("OPENAI_API_KEY"))
async def main() -> None:
async with MCPServerStdio(
name="Codex CLI",
params={
"command": "codex",
"args": ["mcp-server"],
},
client_session_timeout_seconds=360000,
) as codex_mcp_server:
developer_agent = Agent(
name="Game Developer",
instructions=(
"You are an expert in building simple games using basic html + css + javascript with no dependencies. "
"Save your work in a file called index.html in the current directory. "
"Always call codex with \"approval-policy\": \"never\" and \"sandbox\": \"workspace-write\"."
),
mcp_servers=[codex_mcp_server],
)
designer_agent = Agent(
name="Game Designer",
instructions=(
"You are an indie game connoisseur. Come up with an idea for a single page html + css + javascript game that a developer could build in about 50 lines of code. "
"Format your request as a 3 sentence design brief for a game developer and call the Game Developer coder with your idea."
),
model="gpt-5",
handoffs=[developer_agent],
)
await Runner.run(designer_agent, "Implement a fun new game!")
if __name__ == "__main__":
asyncio.run(main())スクリプトを実行します:
python codex_mcp.pyCodex はデザイナーの概要を読み、index.html ファイルを作成して、完全なゲームをディスクに書き込みます。生成されたファイルをブラウザーで開いてプレイしてください。実行するたびに、独自のプレイスタイルの工夫と仕上げを備えた異なるデザインが生成されます。
マルチエージェントワークフローへ拡張する
次に、単一エージェントの構成を、オーケストレーションされ、トレース可能なワークフローへ変えます。システムには次の役割が追加されます:
- プロジェクトマネージャー:共有要件を作成し、ハンドオフを調整して、ガードレールを適用します。
- デザイナー、フロントエンド開発者、サーバー開発者、テスター:それぞれにスコープを限定した指示と出力フォルダーを割り当てます。
multi_agent_workflow.py という新しいファイルを作成します:
import asyncio
import os
from dotenv import load_dotenv
from agents import (
Agent,
ModelSettings,
Runner,
WebSearchTool,
set_default_openai_api,
)
from agents.extensions.handoff_prompt import RECOMMENDED_PROMPT_PREFIX
from agents.mcp import MCPServerStdio
from openai.types.shared import Reasoning
load_dotenv(override=True)
set_default_openai_api(os.getenv("OPENAI_API_KEY"))
async def main() -> None:
async with MCPServerStdio(
name="Codex CLI",
params={"command": "codex", "args": ["mcp-server"]},
client_session_timeout_seconds=360000,
) as codex_mcp_server:
designer_agent = Agent(
name="Designer",
instructions=(
f"""{RECOMMENDED_PROMPT_PREFIX}"""
"You are the Designer.\n"
"Your only source of truth is AGENT_TASKS.md and REQUIREMENTS.md from the Project Manager.\n"
"Do not assume anything that is not written there.\n\n"
"You may use the internet for additional guidance or research."
"Deliverables (write to /design):\n"
"- design_spec.md – a single page describing the UI/UX layout, main screens, and key visual notes as requested in AGENT_TASKS.md.\n"
"- wireframe.md – a simple text or ASCII wireframe if specified.\n\n"
"Keep the output short and implementation-friendly.\n"
"When complete, handoff to the Project Manager with transfer_to_project_manager."
"When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}."
),
model="gpt-5",
tools=[WebSearchTool()],
mcp_servers=[codex_mcp_server],
)
frontend_developer_agent = Agent(
name="Frontend Developer",
instructions=(
f"""{RECOMMENDED_PROMPT_PREFIX}"""
"You are the Frontend Developer.\n"
"Read AGENT_TASKS.md and design_spec.md. Implement exactly what is described there.\n\n"
"Deliverables (write to /frontend):\n"
"- index.html – main page structure\n"
"- styles.css or inline styles if specified\n"
"- main.js or game.js if specified\n\n"
"Follow the Designer’s DOM structure and any integration points given by the Project Manager.\n"
"Do not add features or branding beyond the provided documents.\n\n"
"When complete, handoff to the Project Manager with transfer_to_project_manager_agent."
"When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}."
),
model="gpt-5",
mcp_servers=[codex_mcp_server],
)
backend_developer_agent = Agent(
name="Backend Developer",
instructions=(
f"""{RECOMMENDED_PROMPT_PREFIX}"""
"You are the Backend Developer.\n"
"Read AGENT_TASKS.md and REQUIREMENTS.md. Implement the backend endpoints described there.\n\n"
"Deliverables (write to /backend):\n"
"- package.json – include a start script if requested\n"
"- server.js – implement the API endpoints and logic exactly as specified\n\n"
"Keep the code as simple and readable as possible. No external database.\n\n"
"When complete, handoff to the Project Manager with transfer_to_project_manager_agent."
"When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}."
),
model="gpt-5",
mcp_servers=[codex_mcp_server],
)
tester_agent = Agent(
name="Tester",
instructions=(
f"""{RECOMMENDED_PROMPT_PREFIX}"""
"You are the Tester.\n"
"Read AGENT_TASKS.md and TEST.md. Verify that the outputs of the other roles meet the acceptance criteria.\n\n"
"Deliverables (write to /tests):\n"
"- TEST_PLAN.md – bullet list of manual checks or automated steps as requested\n"
"- test.sh or a simple automated script if specified\n\n"
"Keep it minimal and easy to run.\n\n"
"When complete, handoff to the Project Manager with transfer_to_project_manager."
"When creating files, call Codex MCP with {\"approval-policy\":\"never\",\"sandbox\":\"workspace-write\"}."
),
model="gpt-5",
mcp_servers=[codex_mcp_server],
)
project_manager_agent = Agent(
name="Project Manager",
instructions=(
f"""{RECOMMENDED_PROMPT_PREFIX}"""
"""
You are the Project Manager.
Objective:
Convert the input task list into three project-root files the team will execute against.
Deliverables (write in project root):
- REQUIREMENTS.md: concise summary of product goals, target users, key features, and constraints.
- TEST.md: tasks with [Owner] tags (Designer, Frontend, Backend, Tester) and clear acceptance criteria.
- AGENT_TASKS.md: one section per role containing:
- Project name
- Required deliverables (exact file names and purpose)
- Key technical notes and constraints
Process:
- Resolve ambiguities with minimal, reasonable assumptions. Be specific so each role can act without guessing.
- Create files using Codex MCP with {"approval-policy":"never","sandbox":"workspace-write"}.
- Do not create folders. Only create REQUIREMENTS.md, TEST.md, AGENT_TASKS.md.
Handoffs (gated by required files):
1) After the three files above are created, hand off to the Designer with transfer_to_designer_agent and include REQUIREMENTS.md and AGENT_TASKS.md.
2) Wait for the Designer to produce /design/design_spec.md. Verify that file exists before proceeding.
3) When design_spec.md exists, hand off in parallel to both:
- Frontend Developer with transfer_to_frontend_developer_agent (provide design_spec.md, REQUIREMENTS.md, AGENT_TASKS.md).
- Backend Developer with transfer_to_backend_developer_agent (provide REQUIREMENTS.md, AGENT_TASKS.md).
4) Wait for Frontend to produce /frontend/index.html and Backend to produce /backend/server.js. Verify both files exist.
5) When both exist, hand off to the Tester with transfer_to_tester_agent and provide all prior artifacts and outputs.
6) Do not advance to the next handoff until the required files for that step are present. If something is missing, request the owning agent to supply it and re-check.
PM Responsibilities:
- Coordinate all roles, track file completion, and enforce the above gating checks.
- Do NOT respond with status updates. Just handoff to the next agent until the project is complete.
"""
),
model="gpt-5",
model_settings=ModelSettings(
reasoning=Reasoning(effort="medium"),
),
handoffs=[designer_agent, frontend_developer_agent, backend_developer_agent, tester_agent],
mcp_servers=[codex_mcp_server],
)
designer_agent.handoffs = [project_manager_agent]
frontend_developer_agent.handoffs = [project_manager_agent]
backend_developer_agent.handoffs = [project_manager_agent]
tester_agent.handoffs = [project_manager_agent]
task_list = """
Goal: Build a tiny browser game to showcase a multi-agent workflow.
High-level requirements:
- Single-screen game called "Bug Busters".
- Player clicks a moving bug to earn points.
- Game ends after 20 seconds and shows final score.
- Optional: submit score to a simple backend and display a top-10 leaderboard.
Roles:
- Designer: create a one-page UI/UX spec and basic wireframe.
- Frontend Developer: implement the page and game logic.
- Backend Developer: implement a minimal API (GET /health, GET/POST /scores).
- Tester: write a quick test plan and a simple script to verify core routes.
Constraints:
- No external database—memory storage is fine.
- Keep everything readable for beginners; no frameworks required.
- All outputs should be small files saved in clearly named folders.
"""
result = await Runner.run(project_manager_agent, task_list, max_turns=30)
print(result.final_output)
if __name__ == "__main__":
asyncio.run(main())スクリプトを実行し、生成されるファイルを確認します:
python multi_agent_workflow.py
ls -Rプロジェクトマネージャーエージェントは REQUIREMENTS.md、TEST.md、AGENT_TASKS.md を作成した後、デザイナー、フロントエンド、サーバー、テスターの各エージェント間のハンドオフを調整します。各エージェントは、プロジェクトマネージャーへ制御を戻す前に、担当範囲の成果物をそれぞれのフォルダーへ書き込みます。
ワークフローをトレースする
Codex は、すべてのプロンプト、ツール呼び出し、ハンドオフを記録するトレースを自動的に作成します。マルチエージェントの実行が完了したら、Traces dashboard を開いて実行タイムラインを確認します。
概要トレースでは、プロジェクトマネージャーが次へ進む前にハンドオフを検証する流れが強調されます。個々のステップを開くと、プロンプト、Codex MCP の呼び出し、書き込まれたファイル、実行時間を確認できます。これらの詳細により、各ハンドオフを容易に監査し、ワークフローが turn ごとにどのように変化したかを把握できます。これらのトレースを使用すると、追加の計測を導入しなくても、ワークフローの問題を簡単にデバッグし、エージェントの動作を監査し、時間の経過に伴うパフォーマンスを測定できます。