从写代码,到创作下一幕

探索 字节跳动 - 火山方舟 的 AI 编程与视频创作活动。

Agent Plan & Coding Plan

一站体验多款热门模型,为 AI 编程与智能体开发提供更多选择。新用户可联系(微信: goo_lvyouyou)免费体验 9.9 agent plan。

Seedance 2.5

让创意,跃然成片。探索 30 秒视频、多模态参考与局部编辑,把脑海中的画面变成下一支作品。

中文

子智能体

在 ChatGPT 和 Codex 中使用子智能体,并配置自定义 Codex 智能体

ChatGPT 桌面应用

ChatGPT Work 和 Codex 可以生成多个专用智能体并行执行任务, 然后将它们的结果汇总到一个响应中,从而运行子智能体工作流。这对于 高度并行的复杂任务尤其有用,例如探索代码库或实施多步骤功能计划。

在本地 Codex 客户端中,你还可以针对不同任务定义具有不同模型 配置和指令的自定义智能体。

可用性

当前 Codex 版本默认启用子智能体工作流。子智能体活动 会显示在 ChatGPT 桌面应用、Codex CLI 和 IDE 扩展中。

由于每个子智能体都会独立执行模型和工具相关工作,子智能体工作流 比同类单智能体运行消耗更多 token。

在应用聊天中,请 Codex 将相互独立的工作部分委派给 子智能体。当前本地 Codex 版本会在你直接提出要求,或适用的 AGENTS.md 或技能指令要求委派时执行委派。应用会显示每个 子智能体线程,方便你检查其工作以及返回给主聊天的摘要。

子智能体工作流为何有用

即使上下文窗口很大,模型仍有其限制。如果你让主聊天(你在其中定义需求、约束和决策)充斥探索笔记、测试日志、堆栈跟踪和命令输出等嘈杂的中间输出,会话的可靠性可能会随时间推移而下降。

这种现象通常称为:

  • 上下文污染:有用信息被淹没在嘈杂的中间输出中。
  • 上下文腐化:随着聊天中不太相关的细节不断增多,性能逐渐下降。

有关背景信息,请参阅 Chroma 关于上下文腐化的文章。

子智能体工作流通过将嘈杂的工作移出主线程来提供帮助:

  • 让主智能体 专注于需求、决策和最终输出。
  • 并行运行专用子智能体,用于探索、测试或日志分析。
  • 让子智能体返回摘要,而不是原始中间输出。

当工作可以相互独立地并行执行时,它们还能节省时间;通过将 规模更大的任务拆分成边界明确的部分,也能使其更易处理。 例如,Codex 可以将对数百万 token 文档的分析拆分成 多个较小的问题,并向主线程返回提炼后的要点。

作为起点,可将并行智能体用于以读取为主的任务,例如 探索、测试、问题分类和总结。对于以写入为主的并行 工作流则应更加谨慎,因为多个智能体同时编辑代码可能会产生 冲突并增加协调开销。

核心术语

Codex 在子智能体工作流中使用以下几个相关术语:

  • 子智能体工作流:Codex 运行多个并行智能体并合并其结果的工作流。
  • 子智能体:Codex 启动并委派其处理特定任务的智能体。
  • 智能体线程:子智能体执行工作的线程。受支持的客户端允许你打开这些线程以检查进度或结果。

触发子智能体工作流

直接要求使用子智能体或并行智能体工作。当适用的项目或技能指令 要求委派时,Codex 也可以执行委派。

在实践中,手动触发是指使用直接指令,例如 “生成两个智能体”、“并行委派这项工作”或“每个要点使用一个智能体”。 子智能体工作流比同类单智能体运行消耗更多 token, 因为每个子智能体都会独立执行模型和工具相关工作。

一条好的子智能体提示词应说明如何划分工作、Codex 是否应 等待所有智能体完成后再继续,以及需要返回什么摘要或输出。

Review this branch with parallel subagents. Spawn one subagent for security risks, one for test gaps, and one for maintainability. Wait for all three, then summarize the findings by category with file references.

选择模型和推理级别

不同的智能体需要不同的模型和推理设置。

如果你没有配置子智能体模型或 model_reasoning_effort, 子智能体会继承父智能体的模型和推理强度。如果显式 创建请求或 [agents] 默认设置选择了模型,但没有 显式指定或配置推理强度,子智能体会使用该模型的默认 推理强度。要为每项任务平衡智能水平、速度和价格, 可以在提示词中请求使用特定模型或推理强度, 在 config.toml 中配置 [agents] 默认设置,或直接在自定义智能体文件中设置 model 和 model_reasoning_effort。 例如,使用 gpt-6-luna 进行快速扫描,或使用推理强度更高的 gpt-6.1-sol 配置来处理要求更高的推理任务。

模型选择

  • gpt-6.1-sol:对于需要处理高难度任务的智能体,请优先使用此模型。它适合需求不明确、包含多个步骤的工作,这类工作需要在较大的上下文中进行规划、使用工具、验证并持续跟进。
  • gpt-6-luna:适合需要快速处理明确、重复性或大批量工作的智能体,任务范围应较窄。

推理强度(model_reasoning_effort)

对于 GPT-6.1 Sol,请使用你的客户端和所选模型支持的 推理强度。显式设置模型时,GPT-6 Luna 可先使用 high,GPT-6 Astra 可先使用 low。 再根据任务调整为所选模型支持的级别。

  • ultra:在所选模型支持时,用于最深入的 推理。
  • max 和 xhigh:在所选模型支持这些级别时,用于要求特别高的 推理任务。
  • high:用于智能体需要梳理复杂逻辑、检查假设或处理边界情况的任务(例如评审智能体或专注于安全的智能体)。
  • medium:平衡速度与深度。
  • low:用于任务简单且速度最重要的情况。

更高的推理强度会增加响应时间和 token 用量,但可以提高复杂工作的质量。有关详细信息,请参阅模型、配置基础和配置参考。

编排和线程控制

ChatGPT 或 Codex 负责跨智能体编排,包括生成新的 子智能体、转发后续指令、等待结果以及关闭 智能体线程。

当多个智能体正在运行时,Codex 会等待所有请求的结果 就绪,然后返回合并后的响应。

当前本地 Codex 版本会在收到直接请求或适用的 项目或技能指令后生成智能体。

要查看实际效果,请在你的项目中尝试以下提示词:

I would like to review the following points on the current PR (this branch vs main). Spawn one agent per point, wait for all of them, and summarize the result for each point.
1. Security issue
2. Code quality
3. Bugs
4. Race
5. Test flakiness
6. Maintainability of the code

管理子智能体

  • 从主线程中显示的活动打开子智能体线程,以检查 其工作。
  • 直接要求 Codex 引导正在运行的子智能体、停止它,或关闭已完成的 子智能体线程。

审批和沙箱控制

子智能体会继承你当前的沙箱策略。

子智能体会继承输入框下方选择的权限模式。在要求 Codex 委派工作之前, 请先为父轮次选择权限模式。

你还可以覆盖单个自定义智能体的沙箱配置,例如明确将某个智能体标记为只读模式。

自定义智能体

Codex 随附以下内置智能体:

  • default:通用后备智能体。
  • worker:面向实施和修复、专注执行的智能体。
  • explorer:以读取为主的代码库探索智能体。

要定义自己的自定义智能体,请将独立 TOML 文件添加到 ~/.codex/agents/(个人智能体)或 .codex/agents/(项目范围的 智能体)下。

每个文件定义一个自定义智能体。Codex 会将这些文件作为生成会话的配置 层加载,因此自定义智能体可以覆盖与普通 Codex 会话配置相同的 设置。相比专用的智能体清单,这可能显得较为繁重;随着创作和共享机制日趋成熟, 其格式也可能发生变化。

每个独立的自定义智能体文件都必须定义:

  • name
  • description
  • developer_instructions

如果自定义智能体文件设置了 model 或 model_reasoning_effort,则 文件中的值优先。在应用该文件之前,Codex 会依次从显式创建参数、 对应的 [agents] 默认设置,以及 父智能体的值中确定各项设置。如果显式创建请求或 [agents] 默认设置 选择了模型,但两者均未提供推理强度,Codex 会使用 该模型的默认推理强度。仅设置 model 的自定义智能体文件 会保留此前确定的推理强度。如果所选模型不支持该强度,或你希望使用 其他强度,请同时在文件中设置 model_reasoning_effort。其他会话设置,例如 sandbox_mode、mcp_servers 和 skills.config,在自定义智能体文件未指定时,会从父智能体 继承。

全局设置

全局子智能体设置仍位于配置中的 [agents] 下。

字段 类型 必填 用途
agents.enabled boolean 否 启用或禁用多智能体工具。
agents.max_concurrent_threads_per_session number 否 限制并发打开的已生成智能体线程数量,不包括主智能体。
agents.default_subagent_model string 否 设置已生成智能体的默认模型。
agents.default_subagent_reasoning_effort string 否 设置已生成智能体的默认推理强度。
agents.interrupt_message boolean 否 智能体轮次中断时记录一条模型可见消息。

注意:

  • agents.enabled 默认为 true。将其设为 false 可禁用多智能体工具。
  • 如果未设置 agents.max_concurrent_threads_per_session,Codex 会选择默认值。现有配置可以继续使用 agents.max_threads 作为旧版别名。
  • 显式生成值会覆盖 agents.default_subagent_model 和 agents.default_subagent_reasoning_effort。
  • agents.interrupt_message 默认为 true。将其设为 false 可从智能体上下文中省略模型可见的中断消息。
  • 如果自定义智能体名称与 explorer 等内置智能体相同,则以自定义智能体为准。

自定义智能体文件格式

字段 类型 必填 用途
name string 是 Codex 在生成或引用该智能体时使用的智能体名称。
description string 是 面向用户的指南,说明 Codex 应在何时使用此智能体。
developer_instructions string 是 定义智能体行为的核心指令。

你还可以在自定义智能体文件中包含其他受支持的 config.toml 键,例如 model、model_reasoning_effort、sandbox_mode、mcp_servers 和 skills.config。

Codex 通过自定义智能体的 name 字段识别它。让文件名与 智能体名称保持一致是最简单的约定,但 name 字段才是最终 依据。

自定义智能体示例

最优秀的自定义智能体应当范围明确且有鲜明倾向。为每个智能体分配清晰的职责, 提供与该职责匹配的工具范围,并通过指令避免其 偏离到相邻工作。

如果你的已登录账户或工作区拥有访问权限,这些示例会使用 GPT-6.1 Sol。 如果该模型不可用,请选择你可以 使用的模型。

示例 1:PR 评审

此模式将评审工作拆分给三个各有侧重的自定义智能体:

  • pr_explorer 梳理代码库并收集证据。
  • reviewer 查找正确性、安全性和测试风险。
  • docs_researcher 通过专用 MCP server检查框架或 API 文档。

项目配置(.codex/config.toml):

[agents]
max_concurrent_threads_per_session = 8

.codex/agents/pr-explorer.toml:

name = "pr_explorer"
description = "Read-only codebase explorer for gathering evidence before changes are proposed."
model = "gpt-6-luna"
model_reasoning_effort = "high"
sandbox_mode = "read-only"
developer_instructions = """
Stay in exploration mode.
Trace the real execution path, cite files and symbols, and avoid proposing fixes unless the parent agent asks for them.
Prefer fast search and targeted file reads over broad scans.
"""

.codex/agents/reviewer.toml:

name = "reviewer"
description = "PR reviewer focused on correctness, security, and missing tests."
model = "gpt-6.1-sol"
model_reasoning_effort = "medium"
sandbox_mode = "read-only"
developer_instructions = """
Review code like an owner.
Prioritize correctness, security, behavior regressions, and missing test coverage.
Lead with concrete findings, include reproduction steps when possible, and avoid style-only comments unless they hide a real bug.
"""

.codex/agents/docs-researcher.toml:

name = "docs_researcher"
description = "Documentation specialist that uses the docs MCP server to verify APIs and framework behavior."
model = "gpt-6-luna"
model_reasoning_effort = "high"
sandbox_mode = "read-only"
developer_instructions = """
Use the docs MCP server to confirm APIs, options, and version-specific behavior.
Return concise answers with links or exact references when available.
Do not make code changes.
"""

[mcp_servers.openaiDeveloperDocs]
url = "https://developers.openai.com/mcp"

此设置非常适合以下提示词:

Review this branch against main. Have pr_explorer map the affected code paths, reviewer find real risks, and docs_researcher verify the framework APIs that the patch relies on.

示例 2:前端集成调试

此模式适用于 UI 回归、间歇性浏览器流程,或横跨应用代码和运行中产品的集成错误。

项目配置(.codex/config.toml):

[agents]
max_concurrent_threads_per_session = 6

.codex/agents/code-mapper.toml:

name = "code_mapper"
description = "Read-only codebase explorer for locating the relevant frontend and backend code paths."
model = "gpt-6-luna"
model_reasoning_effort = "high"
sandbox_mode = "read-only"
developer_instructions = """
Map the code that owns the failing UI flow.
Identify entry points, state transitions, and likely files before the worker starts editing.
"""

.codex/agents/browser-debugger.toml:

name = "browser_debugger"
description = "UI debugger that uses browser tooling to reproduce issues and capture evidence."
model = "gpt-6.1-sol"
model_reasoning_effort = "medium"
sandbox_mode = "workspace-write"
developer_instructions = """
Reproduce the issue in the browser, capture exact steps, and report what the UI actually does.
Use browser tooling for screenshots, console output, and network evidence.
Do not edit application code.
"""
[mcp_servers.chrome_devtools]
url = "http://localhost:3000/mcp"
startup_timeout_sec = 20

.codex/agents/ui-fixer.toml:

name = "ui_fixer"
description = "Implementation-focused agent for small, targeted fixes after the issue is understood."
model = "gpt-6-luna"
model_reasoning_effort = "high"
developer_instructions = """
Own the fix once the issue is reproduced.
Make the smallest defensible change, keep unrelated files untouched, and validate only the behavior you changed.
"""

[[skills.config]]
path = "/Users/me/.agents/skills/docs-editor/SKILL.md"
enabled = false

此设置非常适合以下提示词:

Investigate why the settings modal fails to save. Have browser_debugger reproduce it, code_mapper trace the responsible code path, and ui_fixer implement the smallest fix once the failure mode is clear.