← 返回工具箱
Codex Field Manual

如何高效使用 Codex

一份面向实际开发、科研、机器学习与多 Agent 协作的实战使用手册。

0. 核心结论

高效使用 Codex,最重要的转变是:把 Codex 当作一个能够读取代码库、执行命令、修改文件、运行测试、检查结果并继续迭代的软件工程 Agent。

你的主要工作逐渐从“告诉它每一行代码怎么写”,转变为“定义目标、提供必要上下文、设置边界、规定验收标准,然后检查结果”。

一个高质量 Codex 任务通常应该包含四件事:

目标
上下文
约束
验收标准

例如,与其说:

帮我改一下 train.py。

更好的说法是:

修复 train.py 中 resume training 后 learning-rate scheduler 状态没有正确恢复的问题。

背景:
- checkpoint 已保存 optimizer 和 scheduler state
- 当前 resume 后 loss 曲线会突然异常
- 不希望改变正常从头训练的行为

要求:
1. 先定位问题,不要直接猜。
2. 尽量做最小修改。
3. 增加一个可以复现问题的测试。
4. 运行相关测试。
5. 最后告诉我:
   - 根因
   - 修改了哪些文件
   - 如何验证
   - 是否还存在风险

1. 先理解 Codex 的工作方式

Codex 最有效的工作方式大致是:

理解目标
    ↓
读取相关代码
    ↓
形成假设 / 计划
    ↓
修改代码
    ↓
运行测试 / build / benchmark
    ↓
查看结果
    ↓
修复问题
    ↓
再次验证
    ↓
完成

长任务尤其依赖这个循环。真正决定产出的,往往不是 prompt 写得多复杂,而是 Codex 是否能进入一个稳定的“行动—反馈—修正”闭环。

2. 选择正确的 Codex 使用入口

2.1 Desktop / GUI

适合日常开发、查看 diff、多任务并行、Worktree 管理、长任务以及不想一直盯着 Terminal 的工作。

2.2 CLI

codex

适合 SSH、远程服务器、GPU 训练机、Linux 环境、快速检查代码、直接跑测试和日志分析。

2.3 IDE Extension

适合一边人工写代码一边使用 Codex,尤其适合精确查看函数、小范围 refactor、debugging 和代码理解。

2.4 Codex Cloud

适合本机不想被长期占用、任务可能运行较久、希望电脑休眠后任务仍可继续的场景。

Local Codex
→ 在你的机器上执行

Codex Cloud
→ 在托管环境中执行

2.5 Remote

如果开发机位于宿舍、办公室或服务器旁边,而你手上只有另一台设备,可以把 Remote 理解为远程控制平面:代码仍运行在开发机上,你负责启动、审批、steering、review diff 和查看结果。

3. 仓库初始化

第一次让 Codex 接触一个项目时,不要马上扔给它复杂任务。先确认 Git 状态与项目运行方式。

git status

推荐第一条任务:

先不要修改代码。

请阅读这个仓库,告诉我:

1. 项目的主要目录结构。
2. 主程序入口。
3. 测试方式。
4. build / lint / typecheck 命令。
5. 哪些文件最重要。
6. 如果以后让你开发功能,你认为还缺哪些项目说明。

不要逐文件复述。

4. AGENTS.md

AGENTS.md 是 Codex 项目里最重要的长期上下文文件之一,适合保存项目结构、常用命令、约束和编码规范。

项目结构

## Architecture

- `src/model/`: model definitions
- `src/train/`: training loop
- `src/eval/`: evaluation
- `scripts/`: experiment scripts

常用命令

## Commands

Install:
pip install -e .

Tests:
pytest tests/

Lint:
ruff check .

Type check:
pyright

约束

## Constraints

- Do not change dataset preprocessing unless explicitly requested.
- Preserve checkpoint backward compatibility.
- Do not commit generated checkpoints.

5. 不要把 AGENTS.md 写成百科全书

长期积累的大量旧规则很容易产生反作用。

不推荐:

每次修改代码之前:
阅读 architecture.md
阅读 database.md
阅读 deployment.md
阅读 README.md
读取全部 tests

更合理:

Use architecture.md when changing service boundaries.
Use database.md when changing schemas.
Use deployment.md for deployment-related work.
告诉 Codex 信息在哪里,而不是强迫它每次把所有信息全部读一遍。

6. 如何写一个高质量 Codex Prompt

任务
背景
约束
验收
输出要求
任务:
<你希望最终发生什么>

背景:
<为什么要做>
<相关模块>
<已有情况>

约束:
- 不要修改……
- 保持……
- 尽量……
- 可以修改……

验收:
- 测试 X 通过
- benchmark 达到 X
- 行为 X 保持不变
- case X 可以正确工作

完成后:
1. 总结根因/设计。
2. 列出修改文件。
3. 告诉我运行了哪些验证。
4. 报告仍然存在的风险。

7. 少告诉 Codex“怎么做”,多告诉它“什么算完成”

低效:

打开 a.py。
找到 foo。
增加一个 if。
然后修改 b.py。

更好:

解决用户 logout 后 websocket 连接仍然存在的问题。

要求:
- logout 后 1 秒以内关闭连接
- 正常断线重连逻辑不能受影响
- 添加 regression test

先调查目前 websocket lifecycle,再选择最合理的实现。
你通常最清楚的是“要什么结果”;Codex 更擅长在代码库中寻找“怎么达到结果”。

8. 什么时候使用 Plan

复杂任务可以先让 Codex 调查而不修改:

先调查,不要修改代码。

请:
1. 找到相关实现。
2. 描述当前数据流。
3. 找出最可能的问题。
4. 提出修改方案。
5. 说明涉及哪些文件。
6. 说明测试方案。

等方案明确后再实施。

Plan 特别适合架构变化、数据库 migration、大范围 refactor、陌生代码库、高风险功能和跨多个 subsystem 的任务。

9. Goals:长任务的关键工具

Goal 更适合“路径未知,但终点明确”的任务,例如性能优化、flaky test、benchmark 优化、debugging、大型 refactor 和多轮实验。

/goal Reduce p95 inference latency below 100 ms while keeping all correctness tests green.

机器学习任务尤其适合:

/goal Reduce 4-NFE PPL below 350 on the fixed validation configuration,
while preserving entropy and distinct-2 within the baseline range.
Use the existing evaluation scripts.
Do not change the evaluation dataset.
Record each experiment and its result.

一个优秀 Goal 通常包含:

Outcome
Verification
Constraints
Boundaries
Iteration policy
Stop condition

10. Worktree:并行开发最重要的工具之一

main
├── worktree A
├── worktree B
└── worktree C

适合 Worktree 的场景包括:多个 Codex 并行、Codex + Claude Code、实验性 refactor、长任务以及不确定是否最终采用的方案。

一个独立目标对应一个 worktree。

11. 如何并行使用多个 Agent

不要让多个 Agent 同时改同一个核心文件。更好的并行方式是按问题或模块划分。

Agent A → 调查训练速度瓶颈
Agent B → 检查 dataloader
Agent C → 分析 profiler 输出
Agent D → 检查 GPU utilization

或者:

Agent A → frontend
Agent B → backend
Agent C → tests

多 Agent 最适合:独立调查、独立模块、独立实验、独立 review。

12. Codex 最擅长的任务

12.1 Bug investigation

Investigate this bug.

Observed:
<现象>

Expected:
<预期>

Reproduction:
<复现方法>

Do not patch immediately.

First:
1. reproduce it;
2. trace the relevant execution path;
3. identify the root cause.

Then implement the smallest robust fix and add a regression test.

12.2 新功能实现

Implement <feature>.

Before editing:
- inspect the existing architecture;
- reuse existing abstractions when appropriate;
- identify the smallest integration point.

Requirements:
...

Acceptance criteria:
...

Run relevant tests after implementation.

12.3 Refactor

Refactor the evaluation pipeline to separate:
- generation
- metric computation
- result persistence

Do not change observable evaluation behavior.

After refactoring:
- existing tests must pass;
- add tests if current coverage is insufficient;
- compare outputs on a small fixed fixture.

12.4 读代码

Trace what happens from:

python train.py ...

until the first optimizer.step().

Focus only on:
- config loading
- dataset creation
- model creation
- loss computation
- backward
- optimizer update

Give me file/function references.

13. 让 Codex 使用真实证据

重要任务中应明确要求:

Do not consider the task complete until the relevant verification has been run.

验收证据优先级:

真实运行结果
>
自动化测试
>
静态检查
>
代码推理
>
“看起来应该可以”

14. 让测试成为 Definition of Done

完成条件:

pytest tests/test_checkpoint_resume.py

必须通过。

另外运行一个 100-step smoke training,
确认 resume 前后的 LR 连续。

一旦完成条件明确,Agent 的搜索空间会明显缩小。

15. 长任务一定要外部化状态

长任务不要把所有状态只留在聊天记录里。建议把关键状态写入 repo:

docs/
  architecture.md

experiments/
  README.md
  results.csv

PLAN.md
STATUS.md

特别长的任务可以维护:

# Goal

# Current understanding

# Completed

# Current hypothesis

# Failed approaches

# Next step

# Verification

16. Side Chat

当主线程正在执行长任务,而你只是想问一个解释性问题时,Side Chat 可以避免主任务被打断。

/side

适合:问解释、问架构、问错误、讨论想法、分析某个 diff。

17. 什么时候开新 Chat

同一个目标:继续原 chat。

新的独立目标:开新 chat。

原因是旧线程会积累大量与新任务无关的上下文。

18. 上下文越来越长怎么办

/compact

也可以主动要求 Codex 把当前状态写入 STATUS.md,然后新线程从文件继续。

19. 使用 /review

Review the current diff as if you were reviewing a pull request.

Focus on:
- correctness
- regressions
- edge cases
- concurrency
- state management
- performance
- test coverage

Do not modify anything yet.

Rank findings by severity.

一个非常有效的结构:

Agent A 写代码
        ↓
Agent B review
        ↓
Agent A 修复

20. Git 工作流

git status
git diff
git log

推荐流程:

main
 ↓
create branch/worktree
 ↓
Codex implementation
 ↓
tests
 ↓
review diff
 ↓
commit
 ↓
merge / PR

大型 refactor、实验代码、依赖升级、migration 等任务,不建议直接让 Agent 在 main 上无限修改。

21. 权限和安全

本地 Codex 可以执行 shell command,因此权限设计很重要。

普通任务通常只需要允许:

读取项目
修改项目
运行测试

以下操作应谨慎:

删除大量文件
修改系统配置
访问 secrets
push production
数据库 destructive migration
sudo

22. Secret 管理

不要把 API key、密码、access token、private key 直接写入 AGENTS.md、prompt、源码或 commit。

优先使用:

environment variables
secret manager
.env(并加入 .gitignore)

Agent 通常只需要知道 secret 从哪里读取,而不需要知道 secret 本身。

23. Skills

如果经常重复同一种复杂流程,可以把它做成 Skill,例如:

run-training-evaluation
prepare-release
review-migration
benchmark-model
generate-experiment-report

可以这样理解:

AGENTS.md
= 项目宪法

Skill
= 标准作业流程

Prompt
= 当前工单

24. MCP / Plugins

当 Codex 需要频繁访问外部系统时,可以接入 GitHub、documentation、issue tracker、database、internal services 等。

这通常比把大量网页内容复制到 prompt 中更适合长期项目。

25. 模型如何选择

简单任务:改变量名、增加 log、修 typo、简单测试。

中型任务:普通功能、debug、模块重构。

困难任务:架构、复杂 race condition、大型 migration、科研代码、复杂数学逻辑、长任务。

原则是:不要把大量时间花在“哪个模型理论上聪明 3%”,而应把更强模型留给真正需要复杂推理和长程规划的任务。

26. 减少 Token 浪费

  1. 不要每次都让 Codex 读取整个 repo。
  2. 长期背景放 AGENTS.md / docs,不要反复塞进 prompt。
  3. 不要不断发送“继续”;如果终点明确,考虑 Goal。
  4. 简单任务不要要求长篇解释,只保留 changed files / verification / remaining risks。

27. 最容易失败的 Codex 使用方式

27.1 任务过于模糊

优化一下代码。

更好:

Reduce peak GPU memory during evaluation by at least 20%
without changing evaluation outputs.

27.2 同时塞十个无关任务

UI、API、PyTorch 升级、CUDA 优化、数据库重构、README 最好拆成独立任务。

27.3 没有 Definition of Done

没有验收标准时,Agent 很容易停在“看起来合理”。

27.4 用户自己规定错误实现

如果你不确定实现方式,不要强制规定某个技术方案;可以表达假设,让 Codex 先验证。

27.5 没有先复现 bug

Debug 的黄金规则:先复现,再修复。

28. 科研 / Machine Learning 项目的推荐结构

project/
├── src/
├── configs/
├── scripts/
├── tests/
├── experiments/
│   ├── README.md
│   ├── results.csv
│   └── notes/
├── docs/
│   └── architecture.md
├── AGENTS.md
└── README.md

AGENTS.md 可加入实验约束:

## Experiment policy

- Do not overwrite existing experiment results.
- Every new experiment must record its config.
- Use fixed evaluation seeds unless explicitly testing variance.
- Never change evaluation protocol when comparing experiments.
- Record failed experiments as well as successful ones.

29. 让 Codex 做实验

Goal:

Improve 4-step generation quality.

Success criteria:
- PPL < 350
- entropy remains within baseline ±5%
- distinct-2 does not decrease more than 5%

Constraints:
- same validation set
- same tokenizer
- same evaluation script
- no cherry-picking samples

Workflow:
1. inspect current implementation;
2. propose one hypothesis at a time;
3. make the smallest testable change;
4. run the defined experiment;
5. record result;
6. keep or revert based on evidence.

Stop if three consecutive hypotheses fail to improve the primary metric and report findings.

30. Codex + ChatGPT 的推荐分工

ChatGPT:理论分析、架构讨论、阅读论文、设计实验、比较方案。

Codex:读取 repo、实现、debug、测试、benchmark、Git、真实运行。

ChatGPT
    ↓
确定方向
    ↓
SPEC.md
    ↓
Codex
    ↓
实现 + 测试
    ↓
实验结果
    ↓
ChatGPT
    ↓
分析结果

31. Codex + Claude Code

不推荐让 Codex 和 Claude Code 同时编辑同一个 worktree。

更好的方式:

worktree-codex
worktree-claude

或者:

Codex → implementation
Claude → review

再或者:

Claude → architecture investigation
Codex → implementation

32. 一个完整的标准工作流

  1. 创建 Worktree:一个任务一个 worktree。
  2. 让 Codex 调查:Investigate first. Do not edit yet.
  3. 确认目标、约束、验收。
  4. 复杂任务先 Plan。
  5. 执行 implementation。
  6. 运行 tests / lint / typecheck / benchmark。
  7. 使用 review。
  8. 修复 review 问题。
  9. 查看 git diff。
  10. Commit。

33. 常用 Prompt 模板库

模板 1:理解项目

Do not modify anything yet.

Inspect this repository and explain:

1. high-level architecture;
2. major execution entry points;
3. important modules;
4. test/build/lint workflow;
5. where configuration lives;
6. areas that appear fragile or unusually complex.

Do not summarize every file.
Focus on information useful for future development.

模板 2:Bug

Investigate the following bug:

Observed:
...

Expected:
...

Reproduction:
...

First reproduce and identify the root cause.
Do not patch based only on speculation.

Then:
1. implement the smallest robust fix;
2. add a regression test;
3. run relevant tests;
4. report root cause, changes and remaining risks.

模板 3:Feature

Implement:

...

Requirements:
...

Constraints:
...

Acceptance criteria:
...

Before editing, inspect the existing architecture and reuse existing abstractions where appropriate.

After implementation, run the relevant tests and report the verification results.

模板 4:Refactor

Refactor:

...

Goal:
...

The observable behavior must remain unchanged.

Before editing:
- identify current interfaces;
- identify dependencies;
- identify tests covering the behavior.

After editing:
- run regression tests;
- compare behavior before/after where practical;
- report any assumptions.

模板 5:Performance

Goal:

Improve <metric> from X to Y.

Constraints:
- correctness must remain unchanged;
- do not change benchmark methodology;
- do not reduce workload;
- do not disable safety checks.

Workflow:
1. establish baseline;
2. profile;
3. identify bottleneck;
4. make one meaningful optimization;
5. rerun benchmark;
6. compare results.

Continue until the goal is reached or no defensible optimization remains.

模板 6:代码 Review

Review the current diff.

Do not modify files.

Look specifically for:
- correctness bugs
- regressions
- edge cases
- state bugs
- concurrency issues
- performance regressions
- security issues
- inadequate tests

Prioritize findings by severity.

Ignore purely stylistic comments unless they materially affect maintainability.

34. 一个高质量任务应具备的信息

  • Codex 知道我要实现什么吗?
  • Codex 知道为什么要实现吗?
  • Codex 知道什么不能改变吗?
  • Codex 知道如何判断完成吗?
  • Codex 可以运行验证吗?
  • 这是一个独立任务吗?
  • 是否应该创建 worktree?
  • 是否复杂到应该先 plan?
  • 是否长到应该使用 goal?

35. 最终原则

  1. 给目标,不要过度微操实现。
  2. 明确 Definition of Done。
  3. Bug 先复现,再修。
  4. 重要修改必须运行真实验证。
  5. 一个独立任务一个 Chat / Worktree。
  6. 复杂任务先 Plan。
  7. 路径未知、终点明确的长任务使用 Goal。
  8. 长期信息放 AGENTS.md / docs,不要重复塞进 prompt。
  9. AGENTS.md 保持精简,避免积累过时规则。
  10. 把 Codex 的输出当作需要 review 的工程成果,而不是天然正确的答案。

一句话工作流

定义结果
→ 给必要上下文
→ 设边界
→ 定义验收标准
→ Codex 调查
→ 实现
→ 实际运行
→ Review
→ Commit
当你开始按照这个方式使用 Codex 后,真正限制开发速度的因素往往会从“代码敲得够不够快”,转变为:你能否准确描述问题、设计合理系统、判断结果是否正确,以及提出值得解决的问题。