# 上下文与历史 · 只追加、不改写

> 发给模型的历史由 ContextManager 保管，只追加、不改写。环境、权限、协作模式、AGENTS.md 这些会变的上下文被建模成“世界状态”的各个分区：第一轮注入完整快照，之后只追加变化的部分，连换模型、改 AGENTS.md 也是追加一条新消息。真正改写历史的只有压缩与回滚，而且会同时清掉基准、下一轮重新注入完整上下文。

- 作者：David（道雾轩）
- 专栏：Codex 源码解读（https://daiw.net/manual/codex-source.md）
- 最后更新：2026-09-29
- 原文：https://daiw.net/manual/codex-source/context-history
- 转载与引用：请注明出处并附原文链接（https://daiw.net/about/copyright）

# 上下文与历史 · 只追加、不改写

[模型客户端](https://daiw.net/manual/codex-source/model-client)那篇提到，请求设了 `store: false`，走 HTTP 时每次都把完整历史发给模型。于是“历史长什么样”直接决定了两件事：模型看到什么，以及提示缓存能不能命中。仓库根的 `AGENTS.md` 专门为此立了规矩：

```md
### Model visible context

Codex maintains a context (history of messages) that is sent to the model in inference requests.

1. No history rewrite - the context must be built up incrementally.
2. Avoid frequent changes to context that cause cache misses.
3. No unbounded items - everything injected in the model context must have a bounded size and a hard cap.
4. No items larger than 10K tokens.
5. Highlight new individual items that can cross >1k tokens as P0. These need an additional manual review.
6. All injected fragments must be defined as structs in `core/context` and implement ContextualUserFragment trait
```

（`AGENTS.md:91`）

这一篇看代码怎样兑现这六条。

## 用户看到的样子

界面上的对话只显示你说的话、模型的回复与工具活动；内核悄悄插进历史的环境信息、权限说明、AGENTS.md 等内容不会显示成“用户消息”。想看模型实际收到的输入，可以运行 `codex debug prompt-input`，它以 JSON 列出输入条目（见 [AGENTS.md 项目指令](https://daiw.net/manual/codex/agents-md)）。开头通常是几条 developer 与 user 消息，装着 `<permissions instructions>`、`<environment_context>`、`# AGENTS.md instructions` 这样带标记的片段，后面才是真正的对话。

## ContextManager：历史的保管者

历史放在 `SessionState` 的 `history` 字段里，类型是 `ContextManager`：

```rust
/// Transcript of thread history
#[derive(Debug, Clone, Default)]
pub(crate) struct ContextManager {
    /// The oldest items are at the beginning of the vector. Snapshots share the vector until a
    /// caller needs to mutate it, avoiding deep copies for read-only history consumers.
    items: Arc<Vec<ResponseItemEnvelope>>,
    // ...
    /// Bumped whenever history is rewritten, such as compaction or rollback.
    history_version: u64,
    // ...
    reference_context_item: Option<TurnContextItem>,
    /// World state most recently appended to model-visible history.
    world_state_baseline: Option<WorldStateSnapshot>,
}
```

（`codex-rs/core/src/context_manager/history.rs:79`）

- `items` 是 `ResponseItem` 的有序列表，每项外面包一层 `ResponseItemEnvelope`，附带只给内核看的元数据（截断预算、来源、受理顺序等）。它放在 `Arc` 里写时复制，`run_turn` 每次采样前 `clone_history()` 拿一份快照，代价很小。
- 写入口只有追加：`Session::record_conversation_items` 的注释是 “Appends to history, persists the prepared items, then notifies raw-item observers.”——先追加到内存历史，再写 rollout，最后通知观察者。
- 读出口是 `for_prompt`：它不改原历史，而是在副本上做规范化——每个工具调用都要有输出（缺了就补一条内容为 `aborted` 的输出），每个输出都要有对应的调用，再按模型支持的模态剥掉图片或音频。补出来的输出只进这次的输入、不写回历史，它的 ID 由对应调用条目的 ID 确定性地算出，`normalize.rs` 里固定命名空间的注释写着“改动这个值会改变模型可见的 ID，让提示缓存失效”。
- `history_version` 只在改写时递增。`reference_context_item` 与 `world_state_baseline` 是算“差异”的两个基准：前者的注释说它是“下一个普通轮次的基准”，为 `None` 时下一轮会重新注入完整上下文；后者记着最近一次追加进历史的世界状态，下文细说。

## 注入的片段：ContextualUserFragment

内核往历史里插的每一段上下文，都是实现了 `ContextualUserFragment` 的结构体。trait 定义在独立的 `codex-context-fragments` crate，由 `core/src/context/mod.rs` 重新导出；具体片段大多按规则第 6 条放在 `core/src/context/` 下（57 个实现），也有少数定义在别处，比如权限说明 `PermissionsInstructions` 在 `codex-prompts`，技能说明在技能扩展里：

```rust
pub trait ContextualUserFragment {
    fn role(&self) -> &'static str;

    /// Returns a stable `<feature>.<name>` classification, using `generic` for shared fragments.
    fn content_kind(&self) -> ContentItemKind;

    /// Whether this fragment must be recorded as its own response item.
    fn requires_separate_message(&self) -> bool {
        false
    }

    fn markers(&self) -> (&'static str, &'static str);

    fn body(&self) -> String;
```

（`codex-rs/context-fragments/src/fragment.rs:64`）

`render()` 把起止标记夹在正文两边。标记有两个用途：一是识别，`event_mapping.rs` 把历史条目翻成界面上的 `TurnItem` 时，带这些标记的 user 消息都不当作用户消息展示；二是回滚时认出哪些是可以一并裁掉的“轮前上下文”。`content_kind` 给每段内容一个稳定分类（如 `agents_md.instructions`），作为元数据跟着消息走。几个常见片段：

| 片段 | 角色 | 标记 |
| --- | --- | --- |
| AGENTS.md 指令 `UserInstructions` | user | `# AGENTS.md instructions` 与 `</INSTRUCTIONS>` |
| 环境信息 | user | `<environment_context>` |
| 权限说明 | developer | `<permissions instructions>` |
| 协作模式 | developer | `<collaboration_mode>` |
| 换模型说明 `ModelSwitchInstructions` | developer | `<model_switch>` |
| 中断提示 `TurnAborted` | user | `<turn_aborted>` |
| 用户 shell 命令输出 | user | `<user_shell_command>` |

## 世界状态：第一轮全量，之后只发差异

环境、权限、协作模式、AGENTS.md 这些东西会在会话中途变化，又必须让模型知道。Codex 把它们建模成“世界状态”：`WorldState`（`codex-rs/core/src/context/world_state/mod.rs:285`）由一组分区组成，每个分区实现 `WorldStateSection`，有一个写进 rollout 的稳定 `ID`、一个用于比较的 `Snapshot`，以及 `render_diff(previous)`——给定上一次模型看到的快照，决定这次要追加什么，或者什么都不追加。`build_world_state_for_step`（`codex-rs/core/src/session/world_state.rs:40`）每个 step 都按当前设置重建一遍：模型、实时语音、AGENTS.md、权限、协作模式、环境、应用与插件说明、多 agent 模式、托管的开发者指令，外加扩展贡献的分区。

```mermaid
flowchart TB
  A[run_turn 开轮 · 捕获 step] --> B[build_world_state_for_step]
  B --> C{reference_context_item<br/>有基准吗}
  C -->|没有 · 第一轮或压缩后| D[render_full 全量渲染<br/>拼成初始上下文消息]
  C -->|有| E[对每个分区 render_diff<br/>只渲染变化的部分]
  E --> F[merge_contextual_fragments<br/>同角色的片段合成一条消息]
  D --> G[追加进历史]
  F --> G
  G --> H[rollout 写入 WorldState 快照或补丁<br/>与本轮 TurnContextItem]
  H --> I[后续每个 step<br/>record_step_world_state_if_changed]
```

没有基准时（第一轮，或压缩清掉了基准），`build_initial_context_with_world_state` 把全部分区渲染出来，按角色分组：developer 片段合成一条大的 developer 消息（`<model_switch>` 要排在最前），AGENTS.md 与环境信息合成一条 user 消息，另有几类需要单独成条的放在后面。有基准时，每个分区只和上次的快照比，变了才出片段。快照写进 rollout（`RolloutItem::WorldState`），全量注入时写完整快照，之后写 RFC 7386 合并补丁；恢复会话时据此还原基准，继续只发差异。一轮之内，后续每个 step 还会调用 `record_step_world_state_if_changed`，把中途发生的变化同样以追加的方式补上。

`AgentsMdState` 的 `render_diff` 是个好例子：

```rust
        let previous_may_contain_instructions = match previous {
            PreviousSectionState::Known(previous) => previous.text.is_some(),
            PreviousSectionState::Unknown => true,
            PreviousSectionState::Absent => false,
        };
        let instructions = match (&self.instructions, previous_may_contain_instructions) {
            (Some(instructions), true) => UserInstructions {
                directory: instructions.directory.clone(),
                text: format!("{REPLACEMENT_NOTICE}\n\n{}", instructions.text),
            },
            (Some(instructions), false) => instructions.clone(),
            (None, true) => UserInstructions {
                directory: None,
                text: REMOVAL_NOTICE.to_string(),
            },
            (None, false) => return None,
        };
        Some(Box::new(instructions))
```

（`codex-rs/core/src/context/world_state/agents_md.rs:61`）

这段之前还有一道判断：快照与上次完全相同就返回 `None`，什么都不追加。内容变了，不去改历史里那条旧的 AGENTS.md 消息，而是在末尾追加一条新的，开头声明 “These AGENTS.md instructions replace all previously provided AGENTS.md instructions.”；文件被删了，就追加一句 “The previously provided AGENTS.md instructions no longer apply.”。其他分区同理：权限分区若只是多了几条已批准的命令前缀，只追加这几条前缀；环境分区只渲染变化的字段。

## 换模型、改推理强度，也只追加

基础指令在线程创建时就定下，存进会话元数据，此后 `SessionConfiguration.base_instructions` 再不改动。用户中途换了模型怎么办？模型分区 `ModelInstructionsState` 发现模型变了，就追加一条 `<model_switch>` developer 消息，正文是 “The user was previously using a different model. Please continue the conversation according to the following instructions:” 加上新模型的指令。请求开头的基础指令保持原样，这一段的字节就不会因为换模型而变。

推理强度同理：在开发中的 `reasoning_effort_override` 功能下，若提供方是 OpenAI 且模型支持，请求里的 `reasoning.effort` 在一个上下文窗口内固定为最初的值，中途的调整改为向历史追加一条 `ConfigurationUpdate` 条目（`codex-rs/core/src/session/reasoning_effort.rs`，文件头注释就叫“Cache-preserving effort updates”）。

## 有界：每一项都有上限

规则第 3、4 条要求注入的内容有硬上限，代码里随处可见对应的常量。工具输出写进历史时按模型的截断策略截断（随附目录里都是 10000 token），只截内存里给模型看的这份，rollout 保留完整输出；项目级 AGENTS.md 合计不超过 `project_doc_max_bytes`（默认 32 KiB），宿主提供的线程级指令超过 10000 个估算 token 直接拒绝（`MAX_THREAD_INSTRUCTIONS_TOKENS`）；环境信息里的子代理清单最多列 8 个、不超过 1024 字节（`codex-rs/core/src/session/world_state.rs:35`）；压缩后保留的用户消息也有预算，见[下一篇](https://daiw.net/manual/codex-source/compaction)。

## 什么时候会改写

真正改写历史的只有两处。一是压缩：`replace_compacted` 用摘要后的新历史整体替换旧历史，`history_version` 加一，世界状态基准随之重置，按注入方式决定是否同时写入新的 `reference_context_item`；二是回滚：从 rollout 重建历史时遇到 `ThreadRolledBack` 标记，`drop_last_n_user_turns` 裁掉最后几轮，若裁到了第一轮那条混合的初始上下文消息，就清掉 `reference_context_item`，让下一轮重新注入完整上下文。两者都是一次性的、显式的替换，替换后仍按只追加的规则继续生长。

## 和《从 LLM 到 Coding Agent》对照

[系统提示与上下文注入](https://daiw.net/manual/llm-to-agent/context-injection)给出的黄金法则是“静态放高处，动态放低处”，并警告不要原地改要复用的消息。Codex 把这条法则推到了底：会变的上下文一律不进系统提示，而是作为消息追加在历史末尾，连“换模型”“AGENTS.md 改了”这种看似要改开头的事，也用一条新的声明消息来表达。[Prompt 缓存](https://daiw.net/manual/llm-to-agent/prompt-caching)讲的“字节稳定性红线”，在这里落实为世界状态的差异渲染、确定性的合成 ID，以及把改写收窄到压缩与回滚两个出口。

---

上一篇：[模型客户端 · Responses API 的流式调用](https://daiw.net/manual/codex-source/model-client) · 下一篇：[上下文压缩 · 窗口快满时怎么办](https://daiw.net/manual/codex-source/compaction)
