Special Topic · Inside DeepSeek Harness

Model-visible ⟺ logged: An Invariant That Crashes in Your Face

Starting from the byte-for-byte check in invariant.ts: why everything the model sees must be reconstructable from the log

THE QUESTION THIS PAGE ANSWERS

ANSWER FIRST

What is the key idea behind “Model-visible ⟺ logged: An Invariant That Crashes in Your Face”?

Starting from the byte-for-byte check in invariant.ts: why everything the model sees must be reconstructable from the log

DECISION RULE

Make the claim earn its place. Use this page as a decision aid, not a definition to memorize. Connect the idea to one real task, one observable result, and one failure that would change your mind.

TRY NEXT

Write one question you could answer with evidence after trying this idea.

WATCH FOR

A conclusion that sounds complete but leaves the key assumption untested.

Course goalAfter this lesson you can explain three things: why DSH insists a conversation has only one truth, so everything the model sees must be reconstructable from the session log; how it rebuilds and byte-compares on the spot before sending a request; and why, when the check fails, it crashes immediately—too stubborn even to log a warning.
Interactive demo · Log-tampering lab

Play first, then talk. On the left: an append-only event log. On the right: the message array rebuilt from the log, and the request about to go to the model. Hit Play to watch events stream in. When it finishes, play the villain: delete a log entry, or bypass the log and mutate the request, then hit “Send next request”—watch the check tick items one by one, plant a red X at the fork, then crash in your face.

Append-only event logsession-1
(empty log)
Message array rebuilt from the logderiveMessages()
(nothing for the model to see yet)
Request about to be sent to the model
(request not built yet)
Hit “Play” and watch events stream into the log one by one.

The lab’s check order matches the source; the red banner keeps the original source error text: comparison logic in packages/core/agent-loop/src/invariant.ts lines 31–42; error-prefix assembly in packages/runtime-diagnostics/invariants/src/index.ts line 62. Verified on 2026-08-13.

Idea 1 · Single source of truth

What problem it solves.Most chat programs keep two copies of the conversation: an in-memory array and an on-disk archive, each written on its own. One day the process crashes; you restore from the archive, and the recovered history is missing a tool result the model actually saw. Everything the model says next is off, and you can’t prove why—neither copy can vouch for the other. When truth has two homes, they drift.

What the idea is.DSH collapses truth to one copy. The rule sits in the repo-root AGENTS.md line 107:

Model-visible ⟺ logged: anything that reaches a model request must be reconstructable from the session log; a new model-visible input requires a session event. Source: deepseek-harness-master repo AGENTS.md line 107, verified on 2026-08-13

Unpack it. The session log is an append-only event stream; anything you want the model to see must first become an event written into the log. Message history is derived from the log—the official docs say it is “never stored separately.” So in DSH, the log is the conversation itself; there is no second conversation state.

Source: docs/subsystems/session.zh.md line 5, verified on 2026-08-13.

Then the bidirectional arrow ⟺ in the title. The log must imply the request, and the request must be explainable by the log. Writing the log alone isn’t enough—someone has to stand guard at the outbound gate. Before every request leaves, DSH re-derives the expected message array from the log and does a full-string compare against what’s actually on the request; System Prompt, model name, sampling params, and the tool list must also match the request-header snapshot in the log field by field. Only a full match gets through.

Append-only event log Request-header snapshot User message Assistant message Tool call · Tool result Single truth—no copy elsewhere Rebuild on the spot Expected request derived from the log Message array + request-header snapshot Request the loop actually built Deep-frozen—immutable after the check Outbound byte-for-byte compare Match—allow and send Mismatch—crash on the spot
Teaching diagram: nodes and edges explain source relationships; content is course-adapted.

The guard’s core is four lines—short enough to tape on the wall. It proves one thing: before every dispatch, DSH really re-derives messages from the log, full-string-compares them to the actual request, and fails in place on mismatch:

packages/core/agent-loop/src/invariant.tslines 39–42
    const expected = session.deriveMessages()
    if (JSON.stringify(options.messages) !== JSON.stringify(expected)) {
      fail(`llm request for session "${String(session.id)}" diverges from the dispatch-time durable derivation (log-reconstruction desync)`)
    }
Source snapshot note: Based on the local deepseek-harness-master repo; verified against packages/core/agent-loop/src/invariant.ts, verified on 2026-08-13. Code blocks keep the original source text.

One more detail. The derive function used for the check is the same public set used for restore and replay—checker and checked share one rebuild rule; nobody has a private path. Request-header rebuild is just a seven-line pure function: scan events, take the last snapshot.

Source: full check order in packages/core/agent-loop/src/invariant.ts lines 22–52; request-header rebuild in packages/core/session/src/request-header.ts lines 65–71. Verified on 2026-08-13.

Why it lasts.Single source of truth is a fifty-year-old database rule: one ledger only; everything else is a view of that ledger—views can be dropped and rebuilt at will. DSH simply brings that discipline into agent conversation management. Rewrite the source in Rust, triple the event types—still one ledger.

The log is the conversation itself; everything else is a view.
Idea 2 · Fail the check, then crash—don’t keep running while sick

What problem it solves.Imagine the guard finds a mismatch and only logs a warning. A warning means the bad request already reached the model: some plugin bypassed the log and quietly mutated messages; from that moment, what the model sees and what the log records are two different stories. The log keeps writing—wrong books. Three days later someone replays that log to debug and can’t reproduce the weird production behavior. Silent drift is scarier than a crash; it quietly dumps the investigation cost on the future.

What the idea is.So DSH takes the hardest path: fail the compare, throw, void this request on the spot—it never leaves. Design notes explicitly rejected the soft option; the verdict on “compare consecutive requests, warn on divergence” was “rejected because violations must be inexpressible at the interface.”

Source: .agents/notes/implemented/architecture/2026-07-05-reconstructable-requests.zh.md, “Alternatives considered” section, verified on 2026-08-13.

Two small mechanisms ride along. The checker registers at the head of the event-listener queue so no other listener can short-circuit and silently skip the check—the guard sees the request first. Then the request object and message array must be deep-frozen, closing the back door of pass-then-mutate.

Source: head-of-queue registration in invariant.ts lines 20–21 and 54; freeze and session checks in lines 22–29. Verified on 2026-08-13.

Warn path (rejected in design notes) Request diverges from the log Log a warning Bad request still goes out Log books go wrong from here Replay and audit all go false DSH’s choice Request diverges from the log Throw—crash on the spot Request voided—never sent Log stays clean Fix the bug and rerun
Same divergence, two handling styles and their consequences. Teaching schematic.

Why it lasts.That’s fail-fast. A crash locks loss at zero: no poisoned turns in the log—fix the bug and rerun. Keeping running while sick compounds: the longer you go, the more bad data, until you can’t even find when it started. One crash at the boundary is cheaper than archaeology three days into weird symptoms. That judgment doesn’t care about language or framework.

The moment the request leaves, the log can no longer explain the model.
Idea 3 · One log buys you the whole combo

The first two ideas each cost effort; the payoff lands here together. Once the log is the only truth, a whole capability set comes free around it:

  • Restore: process crashed—re-derive from the log and keep chatting.
  • Fork: branch from any event into a parallel session.
  • Replay: re-derive the log and you get the original request—replay tests need no API key.
  • Audit: the UI trail is what the model saw, backed by a runtime assertion.

Even Compaction (compressing a too-long history into a summary) sits under this invariant: the summary is written as an event too; post-compaction requests still face the outbound compare; a buggy Compaction implementation crashes on the spot—no escape.

Why it lasts.This pattern has a name—event sourcing—used for years in banking and accounting: don’t store the balance, store the ledger; the balance is always computed from the flow. DSH just swaps trade events for conversation events. As long as the event stream is the only truth, these capabilities stay free byproducts.

Side-by-side · Everyone logs—who reverse-checks?

Session persistence is almost table stakes for coding agents; the difference is direction. Grok Build (xAI’s open-source Rust coding agent) persists one-way: in-memory conversation state is primary, disk is secondary. Each message copies the in-memory entry into the persist channel; send results are dropped; persist failure doesn’t interrupt the chat; Compaction can even replace the whole on-disk history in one shot. Reasonable product trade-off—but the persistence layer doesn’t own correctness: no path reverse-derives the log and compares it to the outbound request.

Source: grok-build-main repo crates/codegen/xai-grok-shell/src/session/chat_persistence.rs lines 30–38 (persist_message and replace_history), verified on 2026-08-13.

Claude Code is closed-source; what’s public is .jsonl session logs under ~/.claude plus restore. On published evidence, that’s after-the-fact record persistence; nothing public shows a runtime assertion that rebuilds the request from the log at dispatch time and compares—whether an equivalent exists internally is unknown.

So compare mechanism direction only: Grok Build and Claude Code treat persistence as a restore tool; DSH elevates the log to a first principle that must be proven at runtime.

Classroom Exercise
01

Walk a trigger path

A plugin listens for outbound request events and wants to stuff a system prompt into the message array before send. Walk both cases with this lesson’s check order: what happens if it mutates the frozen array directly? If it clones the whole request, mutates the clone, then forwards—where does it get stopped? Hint: what do the freeze check and the byte-for-byte compare each guard?

Takeaway: DSH’s session log is the conversation’s single source of truth—restore, fork, replay, and audit share one event stream. Before every outbound request it rebuilds from the log and byte-compares; mismatch throws on the spot and the request never leaves. Crash beats warn, because the moment the request leaves, the log can no longer explain the model’s behavior.

The handoffs inside “Interactive demo · Log-tampering lab”

“Play first, then talk.” shows that an Agent is not defined by the model alone. Each handoff between model, context, tools, state, permissions, and people affects both progress and recovery.

Write the state before adding capability

Starting from “The lab’s check order matches the source;”, split the workflow into starting state, next action, tool result, state update, and stop condition. Debugging then means finding the first lost piece of information or authority instead of saying vaguely that the model “got worse”.

  • Restore : process crashed—re-derive from the log and keep chatting
  • Fork : branch from any event into a parallel session
  • Replay : re-derive the log and you get the original request—replay tests need no API key

A happy path is not reliability

Use “A plugin listens for outbound request events and wants to stuff a system prompt into the message array before send.” to replay one successful and one failed run. Record the context, tool result, and owner at each turn; the workflow is maintainable when a second person can follow it without the original builder.

From “Interactive demo · Log-tampering lab” to “Idea 1 · Single source of truth”

“Interactive demo · Log-tampering lab” grounds the problem in “Play first, then talk. On the left: an append-only event log. On the right: the message array rebuilt from the log, and the request about to go to the model. Hit Play to watch events stream in. When it finishes…”. “Idea 1 · Single source of truth” then moves it toward “What problem it solves. Most chat programs keep two copies of the conversation: an in-memory array and an on-disk archive, each written on its own. One day the process crashes; you restore from the archive, and…”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.

Carry the judgment into the next situation

When analyzing an Agent, trace state, action, tool result, and next step in order. Each handoff should explain where information came from, who confirmed it, and where failure stops.

  • “Interactive demo · Log-tampering lab”: Play first, then talk. On the left: an append-only event log. On the right: the message array rebuilt from the log, and the request about to go to the model. Hit Play to watch events stream in. When it finishes…
  • “Idea 1 · Single source of truth”: What problem it solves. Most chat programs keep two copies of the conversation: an in-memory array and an on-disk archive, each written on its own. One day the process crashes; you restore from the archive, and…
  • “The closing point”: Audit : the UI trail is what the model saw, backed by a runtime assertion

The final “The closing point” brings the discussion to “Audit : the UI trail is what the model saw, backed by a runtime assertion”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.

Mark as learned Your reading progress updates automatically
← PreviousNext →

Keep reading

The next useful article in the thread.

ARTICLE DISCUSSION

Leave one useful thought here.

Keep the idea that clicked, the question that stayed open, or a small note for the next learner.

Discussing Model-visible ⟺ logged: An Invariant That Crashes in Your Face Inside DeepSeek Harness
3discussionsArticle discussion · synced with the Circle
View in the learning circle
AM
Asha MorganContent editor
INSIGHTField note

I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.

ARTICLE DISCUSSION7 helpful
LH
Lin HarperIndie developer
INSIGHTInsight

After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.

ARTICLE DISCUSSION5 helpful
KM
Kiki MooreProduct operations
QUESTIONQuestion

When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.

ARTICLE DISCUSSION4 helpful