Special Topic · Inside DeepSeek Harness

DSH-Only Tool Surface: terminal / lsp / jobs

What the tools others lack each solve

THE QUESTION THIS PAGE ANSWERS

ANSWER FIRST

What is the key idea behind “DSH-Only Tool Surface: terminal / lsp / jobs”?

What the tools others lack each solve

DECISION RULE

Make the claim earn its place. Use this page as a decision aid, not a definition to memorize. Connect the idea to one real task, one observable result, and one failure that would change your mind.

TRY NEXT

Write one question you could answer with evidence after trying this idea.

WATCH FOR

A conclusion that sounds complete but leaves the key assumption untested.

Course goalbash plus read/write already gets work done — why does DSH still ship three first-class capability sets? After this you can explain three things: how terminal keeps state alive across a call, why lsp only offers four semantic queries, and how jobs fold unrelated background tasks into one list. Plus their shared design discipline: closed vocabulary, structured degradation, owner-based auth.
Interactive demo · Tool-surface comparison room

Same task: start a Python REPL and debug a snippet in three steps. Scenario A gives only bash on the left and the terminal six-pack on the right — watch the gap in rounds and repeated work. Scenario B demos the jobs panel: three totally different background tasks managed by one list.

Left · one-shot bash onlyState cleared every call
Tool Calls 0 roundsRe-ran 0 lines of old code
Right · terminal six-packPersistent PTY, state kept
Tool Calls 0 roundsRe-ran 0 lines of old code
job_list view (background tasks under the current agent)
Hit “Play” to start the comparison, or scroll here to auto-play.
Teaching simulation. Left/right output and tool names match real schemas for bash and terminal_* in docs/tool-catalog.zh.md; Scenario B task shapes map to the three tools in @deepseek-ai/dsh-tool-jobs. SEND_ACTIVE error behavior maps to line 246 of packages/terminal/terminal/src/index.ts.
Three capability sets, each curing one disease
terminal · cures state amnesia
  • Six tools: terminal_open / send / read / signal / close / list, backed by a persistent PTY session.
  • Stateful programs like REPL, gdb, ssh stay alive across calls.
  • Each session belongs to an exact Agent instance — another agent with the id still can’t operate it.
lsp · cures text search without semantics
  • Exactly four query types: go-to-def, find-refs, go-to-impl, hover — a closed union; adding a fifth is a compile-time breaking change.
  • Deliberately no generic JSON-RPC escape hatch — the model can’t invent protocol tricks.
  • With no provider, schema stays unchanged and returns structured LSP_UNAVAILABLE.
jobs · cures background tasks with no one to clean up
  • Background bash, background PTY sends, background subagents — all register as <kind>-N tasks.
  • One trio — job_list / job_output / job_kill — manages everything.
  • Auth checks the owner session — predictable ids are fine; each owner defaults to at most 10 concurrent.
terminal: keep state alive past one call

Start with the most underrated: terminal. The one-shot bash problem already showed up in the demo: REPL variables don’t survive a call, so the model must resend all old code — burning rounds and tokens twice. DSH’s fix makes the PTY session a first-class resource: terminal_open creates a session and returns an id; later send/read/signal/close all key off that id; the session stays alive across backend and tool-plugin hot reloads (see “Ownership & persistence” in docs/subsystems/terminal.zh.md).

The tool description itself is a prompt-engineering model. In the 1878-line tool catalog, the terminal_send entry clears wait semantics in one sentence:

docs/tool-catalog.zh.md · line 878 (terminal_send tool description)
“Send text to a persistent terminal. By default submits Enter and waits for a prompt, stdin wait, output silence, timeout, or session exit. Background mode returns a job id for job_output / job_kill.”

Then concurrency discipline: one PTY session accepts only one active send at a time. Two calls racing the same terminal — the second gets a structured error; nobody sneaks characters into someone else’s command.

Watch startSend()’s gate order — all three checks are mandatory. First through the door: expectOwned(owner, id) — auth compares the exact Agent instance that owns the session; the id isn’t a secret, the boundary is ownership; another agent that guesses the id still fails. Second: is the session closing? A half-closed session rejects any new input. Third: record.active — if the previous send hasn’t settled, the new one throws TerminalError with SEND_ACTIVE; the model sees a structured failure and should wait for settlement. After all checks pass, this op is registered as the session’s only active send; only when its done promise settles (success or failure) does the registry clear and the next send qualify. Ownership-is-boundary shows up again in jobs.

Source: packages/terminal/terminal/src/index.ts lines 243–254 (startSend()), verified on 2026-08-13.

lsp: four operations, not one more

The lsp tool narrows language-server power to four semantic queries. Why not expose the whole LSP protocol? Because the model would face a bottomless schema, and swapping providers would drift behavior. DSH closes it: seam, provider, and tool share the same four-op union — adding a fifth fails compile until all three layers change. Routing is blunt: find the provider by file extension; one extension belongs to one house:

packages/lsp/lsp/src/index.tslines 143–149 excerpt
  async query(request: LspQueryRequest, signal?: AbortSignal): Promise<LspQueryResult> {
    const route = this.routes.get(finalExtension(request.filePath))
    if (route === undefined) {
      throw new LspError(`no LSP provider handles "${request.filePath}"`, 'LSP_UNAVAILABLE')
    }
    return route.provider.query({ ...request, languageId: route.languageId }, signal)
  }
Source snapshot note: Based on the local deepseek-harness-master repo; verified against packages/lsp/lsp/src/index.ts, verified on 2026-08-13. Code blocks keep the original source text.

The key when there’s no provider: the lsp tool stays in the catalog, schema unchanged, and the call returns a structured failure with LSP_UNAVAILABLE. The model learns this project has no language server — no guessing why a tool vanished. Degradation stays structured, vocabulary stays stable — same spirit as bash reporting [sandbox: …].

jobs: three background task kinds, one list

Finally jobs. Background bash, terminal_send’s background mode, background subagents — three producer shapes, totally different. DSH registers them all into one ctx.jobs registry, issues ids like bash-1 / subagent-2 (kind + sequence), and the model uses one job_list / job_output / job_kill set for all. Producers own execution resources; the registry owns identity, access, and lifecycle state (docs/subsystems/jobs.zh.md).

Auth doesn’t rely on secret ids

Ids issue as <kind>-N in order — fully predictable. The defense is owner auth: read, kill, wait all verify the caller’s session matches the task owner; you can’t even see another’s label.

done waits for resource release

The producer’s done promise resolves only after resources are released — finishing work isn’t enough. When an owner is destroyed, the registry cancels and awaits tasks; no orphan processes.

Completion notices don’t double-ping

Once an interface has delivered a terminal state, the reported flag suppresses duplicate completion notices — so one finished task doesn’t give the model two messages and two wasted turns.

Side-by-side · Products can skip it; runtimes should do it
Grok Build: standalone ptyctl crate

On PTY, Grok Build and DSH landed in the same place: the repo has a standalone ptyctl crate (crates/codegen/ptyctl/ with pty, session, server, term, wait modules, plus ptyctl-cli), making terminal control reusable infrastructure. Both treat stateful terminals as first-class; the difference is composition: Grok links a Rust crate at compile time; DSH mounts a runtime plugin plus six model-facing tools.

Claude Code: no first-class PTY or LSP

Claude Code’s reconstructed tool list (manuscript ch. 2) has no first-class PTY session tool and no LSP tool — based on public reconstructed-source evidence. It does have background power: bash with run_in_background, and agent tasks can auto-background by duration (tengu_auto_background_agents, ch. 2 lines 448–458) — but those are built around bash and subagent separately, with no cross-kind unified job registry. The gap isn’t laziness: CC is a product — interactive debug has the IDE, semantic nav has the editor; DSH is a runtime that must feed these capabilities to the model alone in headless environments. What a product can skip, a runtime should do. One more beat: none of these three sit on the agent-loop trunk — all optional plugins; the minimal preset mounts none and is still a complete coding agent.

Classroom Exercise
01

Walk through two boundary scenarios

Scenario 1: on the same terminal session the agent first sends with run_in_background: true, then immediately a foreground send. What happens to the second, which error code, and what does the job panel show? Scenario 2: the project has no language server; the model calls lsp once for the definition of a.py. Write the result shape the model gets, and why that’s friendlier than yanking lsp from the catalog.

Takeaway: terminal keeps state alive across calls; lsp offers four semantic queries and refuses a protocol escape hatch; jobs fold three background kinds into one owner-authorized list. Shared discipline: closed vocabulary, structured degradation, ownership-is-boundary — and all optional plugins, so the trunk carries no baggage.

The handoffs inside “Interactive demo · Tool-surface comparison room”

“Same task: start a Python REPL and debug a snippet in three steps.” shows that an Agent is not defined by the model alone. Each handoff between model, context, tools, state, permissions, and people affects both progress and recovery.

Write the state before adding capability

Starting from “Start with the most underrated: terminal.”, split the workflow into starting state, next action, tool result, state update, and stop condition. Debugging then means finding the first lost piece of information or authority instead of saying vaguely that the model “got worse”.

  • Six tools: terminal_open / send / read / signal / close / list , backed by a persistent PTY session
  • Stateful programs like REPL, gdb, ssh stay alive across calls
  • Each session belongs to an exact Agent instance — another agent with the id still can’t operate it

A happy path is not reliability

Use “Scenario 1: on the same terminal session the agent first sends with run_in_background: true , then immediately a foreground send.” to replay one successful and one failed run. Record the context, tool result, and owner at each turn; the workflow is maintainable when a second person can follow it without the original builder.

From “Interactive demo · Tool-surface comparison room” to “Three capability sets, each curing one disease”

“Interactive demo · Tool-surface comparison room” grounds the problem in “Same task: start a Python REPL and debug a snippet in three steps. Scenario A gives only bash on the left and the terminal six-pack on the right — watch the gap in rounds and repeated work. Scenario B demos the…”. “Three capability sets, each curing one disease” then moves it toward “terminal · cures state amnesia Six tools: terminal_open / send / read / signal / close / list , backed by a persistent PTY session. Stateful programs like REPL, gdb, ssh stay alive across calls. Each session be…”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.

Carry the judgment into the next situation

When analyzing an Agent, trace state, action, tool result, and next step in order. Each handoff should explain where information came from, who confirmed it, and where failure stops.

  • “Interactive demo · Tool-surface comparison room”: Same task: start a Python REPL and debug a snippet in three steps. Scenario A gives only bash on the left and the terminal six-pack on the right — watch the gap in rounds and repeated work. Scenario B demos the…
  • “Three capability sets, each curing one disease”: terminal · cures state amnesia Six tools: terminal_open / send / read / signal / close / list , backed by a persistent PTY session. Stateful programs like REPL, gdb, ssh stay alive across calls. Each session be…
  • “The closing point”: Deliberately no generic JSON-RPC escape hatch — the model can’t invent protocol tricks

The final “The closing point” brings the discussion to “Deliberately no generic JSON-RPC escape hatch — the model can’t invent protocol tricks”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.

Mark as learned Your reading progress updates automatically
← PreviousNext →

Keep reading

The next useful article in the thread.

ARTICLE DISCUSSION

Leave one useful thought here.

Keep the idea that clicked, the question that stayed open, or a small note for the next learner.

Discussing DSH-Only Tool Surface: terminal / lsp / jobs Inside DeepSeek Harness
3discussionsArticle discussion · synced with the Circle
View in the learning circle
AM
Asha MorganContent editor
INSIGHTField note

I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.

ARTICLE DISCUSSION7 helpful
LH
Lin HarperIndie developer
INSIGHTInsight

After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.

ARTICLE DISCUSSION5 helpful
KM
Kiki MooreProduct operations
QUESTIONQuestion

When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.

ARTICLE DISCUSSION4 helpful