Managed Agent: Brain-Hand Separation
Split thinking and execution into different processes — virtualizing Agents like an operating system
THE QUESTION THIS PAGE ANSWERS
ANSWER FIRSTWhat is the key idea behind “Managed Agent: Brain-Hand Separation”?
Split thinking and execution into different processes — virtualizing Agents like an operating system
Make the claim earn its place. Use this page as a decision aid, not a definition to memorize. Connect the idea to one real task, one observable result, and one failure that would change your mind.
Write one question you could answer with evidence after trying this idea.
A conclusion that sounds complete but leaves the key assumption untested.
Container dies = session lost = can't debug = task completely fails.
Like a pet: if it dies, it's over — no replacement possible.
Container dies = tool call fails = Claude decides to retry = spin up a new container and continue.
Like cattle on a farm: one dies, spin up another — the system keeps running.
-
Containers become tool calls:
execute(name, input) -> string— to the Harness, it's just a regular function - Container dies = tool call returns an error = Claude decides whether to retry = automatically spins up a new container to continue
- Harness can start processing before the container is ready — no need to wait for container startup
TTFT = Time to First Token
- Old approach: Agent-generated code and API keys live in the same container — Prompt Injection can steal keys directly
- Git Token: injected as container environment variables when cloning the repo; usable inside the sandbox, but the Agent never sees the token value
- MCP OAuth Token: stored in an external vault, MCP calls are forwarded via a proxy — the sandbox cannot access the token directly
This means you can have one Brain control multiple Sandboxes simultaneously (parallel execution), or pass the same Sandbox between different Brains (relay execution). Components are fully decoupled.
How “Analogy: OS Virtualization” changes an answer
“This means you can have one Brain control multiple Sandboxes simultaneously (parallel execution), or pass the same Sandbox between different Brains (relay execution).” shows that a model does not process the “word count” we see. It processes Token pieces. Tokenization affects input length, how much context fits, and how much computation a request consumes.
Length, information, and context are different
As “This means you can have one Brain control multiple Sandboxes simultaneously (parallel execution), or pass the same Sandbox between different Brains (relay execution).” grows, separate three questions: how many Tokens the text becomes, which pieces can change the current decision, and whether older material has fallen outside the context window. Removing repetition is often more useful than simply making the window larger.
- Containers become tool calls : execute(name, input) -> string — to the Harness, it's just a regular function
- Container dies = tool call returns an error = Claude decides whether to retry = automatically spins up a new container to continue
- Harness can start processing before the container is ready — no need to wait for container startup
Keep what can change the decision
Use “This means you can have one Brain control multiple Sandboxes simultaneously (parallel execution), or pass the same Sandbox between different Brains (relay execution).” as an A/B test: keep the same question while removing repeated background, compressing format, and trimming irrelevant history. Compare answer quality, latency, and Token count.
From “Analogy: OS Virtualization” to “Three Core Components”
“Analogy: OS Virtualization” grounds the problem in “Operating System Virtualizes Hardware When you call read() , you don't care whether the underlying storage is SSD, HDD, or a network drive — the OS abstracts away hardware details. Managed Agent does the same t…”. “Three Core Components” then moves it toward “Session Event Log An append-only event stream with persistent storage. Records everything that has happened: user input, tool calls, and model output. Harness Brain The loop that calls Claude and routes tool ca…”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.
Carry the judgment into the next situation
For long text, keep what can change the conclusion before compressing format and history. A larger context is worth its cost only when the added information is useful.
- “Analogy: OS Virtualization”: Operating System Virtualizes Hardware When you call read() , you don't care whether the underlying storage is SSD, HDD, or a network drive — the OS abstracts away hardware details. Managed Agent does the same t…
- “Three Core Components”: Session Event Log An append-only event stream with persistent storage. Records everything that has happened: user input, tool calls, and model output. Harness Brain The loop that calls Claude and routes tool ca…
- “The closing point”: Git Token : injected as container environment variables when cloning the repo; usable inside the sandbox, but the Agent never sees the token value
The final “The closing point” brings the discussion to “Git Token : injected as container environment variables when cloning the repo; usable inside the sandbox, but the Agent never sees the token value”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.
I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.
After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.
When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.
No discussion on this article yet.