Docker’s agent runtime, docker/docker-agent, has been quietly accumulating stars since September 2025 and is now a serious piece of engineering: Apache-2.0, actively pushed to this week, and built around a permission model that is genuinely well thought out. Four safety modes. A deny-allow-ask evaluation chain. Argument-level matching so you can approve shell only when the command starts with ls. Global user-level rules that an agent author cannot override. It is the most complete tool-permission system we have seen ship with an agent runtime.
So we pulled the repository and counted how many of the example agent configurations actually use any of it.
TL;DR
- We audited all 196 example agent configurations in
docker/docker-agentat commitdd39af0(8 October 2026): 281 agents, 90 filesystem toolsets, 74 shell toolsets. - 87 configurations hand an agent a host-reaching toolset. 10 declare any guard at all. Five contain a
permissions:block, four declare a safety mode, one setsruntime.sandbox. - The
--sandboxflag, the one Docker’s own documentation calls “the isolation boundary to reach for”, injects--yolointo the agent process running inside the VM unless you already passed a safety flag yourself. - Because the injected flag is an explicit CLI flag on the inner process, it outranks
settings.safetyin your user config, the personal floor Docker tells you to set. - An agent config can opt itself into the sandbox with
runtime.sandbox: trueand reach the same unattended state without the word “autonomous” appearing anywhere in the YAML. - In headless CI, an unconfigured run fails the other way: every tool call is rejected because there is no stdin to prompt. The broken state and the unrestricted state are one flag apart.
What we measured
We downloaded the repository tarball at commit dd39af0 and parsed every YAML file under examples/ with a real YAML parser rather than grep, because grep over configuration files produces exactly the kind of substring false positive that makes an audit worthless. Of 196 files, all 196 parsed as agent configurations. Between them they define 281 agents.
The toolset distribution tells you what these agents are for:
filesystem: 90shell: 74mcp: 57think: 30todo: 12fetch: 9
These are not chatbots. Filesystem and shell are the two most common toolsets in the repository by a wide margin. Counting configurations rather than toolsets, 87 of the 196 grant at least one toolset that reaches the host: shell, filesystem, script or background_jobs.
Against those 87, we looked for any of the four governance mechanisms the runtime provides: a top-level permissions: block, a runtime.safety default, a per-agent safety field, or runtime.sandbox. Ten configurations declare one. Here they are, in full, because a number like that is worth being able to check:
permissions:block, five files:gopher.yaml,llm_judge.yaml,mcp-toolkit.yaml,modernize-go-tests.yaml,permissions.yaml- Per-agent
safety, four files:evaluators.yaml,evaluators-laya.yaml,evaluators-openai.yaml,safety_modes.yaml runtime.sandbox: true, one file:sandbox_agent.yamlruntime.network_allowlist, one file:sandbox_agent.yaml, the same one
Notice the pattern. permissions.yaml, safety_modes.yaml and sandbox_agent.yaml are the examples whose entire purpose is to demonstrate the feature. Strip those out and six working agents in the repository use any guard. The remaining 77 configurations grant shell or filesystem access and declare nothing.
That is not a criticism of the examples as examples. They exist to show a feature each, and loading every one with boilerplate permissions would obscure the thing being demonstrated. It is a statement about what gets copied. Nobody starts an agent config from a blank file. They start from coder.yaml or dev-team.yaml, both of which grant filesystem and shell with no scope, and then they change the instruction text.
The flag that turns the prompts off
Here is the part that matters more than the counting.
Docker’s headless guide is admirably clear about the distinction between two questions: what is allowed to run without asking (safety modes and allow-lists), and what happens if the model runs something it should not have (only the sandbox). It then states, correctly, that --sandbox is the isolation boundary to reach for, and that an allow-list is defence in depth rather than a boundary.
What --sandbox also does is in cmd/root/sandbox.go. When the wrapper builds the argument list for the agent process it will run inside the VM, it ends with this:
if !hasYolo && !hasSafety && !cmd.Flags().Changed("session") {
dockerAgentArgs = append(dockerAgentArgs, "--yolo")
}
--yolo is the legacy spelling of the autonomous safety mode: every tool call runs, no confirmation, safe or destructive or unknown. So the flag you reach for to make an agent safer also, in the same action, removes every prompt that was standing between the model and your working directory. Docker documents this in the headless guide and gives the reasoning: inside a VM the blast radius is contained, so unattended operation is reasonable. That reasoning is sound. The problem is that nobody reads --sandbox as “and also approve everything”.
Two details make it sharper than a naming quibble.
First, the injection beats your own default. Docker tells you to pin a personal floor with settings.safety in ~/.config/cagent/config.yaml, and the permissions documentation is emphatic that author-declared defaults never outrank a user-owned source. That precedence holds, but the sandbox wrapper only inspects flags that were set on the command line. It never consults your user settings. The --yolo it appends arrives at the inner process as an explicit CLI flag, and in explicitCLISafety() an explicit flag wins over the alias and settings default. Your configuration directory is mounted into the VM read-only, so the file is right there, being outranked.
Second, an agent author can trigger it. runtime.sandbox is a boolean field in the config schema, present in every schema version from v9 to current. When the flag is not set on the command line, run.go resolves the sandbox decision from the agent configuration instead, and routes through the same runInSandbox path with the same injection. The permissions documentation carries a warning about pulling a config from a URL or an OCI registry that declares safety: autonomous. A config that declares runtime.sandbox: true reaches the same place without that word appearing. There is a VM in the way, which is the whole point and a real mitigation, but “VM with every tool auto-approved” is not what most people picture.
Worth knowing alongside it: the sandbox VM is retained and reused across runs rather than torn down when the session ends, and for local sandboxes your working directory is mounted read-write. The isolation is around the host, not around your repository.
The other failure mode: nothing runs at all
The opposite mistake is just as common and much more confusing to debug. Run an agent headless with --exec, no safety flag and no allow-list, and the session falls back to the historical default: read-only tools auto-approve, everything else asks. There is no stdin in CI, so the runtime answers no on your behalf, and every tool call the model attempts is rejected. The run completes. The exit code is fine. The agent just quietly did nothing, and the model’s increasingly confused attempts to work around its own blocked tools are the only clue in the log.
So the unconfigured state is broken and the obvious fix is unrestricted, with the correct middle ground, --safety restricted plus a narrow allow-list, being the thing almost nobody writes. Our count says six working examples in the repository write it.
One more trap for anyone wiring this into a pipeline: --on-event does nothing under --exec. Event hooks are installed on the interactive application’s event bus, and the headless path returns before that wiring happens. Docker flags this in its documentation, and the code bears it out, the hook options are appended well inside the interactive branch. If your plan for an audit trail is an --on-event hook appending to a log, in CI you will get an empty log and no error. Parse the --json event stream instead.
What to actually do
None of this is a reason to avoid Docker Agent. The permission model is better than most of what is shipping right now, and the documentation is unusually honest about where the boundaries are. The gap is between what the runtime can enforce and what the configurations in circulation actually ask it to enforce.
- Set a floor you can see. Put a
denylist insettings.permissionsin your user config and treat it as the only guard that survives a copied config, because deny patterns cannot be overridden from either direction. - Pass
--safetyexplicitly whenever you pass--sandbox. Setting it suppresses the injection. If you want the unattended behaviour, write--yoloyourself so the next person reading the command knows it is there. - Review
runtime.sandboxandruntime.safetyin any config you did not write. They are two lines in a file that is mostly prose, and they change how every tool call in that file is gated. - Use
restrictedfor CI, notautonomous. Safe calls run silently, everything else is denied without prompting. Then widen with named allow patterns until the job passes, which gives you a written record of exactly what the agent needs. - Do not treat the allow-list as the boundary. Docker says this itself. Command matching is string matching, and a shell is very good at producing strings that do not look like what they do.
The broader point generalises past Docker. Every agent runtime shipping this year has a permission system, and in every one of them the safe configuration is opt-in, undocumented in the examples, and one flag away from the unsafe one. The default that matters is not the one in the manual. It is the one in the file people copy.
At REPTILEHAUS we build and operate agent infrastructure for clients who need it to run unattended without that being a liability, covering permission policy, sandboxing, CI integration and the DevOps work that makes an agent pipeline auditable rather than merely functional. If you are putting coding agents into a build pipeline and are not certain what they are allowed to touch, get in touch.
Audit performed 8 October 2026 against docker/docker-agent at commit dd39af0. Counts are from a YAML parse of all 196 configurations under examples/; the behaviour described is read from the Go source and Docker’s published documentation at that commit, not from instrumented runs.


